WO2024239284A1 - 时空转录组切片的配准方法及装置 - Google Patents
时空转录组切片的配准方法及装置 Download PDFInfo
- Publication number
- WO2024239284A1 WO2024239284A1 PCT/CN2023/096109 CN2023096109W WO2024239284A1 WO 2024239284 A1 WO2024239284 A1 WO 2024239284A1 CN 2023096109 W CN2023096109 W CN 2023096109W WO 2024239284 A1 WO2024239284 A1 WO 2024239284A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- spot
- slice
- spots
- type
- matrix
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/30—Determination of transform parameters for the alignment of images, i.e. image registration
- G06T7/33—Determination of transform parameters for the alignment of images, i.e. image registration using feature-based methods
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/30—Detection of binding sites or motifs
Definitions
- the present disclosure relates to the field of data processing technology, and in particular to a method and device for registering spatiotemporal transcriptome slices.
- PSTE Probabilistic Alignment of Spatial Transcriptomics Experiments
- Each spot point in the PASTE algorithm has the same weight, and in every two pieces of spatiotemporal expression data, the proportion of structural information and feature information in the objective function is pre-set and cannot be automatically adjusted.
- the present invention provides a method and device for registering spatiotemporal transcriptome slices, an electronic device and a storage medium.
- a method for registering spatiotemporal transcriptome slices comprising:
- each slice contains at least one type of expression matrix
- each expression matrix contains multiple site spots of spatiotemporal group expression data
- each slice contains at least one type of spot
- the probability transfer matrix is calculated respectively to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the method further includes:
- respectively determining the weights corresponding to the spots of the same type in each slice includes:
- the corresponding expression matrix is added along the spot points to obtain the expression matrix corresponding to different types;
- the distance between the expression matrices of each type of spot on the two slices is calculated, and the weight corresponding to each type of spot is determined according to the distance.
- the method further includes:
- the non-highly variable genes in each slice were filtered and the high-variable genes in each slice were retained.
- the method before determining the weight corresponding to each type of spot according to the distance, the method further includes:
- the compressed distance value and the distance value less than the preset distance threshold are used as the final distance of the expression matrix of the spot on the two slices;
- Determining the weight corresponding to each type of spot according to the distance includes:
- the inverse of the final distance is calculated, and the weight corresponding to each type of spot is determined according to the calculation result of the inverse number.
- the method further includes:
- Each probability in the probability transfer matrix is divided by the sum of the probability transfer matrices to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the calculating the probability transmission matrix according to the weight corresponding to each type of spot and the preset regularization coefficient includes:
- a probability matrix with an accuracy score higher than a preset score threshold is selected from different probability transfer matrices as the final probability transfer matrix.
- a device for registering spatiotemporal transcriptome slices comprising:
- An acquisition unit used for acquiring two adjacent slices, each slice comprising at least one type of expression matrix, each expression matrix comprising multiple spots of spatiotemporal group expression data, and each slice comprising at least one type of spot;
- a determination unit used to determine the weights corresponding to the same type of spots in each slice, wherein the same type of spots are assigned the same weight
- the registration unit is used to calculate the probability transfer matrix according to the weight corresponding to each type of spot and the preset regularization coefficient, and obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the device further comprises:
- the processing unit is used to remove the spots that are not common to the two adjacent slices before obtaining the two adjacent slices.
- the determining unit includes:
- the summation module is used to sum the corresponding expression matrices along the spot points for the same type of spots in each slice, and obtain the expression matrices corresponding to different types;
- a calculation module used to calculate the distance between the expression matrix of each type of spot on the two slices
- the first determination module is used to determine the weight corresponding to each type of spot according to the distance.
- the determining unit further includes:
- the filtering module is used to filter the corresponding expression matrix along the same type of spot in each slice. Before adding the spot points to obtain the expression matrix corresponding to different types, the non-highly variable genes in each slice are filtered to retain the high-variable genes in each slice.
- the determining unit further includes:
- a search module used to search for a distance value greater than a preset distance threshold before determining the weight corresponding to each type of spot according to the distance;
- a processing module used for calling a preset processing function to compress the distance value greater than the preset distance threshold
- a second determination module is used to use the distance value after compression processing and the distance value less than the preset distance threshold as the final distance of the expression matrix of the spot on the two slices;
- the first determination module is further used to calculate the inverse of the final distance, and determine the weight corresponding to each type of spot according to the calculation result of the inverse.
- the device further comprises:
- a search unit configured to calculate the probability transfer matrix according to the weight corresponding to each type of spot and the preset regularization coefficient in the registration unit, obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices, and then search for the probability less than a preset threshold in the probability transfer matrix;
- a setting unit used to set the probability that the value is less than a preset threshold to 0;
- Each probability in the probability transfer matrix is divided by the sum of the probability transfer matrices to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the registration unit comprises:
- a probability matrix with an accuracy score higher than a preset score threshold is selected from different probability transfer matrices as the final probability transfer matrix.
- an electronic device including:
- the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
- a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
- a computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
- the present invention provides a registration method for spatiotemporal transcriptome slices, wherein two adjacent slices are obtained, each slice comprising at least one type of expression matrix, each expression matrix comprising multiple site spots of spatiotemporal group expression data, and each slice comprising at least one type of spot; weights corresponding to spots of the same type in each slice are determined respectively, wherein spots of the same type are assigned the same weight; and probability transmission matrices are calculated respectively according to the weights corresponding to each type of spot and a preset regularization coefficient, so as to obtain the weights corresponding to each category on the two slices.
- the registration score of each spot in the expression matrix is obtained by two adjacent slices, each slice comprising at least one type of expression matrix, each expression matrix comprising multiple site spots of spatiotemporal group expression data, and each slice comprising at least one type of spot; weights corresponding to spots of the same type in each slice are determined respectively, wherein spots of the same type are assigned the same weight; and probability transmission matrices are calculated respectively according to the weights corresponding to each type
- the embodiment of the present disclosure assigns the same weight to the same type of spots in each slice, and different types of spots are assigned different weights, which improves the accuracy of the probability transfer matrix calculated based on the weights, and further improves the accuracy of the registration achieved based on the probability transfer matrix.
- FIG1 is a schematic diagram of a flow chart of a method for registering spatiotemporal transcriptome slices provided by an embodiment of the present disclosure
- FIG4 is a schematic diagram of the structure of a device for registering spatiotemporal transcriptome slices provided by an embodiment of the present disclosure
- FIG. 6 is a schematic block diagram of an example electronic device provided by an embodiment of the present disclosure.
- FIG1 is a schematic flow chart of a method for registering spatiotemporal transcriptome slices provided in an embodiment of the present disclosure.
- the method comprises the following steps:
- Step 101 obtain two adjacent slices, each slice contains at least one type of expression matrix, each expression matrix contains multiple spots of spatiotemporal group expression data, and each slice contains at least one type of spot.
- the method described in the embodiment of the present application is not only applicable to the registration between two adjacent slices, but in a registration calculation process, two slices need to be registered in pairs. For example, there are 5 slices that need to be registered.
- the current calculation is the registration between slice 1 and the adjacent slice 2.
- the registration between slice 2 and the adjacent slice 3 is continued, and then the registration between slice 3 and the adjacent slice 4 is performed, and finally the registration between slice 4 and the adjacent slice 5 is performed.
- the above examples are only for the convenience of quantitative description, and the number of slice registrations is not limited in actual applications.
- the spot can represent a cell or an expression detection area of n site width multiplied by n site height that is artificially divided on the spatiotemporal chip site and is close to the size of a cell.
- Step 102 respectively determine the weights corresponding to the spots of the same type in each slice, wherein the spots of the same type are assigned the same weights.
- spots of the same type have the same weight value, which depends on the similarity between the expression of the type on one slice and the expression on another slice, that is, the greater the similarity, the greater the weight value, and the smaller the similarity, the smaller the weight value.
- Step 101 When determining whether the types are the same, the clustering result or annotation result of the expression data in step 101 can be used to determine.
- the description of the clustering algorithm is not repeated in this embodiment of the present application.
- Step 103 according to the weight corresponding to each type of spot and the preset regularization coefficient, the probability transfer matrix is calculated respectively to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the embodiment of the present disclosure introduces an automatic parameter adjustment mechanism: in the calculation, the algorithm uses different regularization coefficients with exponential changes in turn, calculates the probability transfer matrix respectively, and obtains its corresponding accuracy score. After using all the regularization coefficients specified for calculation in turn, the probability transfer matrix with the highest accuracy score is selected as the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the present disclosure provides a registration method for spatiotemporal transcriptome slices, obtaining two adjacent slices, each slice containing at least one type of expression matrix, each expression matrix containing multiple site spots of spatiotemporal group expression data, and each slice containing at least one type of spot; respectively determining the weights corresponding to the same type of spots in each slice, wherein the same type of spots are assigned the same weight, and different types of spots are assigned different weights; respectively calculating the probability transfer matrix according to the weight corresponding to each type of spot and the preset regularization coefficient, and obtaining the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the embodiment of the present disclosure assigns the same weight to the same type of spots in each slice, and different types of spots are assigned different weights, thereby improving the accuracy of the probability transfer matrix obtained based on the weight calculation, and thus improving the accuracy of the registration achieved based on the probability transfer matrix.
- the present application example also provides a schematic flow chart of another method for registering spatiotemporal transcriptome slices.
- the method further includes:
- Step 201 remove the spots that are not common to two adjacent slices.
- an operation of removing the incommon spot types is added, that is, the spots that appear only on one piece of the expression data between two pieces are removed to prevent them from becoming noise and interfering with the calculation of the probability transfer matrix.
- Step 202 obtaining two adjacent slices, each slice comprising at least one type of expression matrix, each expression matrix comprising multiple spots of spatiotemporal group expression data, and each slice comprising at least one type of spot;
- step 202 please refer to the detailed description of step 101 .
- Step 203 respectively determining the weights corresponding to the same type of spots in each slice, wherein the same type of spots are assigned the same weight, and different types of spots are assigned different weights;
- Step 2031 filtering the non-highly variable genes in each slice, and retaining the high-variable genes in each slice.
- spot gene expression matrix For the spot gene expression matrix, select and retain highly variable genes and remove other genes that are not highly variable genes. The purpose is to reduce the impact of non-highly variable genes on the calculation of the similarity of spot classes on different slices, making the calculation results more accurate.
- the gene expression matrix records how many entries of each gene are captured in each spot. Therefore, the number of rows in the gene expression matrix is equal to the number of spots, and the number of columns is equal to 10,000.
- the following three steps are performed in order: (1) standardization, (2) calculation of the variance of each gene, and (3) sorting and screening.
- Step 2032 for the spots of the same type in each slice, the corresponding expression matrices are summed along the spot points to obtain expression matrices corresponding to different types;
- the obtained expression matrix is log-normalized to control the scale effect of the data, that is, to reduce the problem that the expression matrices of each spot type are not on the same scale due to the uneven sequencing depth of different capture sites on the chip and the uneven number of spot components in different spot groups.
- spot points there are two categories of spot points: category 1 and category 2 (two points in one category), and 3 genes a, b, and c: count the two spots on category 1, and see how many a genes are expressed in total; count the two spots on category 1, and see how many b genes are expressed in total; count the two spots on category 1, and see how many c genes are expressed in total; count the two spots on category 2, and see how many a genes are expressed in total; count the two spots on category 2, and see how many b genes are expressed in total; count the two spots on category 2, and see how many c genes are expressed in total.
- logarithmic normalization is used to control the scale effect of the expression of the spot class.
- Other methods of eliminating/controlling scale effects may also be used, such as minimum-maximum normalization, etc.
- minimum-maximum normalization etc.
- the category information obtained by clustering and annotating different slices is based on the same reference system: for example, the first slice has two types of spots, a and b, and the second slice has three types of spots, a, b, and c. Then, it is necessary to count the distance between a of the first slice and a of the second slice, and to count the distance between b of the first slice and b of the second slice, without counting the distance between c.
- the preset processing function can also be other functions whose derivatives decrease in the range of [R, + ⁇ ) (where R represents any real number less than positive infinity, + ⁇ represents positive infinity), such as a logarithmic function; specifically, the embodiments of the present disclosure are not limited to this.
- Step 2035 calculate the inverse of the final distance, and determine the weight corresponding to each type of spot according to the calculation result of the inverse.
- the distance value processed in step 2034 is calculated as the opposite number to obtain the weight, and after the weight is shifted to the minimum value of 0, each weight is divided by the sum of all weights.
- Step 204 according to the weight corresponding to each type of spot and the preset regularization coefficient, the probability transfer matrix is calculated respectively to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the MAB matrix stores the similarity of the expression levels of every two spots between the two slices
- the C1 matrix stores the distances between all points in the first slice of the two slices used for calculation
- the C2 matrix stores the distances between all points in the second slice of the two slices used for calculation.
- the MAB matrix stores the similarity of the expression of each two spots between the two slices
- the C1 matrix stores the distance between all points in the first slice of the two slices used for calculation
- the C2 matrix stores the distance between all points in the second slice of the two slices used for calculation.
- the distances between all points in the image are stored in the vectors a and b, respectively, which are the weights corresponding to different cells in the two slices.
- the Fused Gromov-Wasserstein version of the OT algorithm used by the PASTE algorithm will iteratively calculate the probability transfer matrix based on the above matrices and vectors, as well as a fixed regularization coefficient alpha (which is essentially the weight ratio of the structural information item and the expression information item in the cost function).
- the above matrices, vectors and coefficients are collectively referred to as input parameters.
- the user needs to specify an alpha value, and then the OT algorithm will perform calculations based on this fixed value. After multiple iterations, the data stored in the matrix tends to be stable. The probability transfer matrix calculation for these two slices is completed. Then the calculation of the next pair of slices begins.
- the calculation problem of the probability transmission matrix is formed into a mathematical form of an optimization problem, including an optimization objective function and constraint conditions.
- the objective function is:
- MAB matrix corresponds to MAB in the formula
- C1 matrix corresponds to C1 in the formula
- C2 matrix corresponds to C2 in the formula
- a vector corresponds to h in the formula
- b vector corresponds to g in the formula.
- MAB matrix corresponds to MAB in the formula
- C1 matrix corresponds to C1 in the formula
- C2 matrix corresponds to C2 in the formula
- a vector corresponds to ⁇ X in the formula
- b vector corresponds to ⁇ Y in the formula.
- the embodiment of the present application introduces an automatic parameter adjustment mechanism: in the calculation, the algorithm uses different regularization coefficients with exponential changes in turn to calculate the probability transfer matrix and obtain its corresponding accuracy score. After using all the regularization coefficients specified for calculation in turn, the probability transfer matrix with the highest accuracy score is selected as the calculation result.
- Step 205 Find the probability less than a preset threshold in the probability transmission matrix, and set the probability less than the preset threshold to 0.
- This step is also called the post-processing process of the probability transmission matrix.
- the element values (probabilities) whose element values are lower than the preset threshold are uniformly set to 0.
- the calculation of the preset threshold is not limited to the above calculation; it can also be a certain percentage of the maximum probability value as the threshold, etc.
- other methods can also be used, such as clustering the paired spots represented by each spot in the probability transfer matrix result according to the coordinates, retaining the class with a large number of spots, removing the class with a small number, etc.
- Step 206 Divide each probability in the probability transfer matrix by the sum of the probability transfer matrices to obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the present invention also proposes a spatiotemporal transcriptome slice registration device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be repeated in the present invention.
- FIG4 is a schematic diagram of the structure of a spatiotemporal transcriptome slice registration device provided by an embodiment of the present disclosure, as shown in FIG4 , comprising:
- An acquisition unit 41 is used to acquire two adjacent slices, each slice contains at least one type of expression matrix, each expression matrix contains multiple spots of spatiotemporal group expression data, and each slice contains at least one type of spot;
- a determination unit 42 configured to respectively determine weights corresponding to spots of the same type in each slice, wherein spots of the same type are assigned the same weight, and spots of different types are assigned different weights;
- the registration unit 43 is used to calculate the probability transfer matrix according to the weight corresponding to each type of spot and the preset regularization coefficient, and obtain the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the present disclosure provides a registration device for spatiotemporal transcriptome slices, obtaining two adjacent slices, each slice containing at least one type of expression matrix, each expression matrix containing multiple site spots of spatiotemporal group expression data, and each slice containing at least one type of spot; respectively determining the weights corresponding to the same type of spots in each slice, wherein the same type of spots are assigned the same weight, and different types of spots are assigned different weights; respectively calculating the probability transfer matrix according to the weight corresponding to each type of spot and the preset regularization coefficient, and obtaining the registration score of each spot in the expression matrix corresponding to each category on the two slices.
- the embodiment of the present disclosure assigns the same weight to the same type of spots in each slice, and different types of spots are assigned different weights, thereby improving the accuracy of the probability transfer matrix obtained based on the weight calculation, and thus improving the accuracy of the registration achieved based on the probability transfer matrix.
- the device further includes:
- the processing unit 44 is used to remove spots that are not common to the two adjacent slices before acquiring the two adjacent slices.
- the determining unit 42 includes:
- the summing module 421 is used to sum the corresponding expression matrix along the spot points for the same type of spots in each slice, and obtain expression matrixes corresponding to different types;
- a calculation module 422 is used to calculate the distance between the expression matrix of each type of spot on the two slices;
- the first determination module 423 is used to determine the weight corresponding to each type of spot according to the distance.
- the determining unit 42 further includes:
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Theoretical Computer Science (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Evolutionary Biology (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Analytical Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Chemical & Material Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- General Physics & Mathematics (AREA)
- Image Processing (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
本公开提供一种时空转录组切片的配准方法及装置,主要技术方案包括:获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot;分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同(102);根据每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数(103)。与相关技术相比,为每张切片中同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同,提高了基于权重计算得到的概率转移矩阵的准确性,进而提高了基于概率转移矩阵所实现的配准的准确性。
Description
本公开涉及数据处理技术领域,尤其涉及一种时空转录组切片的配准方法及装置。
空间转录组实验的概率配准(Probabilistic Alignment of Spatial Transcriptomics Experiments,PASTE),是一个对来及相邻或相近切片的时空组表达量数据的位点(也称spot),spot既可以表示细胞,也可以表示和细胞大小接近的、人工于时空芯片位点上划分的n个位点宽乘以n个位点高的表达量检测区域之间进行配准的算法。
该PASTE算法中每一个spot点具有相同的权重,且每两片时空表达量数据中,结构信息和特征信息的在目标函数中的占比是预先设定的、无法自动调节的。
发明内容
本公开提供了一种时空转录组切片的配准方法及装置、电子设备和存储介质。
根据本公开的第一方面,提供了一种时空转录组切片的配准方法,包括:
获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;
分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重;
根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
可选地,在获取两张相邻切片之前,所述方法还包括:
去掉两张相邻切片中的不共具的spot。
可选地,所述分别确定每张切片中同一类型的spot对应的权重包括:
对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵;
计算每一类型spot分别在所述两张切片上的表达量矩阵的距离,并根据所述距离确定每一类型spot对应的权重。
可选地,在对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵之前,所述方法还包括:
对每张切片中的非高变基因进行过滤,保留每张切片中的高变基因。
可选地,在根据所述距离确定每一类型spot对应的权重之前,所述方法还包括:
查找大于预设距离阈值的距离值;
调用预设处理函数,将所述大于所述预设距离阈值的距离值进行压缩处理;
将压缩处理后的距离值与小于所述预设距离阈值的距离值,作为所述spot分别在所述两张切片上的表达量矩阵的最终距离;
所述根据所述距离确定每一类型spot对应的权重包括:
计算所述最终距离的相反数,根据相反数计算结果确定每一类型spot对应的权重。
可选地,在所述根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数之后,所述方法还包括:
查找所述概率传输矩阵中小于预设阈值的概率,并将所述小于预设阈值的概率设置为0;
将所述概率传输矩阵中的每个概率,除以所述概率传输矩阵之和,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
可选地,所述根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵包括:
基于不同的预设正则化系数,分别计算对应的概率传输矩阵;
从不同的概率传输矩阵中选择一个准确性分数高于预设分数阈值的概率矩阵作为最终的概率传输矩阵。
第二方面,还提供一种时空转录组切片的配准装置,包括:
获取单元,用于获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;
确定单元,用于分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重;
配准单元,用于根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
可选地,所述装置还包括:
处理单元,用于在获取两张相邻切片之前,去掉两张相邻切片中的不共具的spot。
可选地,所述确定单元,包括:
加和模块,用于对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵;
计算模块,用于计算每一类型spot分别在所述两张切片上的表达量矩阵的距离;
第一确定模块,用于根据所述距离确定每一类型spot对应的权重。
可选地,所述确定单元还包括:
过滤模块,用于在对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿
着spot点进行加和,分别得到不同类型对应的表达量矩阵之前,对每张切片中的非高变基因进行过滤,保留每张切片中的高变基因。
可选地,所述确定单元还包括:
查找模块,用于在根据所述距离确定每一类型spot对应的权重之前,查找大于预设距离阈值的距离值;
处理模块,用于调用预设处理函数,将所述大于所述预设距离阈值的距离值进行压缩处理;
第二确定模块,用于将压缩处理后的距离值与小于所述预设距离阈值的距离值,作为所述spot分别在所述两张切片上的表达量矩阵的最终距离;
所述第一确定模块,还用于计算所述最终距离的相反数,根据相反数计算结果确定每一类型spot对应的权重。
可选地,所述装置还包括:
查找单元,用于在所述配准单元根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数之后,查找所述概率传输矩阵中小于预设阈值的概率;
设置单元,用于将所述小于预设阈值的概率设置为0;
将所述概率传输矩阵中的每个概率,除以所述概率传输矩阵之和,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
可选地,所述配准单元包括:
基于不同的预设正则化系数,分别计算对应的概率传输矩阵;
从不同的概率传输矩阵中选择一个准确性分数高于预设分数阈值的概率矩阵作为最终的概率传输矩阵。
根据本公开的第三方面,提供了一种电子设备,包括:
至少一个处理器;以及
与所述至少一个处理器通信连接的存储器;其中,
所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行前述第一方面所述的方法。
根据本公开的第四方面,提供了一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使所述计算机执行前述第一方面所述的方法。
根据本公开的第五方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现如前述第一方面所述的方法。
本公开提供一种时空转录组切片的配准方法,获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重;根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的
表达量矩阵中每个spot的配准分数。与相关技术相比,本公开实施例为每张切片中同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同,提高了基于权重计算得到的概率转移矩阵的准确性,进而提高了基于概率转移矩阵所实现的配准的准确性。
应当理解,本部分所描述的内容并非旨在标识本申请的实施例的关键或重要特征,也不用于限制本申请的范围。本申请的其它特征将通过以下的说明书而变得容易理解。
附图用于更好地理解本方案,不构成对本公开的限定。其中:
图1为本公开实施例所提供的一种时空转录组切片的配准方法的流程示意图;
图2为本公开实施例所提供的另一种时空转录组切片的配准方法的流程示意图;
图3为本公开实施例所提供的一种确认每张切片中同一类型的spot对应的权重方法的流程示意图;
图4为本公开实施例提供的一种时空转录组切片的配准装置的结构示意图;
图5为本公开实施例提供的另一种时空转录组切片的配准装置的结构示意图;
图6为本公开实施例提供的示例电子设备的示意性框图。
以下结合附图对本公开的示范性实施例做出说明,其中包括本公开实施例的各种细节以助于理解,应当将它们认为仅仅是示范性的。因此,本领域普通技术人员应当认识到,可以对这里描述的实施例做出各种改变和修改,而不会背离本公开的范围和精神。同样,为了清楚和简明,以下的描述中省略了对公知功能和结构的描述。
下面参考附图描述本公开实施例的时空转录组切片的配准方法及装置、电子设备和存储介质。
图1为本公开实施例所提供的一种时空转录组切片的配准方法的流程示意图。
如图1所示,该方法包含以下步骤:
步骤101,获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot。
需要说明的是,本申请实施例所述的方法,并非是指仅适用于两张相邻切片之间的配准,而是在一次配准计算过程中,需要对两张切片进行两两配准,例如,需要进行配准计算的切片有5张,当前计算的为切片1和相邻的切片2之间的配准,在执行完切片1和切片2的配准后,继续执行切片2和相邻的切片3之间的配准,再执行切片3和相邻的切片4之间的配准,最后执行切片4和相邻的切片5之间配准。以上示例仅为便于给出的数量说明,在实际应用中对切片配准的数量不进行限定。
需要说明的是,本申请实施例所述的表达量数据具备聚类或注释结果,以辅助算法的实施;事实上,出于后期分析和展示的需要,绝大部分做spot配准的数据均具备聚类或注释结果。所述spot既可以表示细胞,也可以表示和细胞大小接近的、人工于时空芯片位点上划分的n个位点宽乘以n个位点高的表达量检测区域。
步骤102,分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重。
本申请实施例中,对于相同类型的spot具有相同的权重值,该值取决于该类型在一张切片上的表达量和在另一张切片上的表达量的相似度,即相似度约大,权重值越大,相似度越小,权重值越小。
在判定类型是否相同时,可以有步骤101中表达量数据的聚类结果或注释结果确定。有关聚类算法的说明,本申请实施例在此不再进行赘述。步骤103,根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
为提升配准算法的准确性,本公开实施例引入了自动化调参机制:在计算中,算法依次采用指数变化的不同的正则化系数,分别计算概率传输矩阵,并得到其对应的准确性分数。在轮流使用过所有规定参与计算的正则化系数后,选择准确性分数最高的概率传输矩阵,作为两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
本公开提供一种时空转录组切片的配准方法,获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同;根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。与相关技术相比,本公开实施例为每张切片中同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同,提高了基于权重计算得到的概率转移矩阵的准确性,进而提高了基于概率转移矩阵所实现的配准的准确性。
本申请实施例还提供另一种时空转录组切片的配准方法的流程示意图。
如图2所示,所述方法还包括:
步骤201,去掉两张相邻切片中的不共具的spot。
本申请实施例中增加了去掉不共具的spot类型的操作。将两片表达量数据间,只在一片上出现的spot进行去除,防止其成为噪音,干扰概率传输矩阵的计算。
步骤202,获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;
有关步骤202可参见步骤101的详细说明。
步骤203,分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同;
在确认每张切片中同一类型的spot对应的权重时,可以采用但不局限于以下方式,如图3所示,包括:
步骤2031,对每张切片中的非高变基因进行过滤,保留每张切片中的高变基因。
对于spot基因表达量矩阵,选取和保留高变基因,去掉非高变基因的其他基因。其目的在于削减非高变基因对于计算spot类在不同切片上相似性的影响,使得计算结果更准确。
为了便于对高变基因进行理解,示例性的,某张切片上,一共捕获了10000种不同的基因,基因表达量矩阵上记录了每个spot上,每种基因各自被捕获了多少条,因此,该基因表达量矩阵的行数等于spot个数,列数等于10000。为筛选高变基因,按顺序做:(1)标准化,(2)计算每个基因的方差,(3)排序并筛选这三个步骤。首先,以z-score标准化为例介绍标准化:基因表达量矩阵上,对每一行,计算该行上的所有值的平均值和标准差,让每个值都分别减去该平均值,再除以该标准差;每一行均做完此计算,就完成了步骤(1),注意该矩阵的值由此被更新;之后,对于这10000个基因,计算每个基因的列上的所有数据的方差,作为该基因的方差,就完成了步骤(2);最后,对于这10000个基因,根据(2)计算出的方差将它们进行从大到小的排序,取排名靠前的若干个,就得到了高变基因。这10000个基因中,不是高变基因的,即非高变基因。具体应用该过程中若干个包含但不限于10、20、50、100、200、500等,本申请实施例对具体数量不做限定。
步骤2032,对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵;
对每一张切片上的每一类spot,将其表达量矩阵沿着spot点进行加和,得到整个spot类上,不同基因的表达量矩阵。
对每一张切片上的每一类spot,对得到的表达量矩阵进行对数归一化,以控制数据的规模效应,即减少芯片不同捕获位点测序深度不均一,以及不同spot类群spot组成数目不均一,所带来的各个spot类表达量矩阵不在同一尺度上的问题。
示例性的,假设有4个spot点,spot点共两类:1类和2类(两个点一类),3个基因a,b,c:统计1类上的两个spot,总共表达了多少a基因,统计1类上的两个spot,总共表达了多少b基因,统计1类上的两个spot,总共表达了多少c基因,统计2类上的两个spot,总共表达了多少a基因,统计2类上的两个spot,总共表达了多少b基因,统计2类上的两个spot,总共表达了多少c基因。
在执行归一化处理时,使用对数归一化对spot类的表达量进行规模效应控制,也可以使用其他规模效应消除/控制的方法,例如最小-最大归一化等等,具体的本申请实施例对此不进行限定。
步骤2033,计算每一类型spot分别在所述两张切片上的表达量矩阵的距离,并查
找大于预设距离阈值的距离值。
计算每一类spot在两张切片上的表达量矩阵的距离,将此距离值赋给两张切片上对应spot的距离值。该距离包含但不限于KL散度距离、欧式距离、余弦相似度,詹森-香农距离(Jensen-Shannon,JS)距离,或其他的距离衡量方式。
示例性的,不同切片聚类和注释得到的类别信息,是基于相同的参考系:比如第一张切片,有a,b两类spot,第二张切片,有a,b,c三类spot,那么要统计第一张切片的a和第二张切片的a的距离,要统计第一张切片的b和第二张切片的b的距离,而无需在统计c的距离。
步骤2034,调用预设处理函数,将所述大于所述预设距离阈值的距离值进行压缩处理;将压缩处理后的距离值与小于所述预设距离阈值的距离值,作为所述spot分别在所述两张切片上的表达量矩阵的最终距离。
将各spot类的距离值经过预设处理函数(例如logistic函数)的处理,该函数在一定程度上压缩较大的距离值,保持较小的距离值,从而使下一步的权重计算更能突出具有小距离值spot类,即权重比较大的spot类之间的距离值差异,使对配准更重要的spot类得到更精细的权重分配;logistic函数如下所示:
需要说明的是,预设处理函数除了经过logistic函数外,也可以是其他导数在[R,+∞)范围内递减的函数(其中R表示的是任意一个小于正无穷的实数,+∞表示正无穷),例如对数函数等;具体的,本公开实施例对此不做限定。
步骤2035,计算所述最终距离的相反数,根据相反数计算结果确定每一类型spot对应的权重。
将经过步骤2034处理的距离值计算相反数,得到权重,使权重偏移至最小值为0后,使每个权重除以所有权重的和。过程如下:
weight=-1*dislogistic
weight=weight-min(weight)
weight=weight/Σweight
weight=-1*dislogistic
weight=weight-min(weight)
weight=weight/Σweight
步骤204,根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
基于不同的预设正则化系数,分别计算对应的概率传输矩阵,从不同的概率传输矩阵中选择一个准确性分数高于预设分数阈值的概率矩阵作为最终的概率传输矩阵。
对于每两片切片,首先,计算MAB,C1,C2矩阵:MAB矩阵中存储的是两片的之间每两个spot的表达量相似性,C1矩阵中存储的是用于计算的两片中第一片的所有点之间的距离,C2矩阵中存储的是用于计算的两片中第二片的所有点之间的距离。
对于每两片切片,首先,计算MAB矩阵,C1矩阵,C2矩阵,a向量,b向量:MAB矩阵中存储的是两片的之间每两个spot的表达量相似性,C1矩阵中存储的是用于计算的两片中第一片的所有点之间的距离,C2矩阵中存储的是用于计算的两片中第二片
的所有点之间的距离,a,b向量分别存储的是,这两张切片上不同细胞所对应的权重。
然后,PASTE算法使用的FusedGromov-Wasserstein版的OT算法,会依据上述众矩阵和向量,以及一个固定的正则化系数alpha(其本质上是结构信息项和表达量信息项在代价函数中的权重比例),迭代地计算概率传输矩阵。上述矩阵,向量和系数统称为输入参数。在PASTE算法中,用户需要指定一个alpha值,然后OT算法会依据该固定值,进行计算。多次迭代之后,该矩阵中存储的数据趋于稳定。这两片切片的概率传输矩阵计算结束。然后开始计算下一对切片。
在计算概率传输矩阵时,可以采用但不局限于以下方法,例如:将该概率传输矩阵的计算问题形成一个优化问题的数学形式,包括优化的目标函数,和约束条件。目标函数为:
where
约束条件为:
以上输入参数中,MAB矩阵对应公式中的MAB,C1矩阵对应公式中的C1,C2矩阵对应公式中的C2,a向量对应公式中的h,b向量对应公式中的g。
得到以上数学形式后,使用Fused Gromov Wasserstein的条件梯度算法进行求解。求解的流程如下:
其中,Eq.(7)为:
Eq.(6)为:
其中,
上述输入参数中,MAB矩阵对应公式中的MAB,C1矩阵对应公式中的C1,C2矩阵对应公式中的C2,a向量对应公式中的μX,b向量对应公式中的μY。
综上,为提升算法准确性,本申请实施例引入了自动化调参机制:在计算中,算法依次采用指数变化的,不同的正则化系数,分别计算概率传输矩阵,并得到其对应的准确性分数。在轮流使用过所有规定参与计算的正则化系数后,选择准确性分数最高的概率传输矩阵,作为计算结果。
步骤205,查找所述概率传输矩阵中小于预设阈值的概率,并将所述小于预设阈值的概率设置为0。
本步骤也称为概率传输矩阵的后处理过程,将元素值低于预设阈值的元素值(概率)统一设置为0,
预设阈值的计算方法如下:
thresh=min(pi[pi>0])+( max(pi)–min(pi[pi>0]))/10
thresh=min(pi[pi>0])+( max(pi)–min(pi[pi>0]))/10
实际应用中,对概率传输矩阵进行后处理时,预设阈值的计算不限于上述计算;也可以是取最大的概率值的某个百分比作为阈值等等。除了对比阈值的方法外,也可采用其他方法,如对每个spot在概率传输矩阵结果中表示的与之配对的spot依据坐标进行聚类,保留spot数量较多的类,去掉较少的类等等。
步骤206,将所述概率传输矩阵中的每个概率,除以所述概率传输矩阵之和,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
与上述的时空转录组切片的配准方法相对应,本发明还提出一种时空转录组切片的配准装置。由于本发明的装置实施例与上述的方法实施例相对应,对于装置实施例中未披露的细节可参照上述的方法实施例,本发明中不再进行赘述。
图4为本公开实施例提供的一种时空转录组切片的配准装置的结构示意图,如图4所示,包括:
获取单元41,用于获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;
确定单元42,用于分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同;
配准单元43,用于根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
本公开提供一种时空转录组切片的配准装置,获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同;根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。与相关技术相比,本公开实施例为每张切片中同一类型的spot分配相同的权重,且不同类型的spot分配的权重不同,提高了基于权重计算得到的概率转移矩阵的准确性,进而提高了基于概率转移矩阵所实现的配准的准确性。
在本实施例的一些方式中,如图5所示,所述装置还包括:
处理单元44,用于在获取两张相邻切片之前,去掉两张相邻切片中的不共具的spot。
在本实施例的一些方式中,如图5所示,所述确定单元42,包括:
加和模块421,用于对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵;
计算模块422,用于计算每一类型spot分别在所述两张切片上的表达量矩阵的距离;
第一确定模块423,用于根据所述距离确定每一类型spot对应的权重。
在本实施例的一些方式中,如图5所示,所述确定单元42还包括:
过滤模块424,用于在对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵之前,对每张切片中的非高变基因进行过滤,保留每张切片中的高变基因。
在本实施例的一些方式中,如图5所示,所述确定单元42还包括:
查找模块425,用于在根据所述距离确定每一类型spot对应的权重之前,查找大于预设距离阈值的距离值;
处理模块426,用于调用预设处理函数,将所述大于所述预设距离阈值的距离值进行压缩处理;
第二确定模块427,用于将压缩处理后的距离值与小于所述预设距离阈值的距离值,作为所述spot分别在所述两张切片上的表达量矩阵的最终距离;
所述第一确定模块423,还用于计算所述最终距离的相反数,根据相反数计算结果确定每一类型spot对应的权重。
在本实施例的一些方式中,如图5所示,所述装置还包括:
查找单元45,用于在所述配准单元44根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数之后,查找所述概率传输矩阵中小于预设阈值的概率;
设置单元46,用于将所述小于预设阈值的概率设置为0;
将所述概率传输矩阵中的每个概率,除以所述概率传输矩阵之和,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
在本实施例的一些方式中,如图5所示,所述配准单元44包括:
基于不同的预设正则化系数,分别计算对应的概率传输矩阵;
从不同的概率传输矩阵中选择一个准确性分数高于预设分数阈值的概率矩阵作为最终的概率传输矩阵。
需要说明的是,前述对方法实施例的解释说明,也适用于本实施例的装置,原理相同,本实施例中不再限定。
根据本公开的实施例,本公开还提供了一种电子设备、一种可读存储介质和一种计算机程序产品。
图6示出了可以用来实施本公开的实施例的示例电子设备500的示意性框图。电子设备旨在表示各种形式的数字计算机,诸如,膝上型计算机、台式计算机、工作台、个人数字助理、服务器、刀片式服务器、大型计算机、和其它适合的计算机。电子设备还可以表示各种形式的移动装置,诸如,个人数字处理、蜂窝电话、智能电话、可穿戴设备和其它类似的计算装置。本文所示的部件、它们的连接和关系、以及它们的功能仅仅作为示例,并且不意在限制本文中描述的和/或者要求的本公开的实现。
如图6所示,设备500包括计算单元501,其可以根据存储在ROM(Read-Only Memory,只读存储器)502中的计算机程序或者从存储单元508加载到RAM(Random Access Memory,随机访问/存取存储器)503中的计算机程序,来执行各种适当的动作和处理。在RAM 503中,还可存储设备500操作所需的各种程序和数据。计算单元501、ROM 502以及RAM 503通过总线504彼此相连。I/O(Input/Output,输入/输出)接口505也连接至总线504。
设备500中的多个部件连接至I/O接口505,包括:输入单元506,例如键盘、鼠标等;输出单元507,例如各种类型的显示器、扬声器等;存储单元508,例如磁盘、光盘等;以及通信单元509,例如网卡、调制解调器、无线通信收发机等。通信单元509允许设备500通过诸如因特网的计算机网络和/或各种电信网络与其他设备交换信息/数据。
计算单元501可以是各种具有处理和计算能力的通用和/或专用处理组件。计算单元501的一些示例包括但不限于CPU(Central Processing Unit,中央处理单元)、GPU(Graphic Processing Units,图形处理单元)、各种专用的AI(Artificial Intelligence,人工智能)计算芯片、各种运行机器学习模型算法的计算单元、DSP(Digital Signal Processor,数字信号处理器)、以及任何适当的处理器、控制器、微控制器等。计算单元501执行上文所描述的各个方法和处理,例如时空转录组切片的配准方法。例如,在一些实施例中,时空转录组切片的配准方法可被实现为计算机软件程序,其被有形地包含于机器可读介质,例如存储单元508。在一些实施例中,计算机程序的部分或者全部可以经由ROM 502和/或通信单元509而被载入和/或安装
到设备500上。当计算机程序加载到RAM 503并由计算单元501执行时,可以执行上文描述的方法的一个或多个步骤。备选地,在其他实施例中,计算单元501可以通过其他任何适当的方式(例如,借助于固件)而被配置为执行前述时空转录组切片的配准方法。
本文中以上描述的系统和技术的各种实施方式可以在数字电子电路系统、集成电路系统、FPGA(Field Programmable Gate Array,现场可编程门阵列)、ASIC(Application-Specific Integrated Circuit,专用集成电路)、ASSP(Application Specific Standard Product,专用标准产品)、SOC(System On Chip,芯片上系统的系统)、CPLD(Complex Programmable Logic Device,复杂可编程逻辑设备)、计算机硬件、固件、软件、和/或它们的组合中实现。这些各种实施方式可以包括:实施在一个或者多个计算机程序中,该一个或者多个计算机程序可在包括至少一个可编程处理器的可编程系统上执行和/或解释,该可编程处理器可以是专用或者通用可编程处理器,可以从存储系统、至少一个输入装置、和至少一个输出装置接收数据和指令,并且将数据和指令传输至该存储系统、该至少一个输入装置、和该至少一个输出装置。
用于实施本公开的方法的程序代码可以采用一个或多个编程语言的任何组合来编写。这些程序代码可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理器或控制器,使得程序代码当由处理器或控制器执行时使流程图和/或框图中所规定的功能/操作被实施。程序代码可以完全在机器上执行、部分地在机器上执行,作为独立软件包部分地在机器上执行且部分地在远程机器上执行或完全在远程机器或服务器上执行。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、RAM、ROM、EPROM(Electrically Programmable Read-Only-Memory,可擦除可编程只读存储器)或快闪存储器、光纤、CD-ROM(Compact Disc Read-Only Memory,便捷式紧凑盘只读存储器)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
为了提供与用户的交互,可以在计算机上实施此处描述的系统和技术,该计算机具有:用于向用户显示信息的显示装置(例如,CRT(Cathode-Ray Tube,阴极射线管)或者LCD(LiquidCrystal Display,液晶显示器)监视器);以及键盘和指向装置(例如,鼠标或者轨迹球),用户可以通过该键盘和该指向装置来将输入提供给计算机。其它种类的装置还可以用于提供与用户的交互;例如,提供给用户的反馈可以是任何形式的传感反馈(例如,视觉反馈、听觉反馈、或者触觉反馈);并且可以用
任何形式(包括声输入、语音输入或者、触觉输入)来接收来自用户的输入。
可以将此处描述的系统和技术实施在包括后台部件的计算系统(例如,作为数据服务器)、或者包括中间件部件的计算系统(例如,应用服务器)、或者包括前端部件的计算系统(例如,具有图形用户界面或者网络浏览器的用户计算机,用户可以通过该图形用户界面或者该网络浏览器来与此处描述的系统和技术的实施方式交互)、或者包括这种后台部件、中间件部件、或者前端部件的任何组合的计算系统中。可以通过任何形式或者介质的数字数据通信(例如,通信网络)来将系统的部件相互连接。通信网络的示例包括:LAN(Local Area Network,局域网)、WAN(Wide Area Network,广域网)、互联网和区块链网络。
计算机系统可以包括客户端和服务器。客户端和服务器一般远离彼此并且通常通过通信网络进行交互。通过在相应的计算机上运行并且彼此具有客户端-服务器关系的计算机程序来产生客户端和服务器的关系。服务器可以是云服务器,又称为云计算服务器或云主机,是云计算服务体系中的一项主机产品,以解决了传统物理主机与VPS服务("Virtual Private Server",或简称"VPS")中,存在的管理难度大,业务扩展性弱的缺陷。服务器也可以为分布式系统的服务器,或者是结合了区块链的服务器。
其中,需要说明的是,人工智能是研究使计算机来模拟人的某些思维过程和智能行为(如学习、推理、思考、规划等)的学科,既有硬件层面的技术也有软件层面的技术。人工智能硬件技术一般包括如传感器、专用人工智能芯片、云计算、分布式存储、大数据处理等技术;人工智能软件技术主要包括计算机视觉技术、语音识别技术、自然语言处理技术以及机器学习/深度学习、大数据处理技术、知识图谱技术等几大方向。
应该理解,可以使用上面所示的各种形式的流程,重新排序、增加或删除步骤。例如,本公开中记载的各步骤可以并行地执行也可以顺序地执行也可以不同的次序执行,只要能够实现本公开公开的技术方案所期望的结果,本文在此不进行限制。
上述具体实施方式,并不构成对本公开保护范围的限制。本领域技术人员应该明白的是,根据设计要求和其他因素,可以进行各种修改、组合、子组合和替代。任何在本公开的精神和原则之内所作的修改、等同替换和改进等,均应包含在本公开保护范围之内。
Claims (10)
- 一种时空转录组切片的配准方法,其特征在于,包括:获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重;根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
- 根据权利要求1所述的方法,其特征在于,在获取两张相邻切片之前,所述方法还包括:去掉两张相邻切片中的不共具的spot。
- 根据权利要求1所述的方法,其特征在于,所述分别确定每张切片中同一类型的spot对应的权重包括:对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵;计算每一类型spot分别在所述两张切片上的表达量矩阵的距离,并根据所述距离确定每一类型spot对应的权重。
- 根据权利要求3所述的方法,其特征在于,在对于每张切片中同一类型的spot,分别将对应的表达量矩阵沿着spot点进行加和,分别得到不同类型对应的表达量矩阵之前,所述方法还包括:对每张切片中的非高变基因进行过滤,保留每张切片中的高变基因。
- 根据权利要求3所述的方法,其特征在于,在根据所述距离确定每一类型spot对应的权重之前,所述方法还包括:查找大于预设距离阈值的距离值;调用预设处理函数,将所述大于所述预设距离阈值的距离值进行压缩处理;将压缩处理后的距离值与小于所述预设距离阈值的距离值,作为所述spot分别在所述两张切片上的表达量矩阵的最终距离;所述根据所述距离确定每一类型spot对应的权重包括:计算所述最终距离的相反数,根据相反数计算结果确定每一类型spot对应的权重。
- 根据权利要求1-5中任一项所述的方法,其特征在于,在所述根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数之后,所述方法还包括:查找所述概率传输矩阵中小于预设阈值的概率,并将所述小于预设阈值的概率设置为0;将所述概率传输矩阵中的每个概率,除以所述概率传输矩阵之和,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
- 根据权利要求6所述的方法,其特征在于,所述根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵包括:基于不同的预设正则化系数,分别计算对应的概率传输矩阵;从不同的概率传输矩阵中选择一个准确性分数高于预设分数阈值的概率矩阵作为最终的概率传输矩阵。
- 一种时空转录组切片的配准装置,其特征在于,包括:获取单元,用于获取两张相邻切片,每张切片包含至少一种类型的表达量矩阵,每个表达量矩阵中包含多个时空组表达量数据的位点spot,每张切片上包含至少一个类型的spot;确定单元,用于分别确定每张切片中同一类型的spot对应的权重,其中,同一类型的spot分配相同的权重;配准单元,用于根据所述每一类型spot对应的权重及预设正则化系数,分别计算概率传输矩阵,得到两张切片上每一类别对应的表达量矩阵中每个spot的配准分数。
- 一种电子设备,其特征在于,包括:至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1-7中任一项所述的方法。
- 一种存储有计算机指令的非瞬时计算机可读存储介质,其特征在于,所述计算机指令用于使所述计算机执行根据权利要求1-7中任一项所述的方法。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/096109 WO2024239284A1 (zh) | 2023-05-24 | 2023-05-24 | 时空转录组切片的配准方法及装置 |
| CN202380094961.0A CN120826740A (zh) | 2023-05-24 | 2023-05-24 | 时空转录组切片的配准方法及装置 |
| PCT/CN2023/120768 WO2024239503A1 (zh) | 2023-05-24 | 2023-09-22 | 时空转录组切片的配准方法及装置、电子设备和存储介质 |
| CN202380098639.5A CN121263842A (zh) | 2023-05-24 | 2023-09-22 | 时空转录组切片的配准方法及装置、电子设备和存储介质 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/096109 WO2024239284A1 (zh) | 2023-05-24 | 2023-05-24 | 时空转录组切片的配准方法及装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024239284A1 true WO2024239284A1 (zh) | 2024-11-28 |
Family
ID=93588788
Family Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/096109 Ceased WO2024239284A1 (zh) | 2023-05-24 | 2023-05-24 | 时空转录组切片的配准方法及装置 |
| PCT/CN2023/120768 Ceased WO2024239503A1 (zh) | 2023-05-24 | 2023-09-22 | 时空转录组切片的配准方法及装置、电子设备和存储介质 |
Family Applications After (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/120768 Ceased WO2024239503A1 (zh) | 2023-05-24 | 2023-09-22 | 时空转录组切片的配准方法及装置、电子设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (2) | CN120826740A (zh) |
| WO (2) | WO2024239284A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107194959A (zh) * | 2017-04-25 | 2017-09-22 | 北京海致网聚信息技术有限公司 | 基于切片进行图像配准的方法和装置 |
| CN107545567A (zh) * | 2017-07-31 | 2018-01-05 | 中国科学院自动化研究所 | 生物组织序列切片显微图像的配准方法及装置 |
| US20210035297A1 (en) * | 2019-07-31 | 2021-02-04 | The Joan and Irwin Jacobs Technion-Cornell Institute | System and method for region detection in tissue sections using image registration |
| CN113436238A (zh) * | 2021-08-27 | 2021-09-24 | 湖北亿咖通科技有限公司 | 点云配准精度的评估方法、装置和电子设备 |
| CN114387319A (zh) * | 2022-01-13 | 2022-04-22 | 北京百度网讯科技有限公司 | 点云配准方法、装置、设备以及存储介质 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10467767B2 (en) * | 2016-12-23 | 2019-11-05 | International Business Machines Corporation | 3D segmentation reconstruction from 2D slices |
| CN108038874B (zh) * | 2017-12-01 | 2020-07-24 | 中国科学院自动化研究所 | 面向序列切片的扫描电镜图像实时配准装置及方法 |
| CN110874849B (zh) * | 2019-11-08 | 2023-04-18 | 安徽大学 | 一种基于局部变换一致的非刚性点集配准方法 |
| CN113269813A (zh) * | 2021-06-09 | 2021-08-17 | 志诺维思(北京)基因科技有限公司 | 数字病理图像的配准方法和装置 |
| CN114817363B (zh) * | 2022-04-21 | 2026-01-02 | 北京百度网讯科技有限公司 | 确定数据相似性的方法、装置、电子设备以及存储介质 |
| CN115798587B (zh) * | 2022-10-19 | 2025-10-21 | 上海鹿明生物科技有限公司 | 一种空间转录组与空间代谢组空间信息匹配方法及系统 |
-
2023
- 2023-05-24 CN CN202380094961.0A patent/CN120826740A/zh active Pending
- 2023-05-24 WO PCT/CN2023/096109 patent/WO2024239284A1/zh not_active Ceased
- 2023-09-22 CN CN202380098639.5A patent/CN121263842A/zh active Pending
- 2023-09-22 WO PCT/CN2023/120768 patent/WO2024239503A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107194959A (zh) * | 2017-04-25 | 2017-09-22 | 北京海致网聚信息技术有限公司 | 基于切片进行图像配准的方法和装置 |
| CN107545567A (zh) * | 2017-07-31 | 2018-01-05 | 中国科学院自动化研究所 | 生物组织序列切片显微图像的配准方法及装置 |
| US20210035297A1 (en) * | 2019-07-31 | 2021-02-04 | The Joan and Irwin Jacobs Technion-Cornell Institute | System and method for region detection in tissue sections using image registration |
| CN113436238A (zh) * | 2021-08-27 | 2021-09-24 | 湖北亿咖通科技有限公司 | 点云配准精度的评估方法、装置和电子设备 |
| CN114387319A (zh) * | 2022-01-13 | 2022-04-22 | 北京百度网讯科技有限公司 | 点云配准方法、装置、设备以及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121263842A (zh) | 2026-01-02 |
| WO2024239503A1 (zh) | 2024-11-28 |
| CN120826740A (zh) | 2025-10-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020211242A1 (zh) | 一种基于行为识别的方法、装置及存储介质 | |
| US12437208B2 (en) | Picture searching method and apparatus, electronic device and computer readable storage medium | |
| CN113963197A (zh) | 图像识别方法、装置、电子设备和可读存储介质 | |
| WO2021003803A1 (zh) | 数据处理方法、装置、存储介质及电子设备 | |
| CN114463856A (zh) | 姿态估计模型的训练与姿态估计方法、装置、设备及介质 | |
| CN109146891B (zh) | 一种应用于mri的海马体分割方法、装置及电子设备 | |
| WO2023231184A1 (zh) | 一种特征筛选方法、装置、存储介质及电子设备 | |
| CN112131199B (zh) | 一种日志处理方法、装置、设备及介质 | |
| CN114021650A (zh) | 数据处理方法、装置、电子设备和介质 | |
| CN115359574A (zh) | 人脸活体检测及相应模型的训练方法、装置及存储介质 | |
| CN115171225A (zh) | 图像检测方法和图像检测模型的训练方法 | |
| WO2024239284A1 (zh) | 时空转录组切片的配准方法及装置 | |
| WO2025200320A1 (zh) | 一种训练数据的获取方法、装置及电子设备 | |
| CN118604777A (zh) | 激光点云噪点生成方法及装置、电子设备、存储介质 | |
| CN110929731A (zh) | 一种基于探路者智能搜索算法的医疗影像处理方法及装置 | |
| CN113807413B (zh) | 对象的识别方法、装置、电子设备 | |
| CN114092739B (zh) | 图像处理方法、装置、设备、存储介质和程序产品 | |
| CN115905263A (zh) | 向量数据库的更新方法及基于向量数据库的人脸识别方法 | |
| CN114494817A (zh) | 图像处理方法、模型训练方法、相关装置及电子设备 | |
| CN114494782A (zh) | 图像处理方法、模型训练方法、相关装置及电子设备 | |
| CN119667760B (zh) | 地震数据的处理方法、装置、电子设备及存储介质 | |
| CN112580802B (zh) | 网络模型压缩方法、装置 | |
| CN115455019B (zh) | 一种针对用户行为的分类模型更新方法、装置及设备 | |
| CN114707010A (zh) | 模型训练和媒介信息处理方法、装置、设备及存储介质 | |
| WO2025000230A1 (zh) | 无监督卷积神经网络模型的训练方法及其聚类方法、装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23937958 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202380094961.0 Country of ref document: CN |
|
| WWP | Wipo information: published in national office |
Ref document number: 202380094961.0 Country of ref document: CN |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |