WO2020199337A1 - 一种基因变异识别方法、装置和存储介质 - Google Patents
一种基因变异识别方法、装置和存储介质 Download PDFInfo
- Publication number
- WO2020199337A1 WO2020199337A1 PCT/CN2019/089504 CN2019089504W WO2020199337A1 WO 2020199337 A1 WO2020199337 A1 WO 2020199337A1 CN 2019089504 W CN2019089504 W CN 2019089504W WO 2020199337 A1 WO2020199337 A1 WO 2020199337A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- gene
- site
- base arrangement
- gene mutation
- base
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
Definitions
- the present disclosure relates to the field of computer technology, and in particular to a method, device and storage medium for identifying gene mutations.
- gene sequencing technology greatly improves the efficiency of gene sequencing, reduces the cost of gene sequencing, and maintains the accuracy of gene sequencing. If the first-generation testing technology completes the sequencing of a human genome, it may take three years, while the second-generation sequencing technology can shorten the time to just one week.
- the present disclosure proposes a technical solution for gene mutation identification.
- a method for identifying gene mutations comprising:
- the obtaining the base arrangement characteristics of the gene mutation candidate site includes:
- the base arrangement feature is used to characterize the base arrangement sequence.
- the determining the non-base arrangement characteristics of the gene mutation candidate site based on the non-base arrangement information of the at least one gene sequencing read in a preset site interval includes:
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
- the recognizing the gene variation of the gene variation candidate site based on the base arrangement feature and the non-base arrangement feature of the gene variation candidate site includes:
- a feature matrix of the gene mutation candidate site is obtained; wherein the first dimension feature of the feature matrix corresponds to the gene mutation candidate.
- the base arrangement feature and non-base arrangement feature of the site, the second dimension feature of the feature matrix corresponds to the site in the preset site interval; according to the feature matrix of the gene mutation candidate site, all Identify the genetic variation at the candidate site of the genetic variation.
- the identifying the gene mutation at the gene mutation candidate site according to the feature matrix of the gene mutation candidate site includes:
- the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
- the obtaining the feature matrix of the gene mutation candidate site according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site includes:
- the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
- obtaining at least one gene sequencing read corresponding to the gene mutation candidate site includes:
- a gene mutation identification device comprising:
- the first acquisition module is used to acquire at least one gene sequencing read corresponding to the gene mutation candidate site; the second acquisition module is used to acquire the base arrangement characteristics of the gene mutation candidate site; the determination module is used to The non-base arrangement information of the at least one gene sequencing read in the preset site interval determines the non-base arrangement feature of the gene mutation candidate site; wherein the non-base arrangement feature changes in the base arrangement sequence After that, it remains unchanged; the recognition module is used to identify the gene mutation of the gene mutation candidate site based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site.
- the second acquisition module includes:
- the first determining sub-module is used to determine the preset site interval where the candidate site of the gene mutation is located;
- the second determining sub-module is used to obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used for characterization The sequence of bases.
- the determining module includes:
- the first obtaining submodule is used to obtain the non-base arrangement information of each site in the preset site interval of the at least one gene sequencing read;
- the non-base arrangement information of each site in the site interval determines the non-base arrangement characteristics of the candidate site of the gene mutation.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
- the third determining submodule is specifically configured to:
- the third determining submodule is specifically configured to:
- the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
- the third determining submodule is specifically configured to:
- the third determining submodule is specifically configured to:
- the identification module includes:
- the generation sub-module is used to obtain the feature matrix of the gene mutation candidate site according to the base arrangement feature and non-base arrangement feature of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to The base arrangement feature and non-base arrangement feature of the gene mutation candidate site, the second dimension feature of the feature matrix corresponds to the site in the preset site interval; the identification sub-module is used for The feature matrix of the gene mutation candidate site is used to identify the gene mutation of the gene mutation candidate site.
- the identification submodule is specifically used for:
- the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
- the generating submodule is specifically used for:
- the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
- the first obtaining module includes:
- the second acquisition submodule is used to obtain the gene sequencing reads obtained by performing gene sequencing of somatic genes;
- the comparison submodule is used to compare the base sequence of the gene sequencing reads with the base sequence of the reference genome , Obtain the comparison result;
- the fourth determination sub-module used to determine the abnormal gene mutation candidate site of the gene of the somatic gene according to the comparison result;
- the third acquisition sub-module used to obtain the gene mutation At least one gene sequencing read corresponding to the candidate site.
- a gene mutation identification device including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above method.
- a non-volatile computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor.
- the gene mutation identification solution can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, and obtain the base arrangement characteristics of the gene mutation candidate site, based on the fact that at least one gene sequencing read is in a preset position
- the base arrangement information of the point interval determines the non-base arrangement characteristics of the gene mutation candidate site, so that the gene mutation of the gene mutation candidate site can be based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site Identify it.
- the non-base arrangement feature remains unchanged after the base arrangement sequence is changed, that is, it can be considered that the non-base arrangement feature has the nature of base arrangement invariance.
- the gene mutation of the candidate site of gene mutation is not restricted by the sequence of bases, and the pseudogene mutation caused by the germline gene mutation, noise, error, etc. can be better screened out, so that the gene mutation can be better To identify mutations to improve the accuracy of gene mutation identification,
- Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
- Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
- Fig. 3 shows a flowchart of the base arrangement characteristic process of gene mutation candidate sites according to an embodiment of the present disclosure.
- FIG. 4 shows a flowchart of the non-base arrangement feature process of gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 6 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 7 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
- FIG. 8 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
- the gene mutation identification scheme provided by the embodiments of the present disclosure can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, so that at least one gene sequencing read can be used to identify the gene mutation of the gene mutation candidate site.
- the base arrangement characteristics of the gene mutation candidate site can be determined, and the non-base of the gene mutation candidate site can be determined according to the base arrangement information of at least one gene sequencing read in the preset site interval
- the arrangement feature can then be used to identify the genetic variation at the candidate site of the gene variation through the base arrangement feature and the non-base arrangement feature.
- the non-base arrangement feature here remains unchanged after the base arrangement sequence is changed, that is, it can be considered that whether the genetic mutation at the genetic mutation candidate site is a true mutation is not affected by the base arrangement sequence, so that the When identifying genetic mutations, consider the invariance of the base arrangement of genetic data to improve the accuracy of genetic mutation identification.
- Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
- the gene mutation identification method can be executed by a gene mutation identification device or other processing equipment, where the gene mutation identification device can be User Equipment (UE), mobile equipment, user terminal, terminal, cellular phone, cordless phone, personal digital Processing (Personal Digital Assistant, PDA), handheld devices, computing devices, vehicle-mounted devices, wearable devices, etc., or the gene mutation recognition device may be a server.
- the gene mutation identification method may be implemented by a processor calling computer-readable instructions stored in a memory. As shown in Figure 1, the gene mutation identification method includes:
- Step 11 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
- the gene mutation recognition device can obtain the gene sequencing reads obtained by gene sequencing, and then obtain at least one gene sequencing read corresponding to the gene mutation candidate site from the gene sequencing reads obtained by the gene sequencing.
- the gene sequencing reads here can be understood as base sequences marked with base types after gene sequencing, and the length of each gene sequencing read can be the same or different. In the case of different lengths, the length of each gene sequencing read segment can be within a preset length range, thereby ensuring that the length of each gene sequencing read segment is relatively close.
- the base type may include cytosine (C), guanine (G), adenine (A), and thymine (T), so that the gene sequencing reads may include the base sequence of AGCT.
- the gene mutation candidate site may be a site with an abnormal base sequence.
- the site of the base sequence may indicate the position of the base sequence.
- there may be at least one gene sequencing read that is, at the same site, there may be at least one gene sequencing read obtained by gene sequencing.
- the gene mutation candidate site corresponds to at least one gene sequencing read segment, wherein the at least one gene sequencing read segment covers this locus.
- there may be at least one gene mutation candidate site and each gene mutation candidate site may correspond to at least one gene sequencing read.
- the embodiment of the present disclosure uses a gene mutation candidate site for description.
- Step 12 Obtain the base arrangement characteristics of the gene mutation candidate sites.
- the gene mutation recognition model can be used to extract the base arrangement characteristics of the gene mutation candidate site based on the gene arrangement information of the gene mutation candidate site.
- the base arrangement information here can be information related to the base arrangement sequence. For example, if the base sequence of a certain gene sequencing read in a certain site interval is A, C, G, T, then the base arrangement information Can be ACGT.
- the base arrangement information may include the base type of the reference genome in the preset site interval, the number of genes of each base type, the number of missing genes of each base type, the number of inserted genes of each base type, and so on.
- the base arrangement characteristics obtained from the base arrangement information are related to the base arrangement order.
- Step 13 based on the non-base arrangement information of the at least one gene sequencing read in the preset site interval, determine the non-base arrangement feature of the gene mutation candidate site; wherein, the non-base arrangement feature is The sequence of bases remains unchanged after changing.
- the base of at least one gene sequencing read corresponding to the gene mutation candidate site may be extracted in a preset site interval Arrange the information, and generate the non-base arrangement characteristics of the gene mutation candidate site according to the extracted base arrangement information.
- the non-base arrangement information may be information that is not restricted by the base arrangement order. Therefore, the non-base arrangement characteristics of the gene mutation candidate sites can be determined according to the non-base arrangement information of at least one gene sequencing read in the preset site interval.
- the non-base arrangement information may include information with base arrangement invariance such as the number of gene sequencing reads corresponding to the site, the number of gene sequencing reads that are mutated at the site, and the like.
- non-base arrangement information when extracting non-base arrangement information, several gene sequencing reads corresponding to the candidate site of the gene mutation can be randomly selected, and the non-base arrangement information of several randomly selected gene sequencing reads can be extracted; The non-base arrangement information of each gene sequencing read corresponding to the gene mutation candidate site.
- the non-base arrangement information of at least one gene sequencing read in the preset site interval When extracting the non-base arrangement information of at least one gene sequencing read in the preset site interval, the non-base arrangement information of at least one gene sequencing read in each site within the preset site interval can be extracted, and Several adjacent sites in the preset site interval can be randomly selected, and the non-base arrangement information of at least one gene sequencing read at several adjacent sites can be extracted.
- a gene mutation recognition model obtained based on neural network training can be used.
- Step 14 based on the base arrangement feature and non-base arrangement feature of the gene mutation candidate site, identify the gene mutation of the gene mutation candidate site.
- the characteristic matrix of the gene mutation candidate site can be obtained from the base arrangement characteristics and the non-base arrangement characteristics, and the characteristic matrix Recognition of gene mutations in gene mutation candidate sites.
- the above gene mutation recognition model can be used to determine whether the gene at the gene mutation candidate site is a true mutation caused by a disease or a base sequence abnormality caused by noise or other reasons. Of false mutations.
- the obtained feature matrix of the gene mutation candidate site can be a two-dimensional feature matrix, and the size of the feature matrix can be the number of feature vectors ⁇ the size of the preset site interval, and the feature vector can be based on the base arrangement feature And non-base alignment features.
- the gene mutation at the candidate site of the mutation is caused by the disease is not affected by the sequence of bases, and is more affected by the genetic environment where the candidate site of the mutation is located, for example, by the vicinity of the candidate site of the gene mutation
- the other sites are affected by genetic environment such as mutant genes, so the order of the feature vectors corresponding to the base arrangement feature in the resulting feature matrix can be unlimited, and the order of the feature vectors of the base arrangement feature in the feature matrix It can be changed randomly to improve the efficiency and accuracy of gene mutation identification.
- the gene mutation of the gene mutation candidate site can be identified according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site, so that the invariance of the base arrangement of the gene mutation can be considered, and the Identify genetic variation.
- the gene mutation candidate site at least one gene sequencing read corresponding to the gene mutation candidate site can be obtained.
- the example of the present disclosure also provides a process for obtaining at least one gene sequencing read corresponding to the gene mutation candidate site.
- Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
- obtaining at least one gene sequencing read corresponding to the gene mutation candidate site may include the following steps:
- Step 111 Obtain gene sequencing reads obtained by gene sequencing of somatic genes.
- At least one gene sequencing read segment can be obtained by performing gene sequencing of the somatic cell gene, and the gene sequencing read segment can be a sequence that annotates the base type of the somatic cell gene. After gene sequencing of somatic genes, not only the base type of each gene in the gene sequencing read, but also the gene location information of each gene in the gene sequencing read can be obtained. The same site can correspond to at least one gene sequencing read.
- At least one gene sequencing read can be obtained through gene sequencing of somatic genes, and the gene sequencing reads obtained by gene sequencing can be preprocessed.
- the preprocessing methods here can include cross contamination screening, Sequencing quality screening, comparison quality screening, abnormal read length screening, etc. Through preprocessing, cross-contaminated gene sequencing reads can be screened out, and gene sequencing reads with low sequencing quality and comparison quality and abnormal read length can be screened out.
- Step 112 Align the base sequence of the gene sequencing read with the base sequence of the reference genome to obtain an alignment result.
- the base sequence of the obtained gene sequencing reads can be compared with the base sequence of the reference genome at the same site , Get the comparison result. For example, you can compare the base sequence of each gene sequencing read obtained by gene sequencing with the base sequence of the reference genome at the same site to determine the base sequence of the gene sequencing read that is different from the base sequence of the reference genome. point. It is also possible to compare the base sequence of at least one gene sequencing read with the same site with the base sequence of the reference genome at the same site to determine the site where the base sequence of at least one gene sequencing read is different from the base sequence of the reference genome .
- the reference genome may be a base sequence labeled with a correct base sequence.
- Step 113 According to the comparison result, it is determined that the gene of the somatic gene has an abnormal gene mutation candidate site.
- the base sequence of the gene sequencing read and the reference genome can be determined according to the comparison result. If at least one gene sequencing read corresponding to the locus is in at least one gene sequencing read, the mutation is sent at that position. If the ratio of gene sequencing reads is greater than the preset ratio, it can be determined that the site is a candidate site for gene mutation; otherwise, it can be considered that the site is not a candidate site for gene mutation.
- the base sequence of the gene sequencing read at this position is different from that of the reference genome, which may be caused by sequencing errors. In this way, the abnormality of base sequence caused by gene sequencing errors can be reduced.
- Step 114 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
- At least one gene sequencing read corresponding to the candidate gene mutation site can be obtained.
- the base sequence of the gene mutation candidate site may be different from the base sequence of the reference genome at the same site.
- the base arrangement characteristics of the gene mutation candidate site can be determined according to the base arrangement information of at least one gene sequencing read corresponding to the gene mutation candidate site, so as to identify the gene mutation at the gene mutation candidate site.
- data enhancement processing can be performed on gene recognition based on the base arrangement characteristics. The following describes in detail the process of determining the base arrangement characteristics of gene mutation candidate sites through an example.
- Fig. 3 shows a flowchart of the base arrangement characteristic process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 3, the above step 12 may include the following steps:
- Step 121 Determine the preset site interval where the candidate site of gene mutation is located
- Step 122 Obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used to characterize the base arrangement sequence.
- the base arrangement information may include the base arrangement information of the candidate genome.
- the base arrangement information is the base arrangement information of the candidate genome, it can be considered that the base arrangement information of each gene sequencing read is the same. Base arrangement information of the candidate genome. Therefore, according to the gene location information of the gene mutation candidate site, the preset site interval where the gene mutation candidate site is located can be determined.
- the interval formed by 150 bases before and after the gene mutation candidate site can be used as the gene mutation candidate site.
- the preset site interval where the point is located.
- the base arrangement information of the reference genome in the preset site interval can be obtained, and gene mutation candidate positions can be generated from the base arrangement information of the reference genome in the preset site interval.
- the base arrangement information can refer to the base sequence composition of each site in the preset site interval of the genome.
- the preset site interval includes 4 base sequences, namely A, C, G, and T.
- the base arrangement information may be the base arrangement order of ACGT.
- the base arrangement feature can be represented by the base arrangement feature vector, which can be part of the feature matrix of the gene mutation candidate site.
- a1, a2, a3, and a4 can be the first 4-dimensional features of the feature matrix.
- the base arrangement characteristics corresponding to the gene mutation candidate sites are considered when identifying the gene mutation at the gene mutation candidate sites, but also the base arrangement characteristics of the gene mutation candidate sites are considered.
- Denatured non-base alignment features The following describes in detail the process of determining the non-base arrangement characteristics of gene mutation candidate sites through an example.
- FIG. 4 shows a flowchart of the non-base arrangement feature process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 4, the above step 13 may include the following steps:
- Step 131 Obtain non-base arrangement information of each site in the preset site interval of the at least one gene sequencing read;
- Step 132 based on the non-base arrangement information of each site in the preset site interval, determine the non-base arrangement feature of the gene mutation candidate site.
- Non-base arrangement information may be information with the invariance of base arrangement, for example, the number of gene sequencing reads and the number of mutations corresponding to the site.
- the non-base permutation feature generated by each type of non-base permutation information can form a non-base permutation feature vector, and there can be one or more non-base permutation feature vectors.
- the gene mutation identification scheme provided by the embodiments of the present disclosure can be applied to patients who have been diagnosed with cancer, and the gene mutation identification can guide the patient to take medication. Therefore, a part of the gene sequencing reads in the gene sequencing reads can be derived from normal cells, and normal cells can be considered as cells that are not diseased. Some gene sequencing reads can be derived from diseased cells. Therefore, when determining the non-base arrangement characteristics of gene mutation candidate sites, the non-base arrangement of gene mutation candidate sites can be determined based on gene sequencing reads derived from normal cells and gene sequencing reads derived from diseased cells. feature.
- the non-base arrangement characteristics of gene mutation candidate sites when determining the non-base arrangement characteristics of gene mutation candidate sites, it is possible to determine the gene sequencing reads derived from normal cells in at least one gene sequencing read, and then based on the gene sequencing reads of normal cells Segment the non-base arrangement information of each site in the preset site interval to determine the non-base arrangement characteristics of gene mutation candidate sites. In this way, the non-base arrangement characteristics of gene mutation candidate sites can be determined based on gene sequencing reads derived from normal cells.
- the following provides several examples of determining the non-base arrangement characteristics of gene mutation candidate sites based on the gene sequencing reads of normal cells.
- the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site is consistent with the base type of the reference genome. A gene sequencing read, and then according to the number of the first gene sequencing reads corresponding to each site in the preset site interval, the non-base arrangement characteristics of the gene mutation candidate sites are determined.
- the first gene sequencing read that has not undergone genetic mutation at the gene mutation candidate site can be selected, and for each site in the preset site interval, the first gene sequencing can be counted The number of reads at that location. In other words, you can count how many first gene sequencing reads contain the locus.
- the first gene sequencing read that includes a certain site can be considered as the first gene sequencing read corresponding to the site. Since the length of each gene sequencing read may be different, the position of the gene mutation candidate site relative to each gene sequencing read is different.
- the gene mutation candidate site can be located in the middle of the gene sequencing read, or it can be located in the gene The edge positions of the sequencing reads, so that the number of gene sequencing reads corresponding to each site in the preset site interval is different. From the number of first gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature can be generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding position. The number of first gene sequencing reads of the point.
- the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site is consistent with the base type of the reference genome First gene sequencing reads, and then at each position in the preset site interval, determine the number of first gene sequencing reads whose base types are inconsistent with those of the reference genome. As the variation quantity of the first gene sequencing read segment, the non-base arrangement characteristics of the gene variation candidate site are determined according to the variation quantity of the first gene sequencing read segment.
- the first gene sequencing read that has not undergone genetic mutation at the gene mutation candidate site can be selected, and for each site in the preset site interval, the first gene sequencing can be counted The number of genetic mutations in the read at this locus.
- gene sequencing reads did not undergo genetic mutation at the gene mutation candidate site (that is, the gene mutation candidate site is consistent with the base type of the reference genome), it may occur at other sites other than the gene mutation candidate site Gene mutation (that is, the base type is inconsistent with the reference genome at other sites), so that for each site in the preset site interval, the number of mutations in the first gene sequencing read of that site can be counted .
- the first gene sequencing reads in the gene sequencing reads of normal cells that have not been mutated at the gene mutation candidate site, and then targeting the preset site interval For each site, count the number of first gene sequencing reads corresponding to each site and the number of mutations at that site.
- These two pieces of information can correspond to the fifth-dimensional feature and the sixth-dimensional feature in the above feature matrix. Dimensional characteristics.
- the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site and the gene mutation candidate site The second gene sequencing reads with the same mutation base type of the points, and then according to the number of second gene sequencing reads corresponding to each site in the preset site interval, determine the non-uniformity of the gene mutation candidate site Base arrangement characteristics.
- the second gene sequencing reads that are consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads.
- the second gene sequencing reads can be counted The number at that site. From the number of second gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature is generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding site The number of second gene sequencing reads.
- the gene mutation candidate site and the gene mutation candidate site can be determined in the gene sequencing read. Sequencing reads of the second gene with the same variant base type, and then at each position in the preset site interval, determine the second gene whose base type of the second gene sequencing read is inconsistent with that of the reference genome The number of sequencing reads is used as the number of mutations of the second gene sequencing reads. According to the number of mutations of the second gene sequencing reads, the non-base arrangement characteristics of gene mutation candidate sites are determined.
- the second gene sequencing read that is consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads (the mutation base type of the gene mutation candidate site can be obtained through gene sequencing), and the preset position For each site in the point interval, count the number of mutations of the second gene sequencing read at that site, in other words, count the number of second gene sequencing reads that contain the site and have mutations at that site .
- the number of mutations in the second gene sequencing read corresponding to each site can generate a non-base arrangement feature vector corresponding to the non-base arrangement feature vector.
- Each feature element in the non-base arrangement feature vector can be The variation number of the second gene sequencing read corresponding to the corresponding locus.
- a second gene sequencing read that is consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads of normal cells, and then targeted at the preset site interval For each site, count the number of second gene sequencing reads corresponding to each site and the number of mutations at that site.
- the third gene sequencing read in the gene sequencing read can be determined, and then according to the preset site interval The number of third gene sequencing reads corresponding to each locus in, determines the non-base arrangement characteristics of gene mutation candidate locus.
- the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene sequencing read has the base type at the gene mutation candidate site and the gene mutation candidate site
- the variant base types of the points are inconsistent, that is, the third gene sequence read is the remaining gene sequence read from the gene sequence read except the first gene sequence read and the second gene sequence read.
- the third gene sequencing read may be a gene sequencing read in which there are inserted genes, deleted genes, etc., at candidate sites of gene mutation.
- the remaining third gene sequencing reads can be determined in the gene sequencing reads, and for each site in the preset site interval, the number of third gene sequencing reads at the site can be counted. From the number of third gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature is generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding site The number of sequencing reads of the third gene.
- the third gene sequencing read in the gene sequencing read can be determined, and then in the preset site interval For each site, determine the number of third gene sequencing reads whose base type of the third gene sequencing read is inconsistent with the base type of the reference genome, as the variation number of the third gene sequencing read, according to the third The number of mutations in gene sequencing reads determines the non-base arrangement characteristics of gene mutation candidate sites.
- the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene sequencing read has the base type at the gene mutation candidate site and the gene mutation candidate site
- the variant base types of the points are inconsistent, that is, the third gene sequence read is the remaining gene sequence read from the gene sequence read except the first gene sequence read and the second gene sequence read.
- the remaining third-gene sequencing reads can be determined in the gene-sequencing reads, and for each site in the preset site interval, count the genetic mutations of the third-gene sequencing reads at that site. The amount of variation.
- the number of mutations in the third gene sequencing read corresponding to each site can generate a non-base arrangement feature vector corresponding to the non-base arrangement feature vector, and each feature element in the non-base arrangement feature vector can be The variation number of the third gene sequencing read corresponding to the corresponding locus.
- a third gene sequencing read excluding the first gene sequencing read and the second gene sequencing read can be selected from the gene sequencing reads of normal cells. Then for each site in the preset site interval, count the number of third gene sequencing reads corresponding to each site and the number of mutations at that site.
- the gene sequencing reads derived from diseased cells in at least one gene sequencing read can be determined, and then based on the gene sequencing reads of the diseased cells Segment the non-base arrangement information of each site in the preset site interval to determine the non-base arrangement characteristics of gene mutation candidate sites. In this way, the non-base arrangement characteristics of gene mutation candidate sites can be determined based on the gene sequencing reads derived from diseased cells.
- the process of determining the non-base arrangement characteristics of the gene mutation candidate site based on the gene sequencing reads of the diseased cells can refer to the process of determining the non-base arrangement characteristics of the gene sequencing reads of the normal cells.
- the first gene sequencing reads, second gene sequencing reads, and third gene sequencing reads can be determined in the gene sequencing reads of diseased cells, and then targeted Set each site in the site interval, and count the number of first gene sequencing reads and the number of mutations corresponding to each site, the number of second gene sequencing reads and the number of mutations, and the number of third gene sequencing reads And the number of mutations, this information can correspond to the 11th to 16th dimensional features in the above feature matrix.
- the non-base arrangement information related to the base arrangement of at least one gene sequencing read in the preset site interval can be determined to determine the non-base arrangement characteristics of the gene mutation candidate site, so that the gene mutation can be identified Considering the invariance of the base arrangement of gene data, making gene mutation identification easier and more accurate.
- the following uses an example to illustrate the process of identifying gene mutations at gene mutation candidate sites.
- Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 5, the above step 14 may include the following steps:
- Step 141 Obtain a feature matrix of the gene mutation candidate site according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to the Base arrangement characteristics and non-base arrangement characteristics of gene mutation candidate sites, and the second dimension characteristics of the characteristic matrix correspond to the sites in the preset site interval;
- Step 142 Identify the gene mutation of the gene mutation candidate site according to the feature matrix of the gene mutation candidate site.
- the gene variation recognition model obtained based on neural network can be used to analyze the base arrangement feature and non-base arrangement feature.
- the arrangement feature performs feature integration, and the base arrangement feature vector formed by the base arrangement feature and the non-base arrangement feature vector formed by the non-base arrangement feature are combined into a feature matrix.
- the first dimensional feature of the feature matrix corresponds to base arrangement information and non-base arrangement information
- the second dimensional feature corresponds to a site in the preset site interval.
- the size of the feature matrix is the number of feature vectors ⁇ the size of the preset site interval.
- the size of the feature matrix can be 16 ⁇ 150, where the first-dimensional feature corresponds to the 16-dimensional feature vector, and the first To 4 can correspond to the base arrangement feature, and the 5th to 16th dimension feature vectors can correspond to the non-base arrangement feature, and have the invariance of base arrangement.
- the gene mutation identification model can be used to identify the gene mutation of the mutation candidate site according to the feature matrix.
- the neural network model can be used to integrate base arrangement information and non-base arrangement information corresponding to gene mutation candidate sites, so that gene sequencing data can be analyzed more comprehensively, and gene mutation identification can be more accurate.
- identifying the gene mutation at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site may include: according to the feature matrix of the gene mutation candidate site, Obtain the mutation value of the gene at the gene mutation candidate site, and if the mutation value is greater than or equal to a preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
- the mutation value of the gene mutation may be used to characterize the possibility of true mutation at the candidate site of the gene mutation. For example, if the mutation value is greater, the possibility of true mutation at the candidate site of the gene mutation is greater.
- the above-mentioned gene mutation recognition model can be used to process the obtained two-dimensional feature matrix to obtain the mutation value, and to determine whether the gene mutation at the gene mutation candidate site is a true mutation according to the mutation value.
- the variation value can be between 0 and 1.
- the preset threshold can be set according to the application scenario, for example, 0.3, 0.5. If the mutation value is greater than the preset threshold, the gene mutation at the candidate site of the gene mutation can be regarded as a true mutation, that is, the gene mutation caused by the disease; otherwise, The gene mutation at the candidate site of the gene mutation is a false mutation, that is, a gene abnormality caused by interference.
- the embodiments of the present disclosure can use a gene mutation recognition model to identify gene mutations at candidate sites of gene mutations.
- the base arrangement invariance of the gene data can be used to extract the gene mutation recognition model.
- the feature matrix is subjected to matrix transformation, so that data enhancement can be performed during the model training process, so that the trained gene mutation recognition model has better robustness and reduces over-fitting problems.
- Fig. 6 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
- the data enhancement of base arrangement information can be applied in the training process of the gene mutation recognition model.
- the feature matrix of gene mutation candidate sites is obtained, which may include:
- Step 1411 Generate a feature vector of each first-dimensional feature in the preset site interval according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate sites;
- Step 1412 Determine the base arrangement feature vector formed by the base arrangement feature in the feature vector
- Step 1413 Randomly sort the base arrangement feature vector to obtain the feature matrix of the gene mutation candidate site.
- the first dimension feature corresponds to the base arrangement information of the at least one gene sequencing read in the preset site interval
- the feature vector of the first dimension feature may include a base arrangement feature vector formed by the base arrangement feature and Non-base arrangement feature vector formed by non-base arrangement feature. Since the non-base arrangement feature has base arrangement invariance, the non-base arrangement feature will not be affected after the arrangement order of the base arrangement feature vector is changed. Therefore, the base arrangement feature vector formed by the base arrangement feature in the feature vector can be randomly sorted to obtain the feature matrix of the gene mutation candidate site, realize the data enhancement processing of the base arrangement information, and make the gene mutation recognition obtained after training The model considers the nature of the invariance of the base arrangement and has better performance.
- the first dimension feature corresponds to a 16-dimensional feature vector
- the first to fourth dimension can correspond to the base arrangement feature
- the fifth to 16th dimension feature vector can correspond to a non-base
- the embodiments of the present disclosure extract the base arrangement characteristics and non-base arrangement characteristics of gene mutation candidate sites, so that the invariance of the base arrangement of the gene data can be considered when the gene mutation is identified, so that the recognition result of the gene mutation is more improved. Accurate, screen out germline gene mutations and interference caused by noise and errors, and improve the accuracy of gene mutation recognition.
- the writing order of the steps does not mean a strict execution order but constitutes any limitation on the implementation process.
- the specific execution order of each step should be based on its function and possibility.
- the inner logic is determined.
- Fig. 7 shows a block diagram of a gene mutation recognition device according to an embodiment of the present disclosure. As shown in Fig. 7, the gene mutation recognition device includes:
- the first obtaining module 71 is configured to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
- the second obtaining module 72 is configured to obtain the base arrangement characteristics of the candidate gene mutation sites
- the determining module 73 is configured to determine the non-base arrangement characteristics of the gene mutation candidate site based on the non-base arrangement information of the at least one gene sequencing read in the preset site interval; wherein, the non-base arrangement The arrangement characteristics remain unchanged after the base arrangement sequence is changed;
- the identification module 74 is configured to identify the gene mutation of the gene mutation candidate site based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site.
- the second obtaining module 72 includes:
- the first determining sub-module is used to determine the preset site interval where the candidate site of the gene mutation is located;
- the second determining sub-module is used to obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used for characterization The sequence of bases.
- the determining module 73 includes:
- the first acquisition sub-module is used to acquire the non-base arrangement information of each site of the at least one gene sequencing read in the preset site interval;
- the third determining sub-module is used to determine the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
- the third determining submodule is specifically configured to:
- the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
- the third determining submodule is specifically configured to:
- the third determining submodule is specifically configured to:
- the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
- the third determining submodule is specifically configured to:
- the third determining submodule is specifically configured to:
- the identification module 74 includes:
- the generation sub-module is used to obtain the feature matrix of the gene mutation candidate site according to the base arrangement feature and non-base arrangement feature of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to Base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate sites, and the second-dimensional characteristics of the characteristic matrix correspond to sites in the preset site interval;
- the identification sub-module is used to identify the gene mutation of the gene mutation candidate site according to the feature matrix of the gene mutation candidate site.
- the identification submodule is specifically used for:
- the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
- the generating submodule is specifically used for:
- the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
- the first obtaining module includes:
- the second acquisition submodule is used to obtain the gene sequencing reads obtained by performing gene sequencing of somatic genes;
- the comparison submodule is used to compare the base sequence of the gene sequencing reads with the base sequence of the reference genome , Obtain the comparison result;
- the fourth determination sub-module used to determine the abnormal gene mutation candidate site of the gene of the somatic gene according to the comparison result;
- the third acquisition sub-module used to obtain the gene mutation At least one gene sequencing read corresponding to the candidate site.
- the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
- the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
- Fig. 8 is a block diagram showing a device 1900 for gene mutation identification according to an exemplary embodiment.
- the device 1900 may be provided as a server.
- the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions that can be executed by the processing component 1922, such as application programs.
- the application program stored in the memory 1932 may include one or more modules each corresponding to a set of instructions.
- the processing component 1922 is configured to execute instructions to perform the above-described methods.
- the device 1900 may also include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to the network, and an input output (I/O) interface 1958.
- the device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
- a non-volatile computer-readable storage medium such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to complete the foregoing method.
- the present disclosure may be a system, method, and/or computer program product.
- the computer program product may include a computer-readable storage medium loaded with computer-readable program instructions for enabling a processor to implement various aspects of the present disclosure.
- the computer-readable storage medium may be a tangible device that can hold and store instructions used by the instruction execution device.
- the computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
- Computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) Or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a printer with instructions stored thereon
- RAM random access memory
- ROM read-only memory
- EPROM erasable programmable read-only memory
- flash memory flash memory
- SRAM static random access memory
- CD-ROM compact disk read-only memory
- DVD digital versatile disk
- memory stick floppy disk
- mechanical encoding device such as a printer with instructions stored thereon
- the computer-readable storage medium used here is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (for example, light pulses through fiber optic cables), or through wires Transmission of electrical signals.
- the computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing/processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and/or a wireless network.
- the network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and/or edge servers.
- the network adapter card or network interface in each computing/processing device receives computer-readable program instructions from the network, and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing/processing device .
- the computer program instructions used to perform the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or in one or more programming languages.
- Source code or object code written in any combination, the programming language includes object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" language or similar programming languages.
- Computer-readable program instructions can be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly executed on a remote computer, or entirely on the remote computer or server carried out.
- the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to access the Internet connection).
- LAN local area network
- WAN wide area network
- an electronic circuit such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can be customized by using the status information of the computer-readable program instructions.
- the computer-readable program instructions are executed to realize various aspects of the present disclosure.
- These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device , A device that implements the functions/actions specified in one or more blocks in the flowchart and/or block diagram is produced. It is also possible to store these computer-readable program instructions in a computer-readable storage medium. These instructions make computers, programmable data processing apparatuses, and/or other devices work in a specific manner, so that the computer-readable medium storing instructions includes An article of manufacture, which includes instructions for implementing various aspects of the functions/actions specified in one or more blocks in the flowchart and/or block diagram.
- each block in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and the module, program segment, or part of an instruction contains one or more functions for implementing the specified logical function.
- Executable instructions may also occur in a different order from the order marked in the drawings. For example, two consecutive blocks can actually be executed in parallel, or they can sometimes be executed in the reverse order, depending on the functions involved.
- each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or actions Or it can be realized by a combination of dedicated hardware and computer instructions.
Landscapes
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Biotechnology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Theoretical Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Evolutionary Biology (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Epidemiology (AREA)
- Public Health (AREA)
- Primary Health Care (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Description
Claims (32)
- 一种基因变异识别方法,其特征在于,所述方法包括:获取基因变异候选位点对应的至少一个基因测序读段;获取所述基因变异候选位点的碱基排列特征;基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征碱基排列顺序改变后保持不变;基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求1所述的方法,其特征在于,所述获取所述基因变异候选位点的碱基排列特征,包括:确定所述基因变异候选位点所在的预设位点区间;根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
- 根据权利要求1或2所述的方法,其特征在于,所述基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致 的第二基因测序读段;在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3至9中任意一项所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求3至9中任意一项所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求1至11中任意一项所述的方法,其特征在于,所述基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别,包括:根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求12所述的方法,其特征在于,所述根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别,包括:根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
- 根据权利要求12所述的方法,其特征在于,所述根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵,包括:根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;确定所述特征向量中碱基排列特征形成的碱基排列特征向量;对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
- 根据权利要求1至14中任意一项所述的方法,其特征在于,获取基因变异候选位点对应的至少一个基因测序读段,包括:获取由体细胞基因进行基因测序得到的基因测序读段;将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;获取所述基因变异候选位点对应的至少一个基因测序读段。
- 一种基因变异识别装置,其特征在于,所述装置包括:第一获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;第二获取模块,用于获取所述基因变异候选位点的碱基排列特征;确定模块,用于基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变;识别模块,用于基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求16所述的装置,其特征在于,所述第二获取模块,包括:第一确定子模块,用于确定所述基因变异候选位点所在的预设位点区间;第二确定子模块,用于根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
- 根据权利要求16或17所述的装置,其特征在于,所述确定模块,包括:第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;第三确定子模块,用于基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18至24中任意一项所述的装置,其特征在于,所述第三确定子模块,具体用于,确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求18至24中任意一项所述的装置,其特征在于,所述第三确定子模块,具体用于,确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
- 根据权利要求16至26中任意一项所述的装置,其特征在于,所述识别模块,包括:生成子模块,用于根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;识别子模块,用于根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求27所述的装置,其特征在于,所述识别子模块,具体用于,根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
- 根据权利要求27所述的装置,其特征在于,所述生成子模块,具体用于,根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;确定所述特征向量中碱基排列特征形成的碱基排列特征向量;对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
- 根据权利要求16至29中任意一项所述的装置,其特征在于,所述第一获取模块,包括:第二获取子模块,用于获取由体细胞基因进行基因测序得到的基因测序读段;对比子模块,用于将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;第四确定子模块,用于根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;第三获取子模块,用于获取所述基因变异候选位点对应的至少一个基因测序读段。
- 一种基因变异识别装置,其特征在于,包括:处理器;用于存储处理器可执行指令的存储器;其中,所述处理器被配置为:执行权利要求1至15中任意一项所述的方法。
- 一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其特征在于,所述计算机程序指令被处理器执行时实现权利要求1至15中任意一项所述的方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| SG11202101410WA SG11202101410WA (en) | 2019-03-29 | 2019-05-31 | Genovariation identification method and device, and storage medium |
| JP2021517044A JP7064655B2 (ja) | 2019-03-29 | 2019-05-31 | 遺伝子変異認識方法、装置および記憶媒体 |
| US17/162,465 US20210151124A1 (en) | 2019-03-29 | 2021-01-29 | Genetic variation identification method, genetic variation identification apparatuses, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910252747.9A CN109979531B (zh) | 2019-03-29 | 2019-03-29 | 一种基因变异识别方法、装置和存储介质 |
| CN201910252747.9 | 2019-03-29 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/162,465 Continuation US20210151124A1 (en) | 2019-03-29 | 2021-01-29 | Genetic variation identification method, genetic variation identification apparatuses, and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020199337A1 true WO2020199337A1 (zh) | 2020-10-08 |
Family
ID=67081906
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/089504 Ceased WO2020199337A1 (zh) | 2019-03-29 | 2019-05-31 | 一种基因变异识别方法、装置和存储介质 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US20210151124A1 (zh) |
| JP (1) | JP7064655B2 (zh) |
| CN (1) | CN109979531B (zh) |
| SG (1) | SG11202101410WA (zh) |
| TW (1) | TWI740262B (zh) |
| WO (1) | WO2020199337A1 (zh) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111081313A (zh) * | 2019-12-13 | 2020-04-28 | 北京市商汤科技开发有限公司 | 基因变异的识别方法及装置、电子设备和存储介质 |
| CN111091873B (zh) * | 2019-12-13 | 2023-07-18 | 北京市商汤科技开发有限公司 | 基因变异的识别方法及装置、电子设备和存储介质 |
| CN111899790A (zh) * | 2020-08-17 | 2020-11-06 | 天津诺禾医学检验所有限公司 | 测序数据的处理方法及装置 |
| CN113539357B (zh) * | 2021-06-10 | 2024-04-30 | 阿里巴巴达摩院(杭州)科技有限公司 | 基因检测方法、模型训练方法、装置、设备及系统 |
| CN115458052B (zh) * | 2022-08-16 | 2023-06-30 | 珠海横琴铂华医学检验有限公司 | 基于一代测序的基因突变分析方法、设备和存储介质 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2015062184A1 (en) * | 2013-11-01 | 2015-05-07 | Accurascience, Llc | Method and apparatus for calling single-nucleotide variations and other variations |
| CN106611106A (zh) * | 2016-12-06 | 2017-05-03 | 北京荣之联科技股份有限公司 | 基因变异检测方法及装置 |
| CN108595912A (zh) * | 2018-05-07 | 2018-09-28 | 深圳市瀚海基因生物科技有限公司 | 检测染色体非整倍性的方法、装置及系统 |
Family Cites Families (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2002046462A2 (en) * | 2000-12-07 | 2002-06-13 | Isis Innovation Limited | Functional genetic variants of matrix metalloproteinases (nmps) |
| WO2012149042A2 (en) * | 2011-04-25 | 2012-11-01 | Bio-Rad Laboratories, Inc. | Methods and compositions for nucleic acid analysis |
| AU2013236175A1 (en) * | 2012-03-22 | 2014-10-16 | Kanagawa Prefectural Hospital Organization | Method for identification and detection of mutant gene using intercalator |
| CN104603609B (zh) * | 2013-07-31 | 2016-08-24 | 株式会社日立制作所 | 基因分析装置、基因分析系统以及基因分析方法 |
| WO2015200869A1 (en) * | 2014-06-26 | 2015-12-30 | 10X Genomics, Inc. | Analysis of nucleic acid sequences |
| SG11201703638XA (en) * | 2014-11-10 | 2017-06-29 | Alnylam Pharmaceuticals Inc | Hepatitis b virus (hbv) irna compositions and methods of use thereof |
| JP6675164B2 (ja) | 2015-07-28 | 2020-04-01 | 株式会社理研ジェネシス | 変異判定方法、変異判定プログラムおよび記録媒体 |
| JP6679065B2 (ja) | 2015-10-07 | 2020-04-15 | 国立研究開発法人国立がん研究センター | 稀少突然変異の検出方法、検出装置及びコンピュータプログラム |
| US9988624B2 (en) * | 2015-12-07 | 2018-06-05 | Zymergen Inc. | Microbial strain improvement by a HTP genomic engineering platform |
| US20200199612A1 (en) * | 2016-03-18 | 2020-06-25 | Monsanto Technology Llc | Transgenic plants with enhanced traits |
| CN106529211A (zh) * | 2016-11-04 | 2017-03-22 | 成都鑫云解码科技有限公司 | 变异位点的获取方法及装置 |
| CN106503489A (zh) * | 2016-11-04 | 2017-03-15 | 成都鑫云解码科技有限公司 | 心血管系统对应的基因的突变位点的获取方法及装置 |
| CN106407747A (zh) * | 2016-11-04 | 2017-02-15 | 成都鑫云解码科技有限公司 | 肿瘤对应的基因的突变位点的获取方法及装置 |
| KR101936933B1 (ko) * | 2016-11-29 | 2019-01-09 | 연세대학교 산학협력단 | 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스 |
| JP7350659B2 (ja) * | 2017-06-06 | 2023-09-26 | ザイマージェン インコーポレイテッド | Saccharopolyspora spinosaの改良のためのハイスループット(HTP)ゲノム操作プラットフォーム |
| CN109033751B (zh) * | 2018-07-20 | 2021-07-27 | 东南大学 | 一种非编码区单核苷酸基因组变异的功能预测方法 |
| CN109411016B (zh) * | 2018-11-14 | 2020-12-01 | 钟祥博谦信息科技有限公司 | 基因变异位点检测方法、装置、设备及存储介质 |
-
2019
- 2019-03-29 CN CN201910252747.9A patent/CN109979531B/zh active Active
- 2019-05-31 WO PCT/CN2019/089504 patent/WO2020199337A1/zh not_active Ceased
- 2019-05-31 SG SG11202101410WA patent/SG11202101410WA/en unknown
- 2019-05-31 JP JP2021517044A patent/JP7064655B2/ja not_active Expired - Fee Related
- 2019-11-04 TW TW108139976A patent/TWI740262B/zh active
-
2021
- 2021-01-29 US US17/162,465 patent/US20210151124A1/en not_active Abandoned
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2015062184A1 (en) * | 2013-11-01 | 2015-05-07 | Accurascience, Llc | Method and apparatus for calling single-nucleotide variations and other variations |
| CN106611106A (zh) * | 2016-12-06 | 2017-05-03 | 北京荣之联科技股份有限公司 | 基因变异检测方法及装置 |
| CN108595912A (zh) * | 2018-05-07 | 2018-09-28 | 深圳市瀚海基因生物科技有限公司 | 检测染色体非整倍性的方法、装置及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20210151124A1 (en) | 2021-05-20 |
| TWI740262B (zh) | 2021-09-21 |
| CN109979531B (zh) | 2021-08-31 |
| SG11202101410WA (en) | 2021-03-30 |
| CN109979531A (zh) | 2019-07-05 |
| JP2022502766A (ja) | 2022-01-11 |
| TW202036584A (zh) | 2020-10-01 |
| JP7064655B2 (ja) | 2022-05-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109994155B (zh) | 一种基因变异识别方法、装置和存储介质 | |
| TWI740262B (zh) | 一種基因變異識別方法、裝置和儲存介質 | |
| Kaminski et al. | pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models | |
| Graham et al. | BinSanity: unsupervised clustering of environmental microbial assemblies using coverage and affinity propagation | |
| Schrider et al. | Soft sweeps are the dominant mode of adaptation in the human genome | |
| Peltzer et al. | EAGER: efficient ancient genome reconstruction | |
| Lee et al. | DUDE-Seq: fast, flexible, and robust denoising for targeted amplicon sequencing | |
| Nguyen et al. | TIPP: taxonomic identification and phylogenetic profiling | |
| Yaveroğlu et al. | Proper evaluation of alignment-free network comparison methods | |
| CN111292802A (zh) | 用于检测突变的方法、电子设备和计算机存储介质 | |
| Sarmashghi et al. | Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT | |
| WO2023143016A1 (zh) | 特征提取模型的生成方法、图像特征提取方法和装置 | |
| Qian et al. | MetaCon: unsupervised clustering of metagenomic contigs with probabilistic k-mers statistics and coverage | |
| CN109979530B (zh) | 一种基因变异识别方法、装置和存储介质 | |
| Alganmi et al. | Evaluation of an optimized germline exomes pipeline using BWA-MEM2 and Dragen-GATK tools | |
| CN111933214B (zh) | 用于检测rna水平体细胞基因变异的方法、计算设备 | |
| CN118655989A (zh) | 提示词生成方法及文本处理方法 | |
| Qian et al. | TEtrimmer: a tool to automate the manual curation of transposable elements | |
| CN106326904A (zh) | 获取特征排序模型的装置和方法以及特征排序方法 | |
| Huang et al. | Reveel: large-scale population genotyping using low-coverage sequencing data | |
| Qi et al. | CREATE: a novel attention-based framework for efficient classification of transposable elements | |
| Bonham-Carter et al. | Cellular proliferation biases clonal lineage tracing and trajectory inference | |
| Charitakis et al. | Comparative analysis of packages and algorithms for the analysis of spatially resolved transcriptomics data | |
| Nyström-Persson et al. | Precise and scalable metagenomic profiling with sample-tailored minimizer libraries | |
| CN110570908B (zh) | 测序序列多态识别方法及装置、存储介质、电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19923039 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021517044 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 04/02/2022) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19923039 Country of ref document: EP Kind code of ref document: A1 |