WO2020199337A1 - 一种基因变异识别方法、装置和存储介质 - Google Patents

一种基因变异识别方法、装置和存储介质 Download PDF

Info

Publication number
WO2020199337A1
WO2020199337A1 PCT/CN2019/089504 CN2019089504W WO2020199337A1 WO 2020199337 A1 WO2020199337 A1 WO 2020199337A1 CN 2019089504 W CN2019089504 W CN 2019089504W WO 2020199337 A1 WO2020199337 A1 WO 2020199337A1
Authority
WO
WIPO (PCT)
Prior art keywords
gene
site
base arrangement
gene mutation
base
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/089504
Other languages
English (en)
French (fr)
Inventor
胡志强
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to SG11202101410WA priority Critical patent/SG11202101410WA/en
Priority to JP2021517044A priority patent/JP7064655B2/ja
Publication of WO2020199337A1 publication Critical patent/WO2020199337A1/zh
Priority to US17/162,465 priority patent/US20210151124A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H20/00ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search

Definitions

  • the present disclosure relates to the field of computer technology, and in particular to a method, device and storage medium for identifying gene mutations.
  • gene sequencing technology greatly improves the efficiency of gene sequencing, reduces the cost of gene sequencing, and maintains the accuracy of gene sequencing. If the first-generation testing technology completes the sequencing of a human genome, it may take three years, while the second-generation sequencing technology can shorten the time to just one week.
  • the present disclosure proposes a technical solution for gene mutation identification.
  • a method for identifying gene mutations comprising:
  • the obtaining the base arrangement characteristics of the gene mutation candidate site includes:
  • the base arrangement feature is used to characterize the base arrangement sequence.
  • the determining the non-base arrangement characteristics of the gene mutation candidate site based on the non-base arrangement information of the at least one gene sequencing read in a preset site interval includes:
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the determining the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval includes:
  • the recognizing the gene variation of the gene variation candidate site based on the base arrangement feature and the non-base arrangement feature of the gene variation candidate site includes:
  • a feature matrix of the gene mutation candidate site is obtained; wherein the first dimension feature of the feature matrix corresponds to the gene mutation candidate.
  • the base arrangement feature and non-base arrangement feature of the site, the second dimension feature of the feature matrix corresponds to the site in the preset site interval; according to the feature matrix of the gene mutation candidate site, all Identify the genetic variation at the candidate site of the genetic variation.
  • the identifying the gene mutation at the gene mutation candidate site according to the feature matrix of the gene mutation candidate site includes:
  • the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
  • the obtaining the feature matrix of the gene mutation candidate site according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site includes:
  • the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
  • obtaining at least one gene sequencing read corresponding to the gene mutation candidate site includes:
  • a gene mutation identification device comprising:
  • the first acquisition module is used to acquire at least one gene sequencing read corresponding to the gene mutation candidate site; the second acquisition module is used to acquire the base arrangement characteristics of the gene mutation candidate site; the determination module is used to The non-base arrangement information of the at least one gene sequencing read in the preset site interval determines the non-base arrangement feature of the gene mutation candidate site; wherein the non-base arrangement feature changes in the base arrangement sequence After that, it remains unchanged; the recognition module is used to identify the gene mutation of the gene mutation candidate site based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site.
  • the second acquisition module includes:
  • the first determining sub-module is used to determine the preset site interval where the candidate site of the gene mutation is located;
  • the second determining sub-module is used to obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used for characterization The sequence of bases.
  • the determining module includes:
  • the first obtaining submodule is used to obtain the non-base arrangement information of each site in the preset site interval of the at least one gene sequencing read;
  • the non-base arrangement information of each site in the site interval determines the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the third determining submodule is specifically configured to:
  • the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the third determining submodule is specifically configured to:
  • the identification module includes:
  • the generation sub-module is used to obtain the feature matrix of the gene mutation candidate site according to the base arrangement feature and non-base arrangement feature of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to The base arrangement feature and non-base arrangement feature of the gene mutation candidate site, the second dimension feature of the feature matrix corresponds to the site in the preset site interval; the identification sub-module is used for The feature matrix of the gene mutation candidate site is used to identify the gene mutation of the gene mutation candidate site.
  • the identification submodule is specifically used for:
  • the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
  • the generating submodule is specifically used for:
  • the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
  • the first obtaining module includes:
  • the second acquisition submodule is used to obtain the gene sequencing reads obtained by performing gene sequencing of somatic genes;
  • the comparison submodule is used to compare the base sequence of the gene sequencing reads with the base sequence of the reference genome , Obtain the comparison result;
  • the fourth determination sub-module used to determine the abnormal gene mutation candidate site of the gene of the somatic gene according to the comparison result;
  • the third acquisition sub-module used to obtain the gene mutation At least one gene sequencing read corresponding to the candidate site.
  • a gene mutation identification device including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above method.
  • a non-volatile computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor.
  • the gene mutation identification solution can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, and obtain the base arrangement characteristics of the gene mutation candidate site, based on the fact that at least one gene sequencing read is in a preset position
  • the base arrangement information of the point interval determines the non-base arrangement characteristics of the gene mutation candidate site, so that the gene mutation of the gene mutation candidate site can be based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site Identify it.
  • the non-base arrangement feature remains unchanged after the base arrangement sequence is changed, that is, it can be considered that the non-base arrangement feature has the nature of base arrangement invariance.
  • the gene mutation of the candidate site of gene mutation is not restricted by the sequence of bases, and the pseudogene mutation caused by the germline gene mutation, noise, error, etc. can be better screened out, so that the gene mutation can be better To identify mutations to improve the accuracy of gene mutation identification,
  • Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
  • Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
  • Fig. 3 shows a flowchart of the base arrangement characteristic process of gene mutation candidate sites according to an embodiment of the present disclosure.
  • FIG. 4 shows a flowchart of the non-base arrangement feature process of gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 6 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 7 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
  • FIG. 8 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
  • the gene mutation identification scheme provided by the embodiments of the present disclosure can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, so that at least one gene sequencing read can be used to identify the gene mutation of the gene mutation candidate site.
  • the base arrangement characteristics of the gene mutation candidate site can be determined, and the non-base of the gene mutation candidate site can be determined according to the base arrangement information of at least one gene sequencing read in the preset site interval
  • the arrangement feature can then be used to identify the genetic variation at the candidate site of the gene variation through the base arrangement feature and the non-base arrangement feature.
  • the non-base arrangement feature here remains unchanged after the base arrangement sequence is changed, that is, it can be considered that whether the genetic mutation at the genetic mutation candidate site is a true mutation is not affected by the base arrangement sequence, so that the When identifying genetic mutations, consider the invariance of the base arrangement of genetic data to improve the accuracy of genetic mutation identification.
  • Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
  • the gene mutation identification method can be executed by a gene mutation identification device or other processing equipment, where the gene mutation identification device can be User Equipment (UE), mobile equipment, user terminal, terminal, cellular phone, cordless phone, personal digital Processing (Personal Digital Assistant, PDA), handheld devices, computing devices, vehicle-mounted devices, wearable devices, etc., or the gene mutation recognition device may be a server.
  • the gene mutation identification method may be implemented by a processor calling computer-readable instructions stored in a memory. As shown in Figure 1, the gene mutation identification method includes:
  • Step 11 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
  • the gene mutation recognition device can obtain the gene sequencing reads obtained by gene sequencing, and then obtain at least one gene sequencing read corresponding to the gene mutation candidate site from the gene sequencing reads obtained by the gene sequencing.
  • the gene sequencing reads here can be understood as base sequences marked with base types after gene sequencing, and the length of each gene sequencing read can be the same or different. In the case of different lengths, the length of each gene sequencing read segment can be within a preset length range, thereby ensuring that the length of each gene sequencing read segment is relatively close.
  • the base type may include cytosine (C), guanine (G), adenine (A), and thymine (T), so that the gene sequencing reads may include the base sequence of AGCT.
  • the gene mutation candidate site may be a site with an abnormal base sequence.
  • the site of the base sequence may indicate the position of the base sequence.
  • there may be at least one gene sequencing read that is, at the same site, there may be at least one gene sequencing read obtained by gene sequencing.
  • the gene mutation candidate site corresponds to at least one gene sequencing read segment, wherein the at least one gene sequencing read segment covers this locus.
  • there may be at least one gene mutation candidate site and each gene mutation candidate site may correspond to at least one gene sequencing read.
  • the embodiment of the present disclosure uses a gene mutation candidate site for description.
  • Step 12 Obtain the base arrangement characteristics of the gene mutation candidate sites.
  • the gene mutation recognition model can be used to extract the base arrangement characteristics of the gene mutation candidate site based on the gene arrangement information of the gene mutation candidate site.
  • the base arrangement information here can be information related to the base arrangement sequence. For example, if the base sequence of a certain gene sequencing read in a certain site interval is A, C, G, T, then the base arrangement information Can be ACGT.
  • the base arrangement information may include the base type of the reference genome in the preset site interval, the number of genes of each base type, the number of missing genes of each base type, the number of inserted genes of each base type, and so on.
  • the base arrangement characteristics obtained from the base arrangement information are related to the base arrangement order.
  • Step 13 based on the non-base arrangement information of the at least one gene sequencing read in the preset site interval, determine the non-base arrangement feature of the gene mutation candidate site; wherein, the non-base arrangement feature is The sequence of bases remains unchanged after changing.
  • the base of at least one gene sequencing read corresponding to the gene mutation candidate site may be extracted in a preset site interval Arrange the information, and generate the non-base arrangement characteristics of the gene mutation candidate site according to the extracted base arrangement information.
  • the non-base arrangement information may be information that is not restricted by the base arrangement order. Therefore, the non-base arrangement characteristics of the gene mutation candidate sites can be determined according to the non-base arrangement information of at least one gene sequencing read in the preset site interval.
  • the non-base arrangement information may include information with base arrangement invariance such as the number of gene sequencing reads corresponding to the site, the number of gene sequencing reads that are mutated at the site, and the like.
  • non-base arrangement information when extracting non-base arrangement information, several gene sequencing reads corresponding to the candidate site of the gene mutation can be randomly selected, and the non-base arrangement information of several randomly selected gene sequencing reads can be extracted; The non-base arrangement information of each gene sequencing read corresponding to the gene mutation candidate site.
  • the non-base arrangement information of at least one gene sequencing read in the preset site interval When extracting the non-base arrangement information of at least one gene sequencing read in the preset site interval, the non-base arrangement information of at least one gene sequencing read in each site within the preset site interval can be extracted, and Several adjacent sites in the preset site interval can be randomly selected, and the non-base arrangement information of at least one gene sequencing read at several adjacent sites can be extracted.
  • a gene mutation recognition model obtained based on neural network training can be used.
  • Step 14 based on the base arrangement feature and non-base arrangement feature of the gene mutation candidate site, identify the gene mutation of the gene mutation candidate site.
  • the characteristic matrix of the gene mutation candidate site can be obtained from the base arrangement characteristics and the non-base arrangement characteristics, and the characteristic matrix Recognition of gene mutations in gene mutation candidate sites.
  • the above gene mutation recognition model can be used to determine whether the gene at the gene mutation candidate site is a true mutation caused by a disease or a base sequence abnormality caused by noise or other reasons. Of false mutations.
  • the obtained feature matrix of the gene mutation candidate site can be a two-dimensional feature matrix, and the size of the feature matrix can be the number of feature vectors ⁇ the size of the preset site interval, and the feature vector can be based on the base arrangement feature And non-base alignment features.
  • the gene mutation at the candidate site of the mutation is caused by the disease is not affected by the sequence of bases, and is more affected by the genetic environment where the candidate site of the mutation is located, for example, by the vicinity of the candidate site of the gene mutation
  • the other sites are affected by genetic environment such as mutant genes, so the order of the feature vectors corresponding to the base arrangement feature in the resulting feature matrix can be unlimited, and the order of the feature vectors of the base arrangement feature in the feature matrix It can be changed randomly to improve the efficiency and accuracy of gene mutation identification.
  • the gene mutation of the gene mutation candidate site can be identified according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site, so that the invariance of the base arrangement of the gene mutation can be considered, and the Identify genetic variation.
  • the gene mutation candidate site at least one gene sequencing read corresponding to the gene mutation candidate site can be obtained.
  • the example of the present disclosure also provides a process for obtaining at least one gene sequencing read corresponding to the gene mutation candidate site.
  • Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
  • obtaining at least one gene sequencing read corresponding to the gene mutation candidate site may include the following steps:
  • Step 111 Obtain gene sequencing reads obtained by gene sequencing of somatic genes.
  • At least one gene sequencing read segment can be obtained by performing gene sequencing of the somatic cell gene, and the gene sequencing read segment can be a sequence that annotates the base type of the somatic cell gene. After gene sequencing of somatic genes, not only the base type of each gene in the gene sequencing read, but also the gene location information of each gene in the gene sequencing read can be obtained. The same site can correspond to at least one gene sequencing read.
  • At least one gene sequencing read can be obtained through gene sequencing of somatic genes, and the gene sequencing reads obtained by gene sequencing can be preprocessed.
  • the preprocessing methods here can include cross contamination screening, Sequencing quality screening, comparison quality screening, abnormal read length screening, etc. Through preprocessing, cross-contaminated gene sequencing reads can be screened out, and gene sequencing reads with low sequencing quality and comparison quality and abnormal read length can be screened out.
  • Step 112 Align the base sequence of the gene sequencing read with the base sequence of the reference genome to obtain an alignment result.
  • the base sequence of the obtained gene sequencing reads can be compared with the base sequence of the reference genome at the same site , Get the comparison result. For example, you can compare the base sequence of each gene sequencing read obtained by gene sequencing with the base sequence of the reference genome at the same site to determine the base sequence of the gene sequencing read that is different from the base sequence of the reference genome. point. It is also possible to compare the base sequence of at least one gene sequencing read with the same site with the base sequence of the reference genome at the same site to determine the site where the base sequence of at least one gene sequencing read is different from the base sequence of the reference genome .
  • the reference genome may be a base sequence labeled with a correct base sequence.
  • Step 113 According to the comparison result, it is determined that the gene of the somatic gene has an abnormal gene mutation candidate site.
  • the base sequence of the gene sequencing read and the reference genome can be determined according to the comparison result. If at least one gene sequencing read corresponding to the locus is in at least one gene sequencing read, the mutation is sent at that position. If the ratio of gene sequencing reads is greater than the preset ratio, it can be determined that the site is a candidate site for gene mutation; otherwise, it can be considered that the site is not a candidate site for gene mutation.
  • the base sequence of the gene sequencing read at this position is different from that of the reference genome, which may be caused by sequencing errors. In this way, the abnormality of base sequence caused by gene sequencing errors can be reduced.
  • Step 114 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
  • At least one gene sequencing read corresponding to the candidate gene mutation site can be obtained.
  • the base sequence of the gene mutation candidate site may be different from the base sequence of the reference genome at the same site.
  • the base arrangement characteristics of the gene mutation candidate site can be determined according to the base arrangement information of at least one gene sequencing read corresponding to the gene mutation candidate site, so as to identify the gene mutation at the gene mutation candidate site.
  • data enhancement processing can be performed on gene recognition based on the base arrangement characteristics. The following describes in detail the process of determining the base arrangement characteristics of gene mutation candidate sites through an example.
  • Fig. 3 shows a flowchart of the base arrangement characteristic process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 3, the above step 12 may include the following steps:
  • Step 121 Determine the preset site interval where the candidate site of gene mutation is located
  • Step 122 Obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used to characterize the base arrangement sequence.
  • the base arrangement information may include the base arrangement information of the candidate genome.
  • the base arrangement information is the base arrangement information of the candidate genome, it can be considered that the base arrangement information of each gene sequencing read is the same. Base arrangement information of the candidate genome. Therefore, according to the gene location information of the gene mutation candidate site, the preset site interval where the gene mutation candidate site is located can be determined.
  • the interval formed by 150 bases before and after the gene mutation candidate site can be used as the gene mutation candidate site.
  • the preset site interval where the point is located.
  • the base arrangement information of the reference genome in the preset site interval can be obtained, and gene mutation candidate positions can be generated from the base arrangement information of the reference genome in the preset site interval.
  • the base arrangement information can refer to the base sequence composition of each site in the preset site interval of the genome.
  • the preset site interval includes 4 base sequences, namely A, C, G, and T.
  • the base arrangement information may be the base arrangement order of ACGT.
  • the base arrangement feature can be represented by the base arrangement feature vector, which can be part of the feature matrix of the gene mutation candidate site.
  • a1, a2, a3, and a4 can be the first 4-dimensional features of the feature matrix.
  • the base arrangement characteristics corresponding to the gene mutation candidate sites are considered when identifying the gene mutation at the gene mutation candidate sites, but also the base arrangement characteristics of the gene mutation candidate sites are considered.
  • Denatured non-base alignment features The following describes in detail the process of determining the non-base arrangement characteristics of gene mutation candidate sites through an example.
  • FIG. 4 shows a flowchart of the non-base arrangement feature process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 4, the above step 13 may include the following steps:
  • Step 131 Obtain non-base arrangement information of each site in the preset site interval of the at least one gene sequencing read;
  • Step 132 based on the non-base arrangement information of each site in the preset site interval, determine the non-base arrangement feature of the gene mutation candidate site.
  • Non-base arrangement information may be information with the invariance of base arrangement, for example, the number of gene sequencing reads and the number of mutations corresponding to the site.
  • the non-base permutation feature generated by each type of non-base permutation information can form a non-base permutation feature vector, and there can be one or more non-base permutation feature vectors.
  • the gene mutation identification scheme provided by the embodiments of the present disclosure can be applied to patients who have been diagnosed with cancer, and the gene mutation identification can guide the patient to take medication. Therefore, a part of the gene sequencing reads in the gene sequencing reads can be derived from normal cells, and normal cells can be considered as cells that are not diseased. Some gene sequencing reads can be derived from diseased cells. Therefore, when determining the non-base arrangement characteristics of gene mutation candidate sites, the non-base arrangement of gene mutation candidate sites can be determined based on gene sequencing reads derived from normal cells and gene sequencing reads derived from diseased cells. feature.
  • the non-base arrangement characteristics of gene mutation candidate sites when determining the non-base arrangement characteristics of gene mutation candidate sites, it is possible to determine the gene sequencing reads derived from normal cells in at least one gene sequencing read, and then based on the gene sequencing reads of normal cells Segment the non-base arrangement information of each site in the preset site interval to determine the non-base arrangement characteristics of gene mutation candidate sites. In this way, the non-base arrangement characteristics of gene mutation candidate sites can be determined based on gene sequencing reads derived from normal cells.
  • the following provides several examples of determining the non-base arrangement characteristics of gene mutation candidate sites based on the gene sequencing reads of normal cells.
  • the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site is consistent with the base type of the reference genome. A gene sequencing read, and then according to the number of the first gene sequencing reads corresponding to each site in the preset site interval, the non-base arrangement characteristics of the gene mutation candidate sites are determined.
  • the first gene sequencing read that has not undergone genetic mutation at the gene mutation candidate site can be selected, and for each site in the preset site interval, the first gene sequencing can be counted The number of reads at that location. In other words, you can count how many first gene sequencing reads contain the locus.
  • the first gene sequencing read that includes a certain site can be considered as the first gene sequencing read corresponding to the site. Since the length of each gene sequencing read may be different, the position of the gene mutation candidate site relative to each gene sequencing read is different.
  • the gene mutation candidate site can be located in the middle of the gene sequencing read, or it can be located in the gene The edge positions of the sequencing reads, so that the number of gene sequencing reads corresponding to each site in the preset site interval is different. From the number of first gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature can be generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding position. The number of first gene sequencing reads of the point.
  • the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site is consistent with the base type of the reference genome First gene sequencing reads, and then at each position in the preset site interval, determine the number of first gene sequencing reads whose base types are inconsistent with those of the reference genome. As the variation quantity of the first gene sequencing read segment, the non-base arrangement characteristics of the gene variation candidate site are determined according to the variation quantity of the first gene sequencing read segment.
  • the first gene sequencing read that has not undergone genetic mutation at the gene mutation candidate site can be selected, and for each site in the preset site interval, the first gene sequencing can be counted The number of genetic mutations in the read at this locus.
  • gene sequencing reads did not undergo genetic mutation at the gene mutation candidate site (that is, the gene mutation candidate site is consistent with the base type of the reference genome), it may occur at other sites other than the gene mutation candidate site Gene mutation (that is, the base type is inconsistent with the reference genome at other sites), so that for each site in the preset site interval, the number of mutations in the first gene sequencing read of that site can be counted .
  • the first gene sequencing reads in the gene sequencing reads of normal cells that have not been mutated at the gene mutation candidate site, and then targeting the preset site interval For each site, count the number of first gene sequencing reads corresponding to each site and the number of mutations at that site.
  • These two pieces of information can correspond to the fifth-dimensional feature and the sixth-dimensional feature in the above feature matrix. Dimensional characteristics.
  • the non-base arrangement characteristics of the gene mutation candidate site when determining the non-base arrangement characteristics of the gene mutation candidate site, it can be determined in the gene sequencing read that the gene mutation candidate site and the gene mutation candidate site The second gene sequencing reads with the same mutation base type of the points, and then according to the number of second gene sequencing reads corresponding to each site in the preset site interval, determine the non-uniformity of the gene mutation candidate site Base arrangement characteristics.
  • the second gene sequencing reads that are consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads.
  • the second gene sequencing reads can be counted The number at that site. From the number of second gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature is generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding site The number of second gene sequencing reads.
  • the gene mutation candidate site and the gene mutation candidate site can be determined in the gene sequencing read. Sequencing reads of the second gene with the same variant base type, and then at each position in the preset site interval, determine the second gene whose base type of the second gene sequencing read is inconsistent with that of the reference genome The number of sequencing reads is used as the number of mutations of the second gene sequencing reads. According to the number of mutations of the second gene sequencing reads, the non-base arrangement characteristics of gene mutation candidate sites are determined.
  • the second gene sequencing read that is consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads (the mutation base type of the gene mutation candidate site can be obtained through gene sequencing), and the preset position For each site in the point interval, count the number of mutations of the second gene sequencing read at that site, in other words, count the number of second gene sequencing reads that contain the site and have mutations at that site .
  • the number of mutations in the second gene sequencing read corresponding to each site can generate a non-base arrangement feature vector corresponding to the non-base arrangement feature vector.
  • Each feature element in the non-base arrangement feature vector can be The variation number of the second gene sequencing read corresponding to the corresponding locus.
  • a second gene sequencing read that is consistent with the mutation of the gene mutation candidate site can be selected from the gene sequencing reads of normal cells, and then targeted at the preset site interval For each site, count the number of second gene sequencing reads corresponding to each site and the number of mutations at that site.
  • the third gene sequencing read in the gene sequencing read can be determined, and then according to the preset site interval The number of third gene sequencing reads corresponding to each locus in, determines the non-base arrangement characteristics of gene mutation candidate locus.
  • the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene sequencing read has the base type at the gene mutation candidate site and the gene mutation candidate site
  • the variant base types of the points are inconsistent, that is, the third gene sequence read is the remaining gene sequence read from the gene sequence read except the first gene sequence read and the second gene sequence read.
  • the third gene sequencing read may be a gene sequencing read in which there are inserted genes, deleted genes, etc., at candidate sites of gene mutation.
  • the remaining third gene sequencing reads can be determined in the gene sequencing reads, and for each site in the preset site interval, the number of third gene sequencing reads at the site can be counted. From the number of third gene sequencing reads corresponding to each site, a non-base alignment feature vector corresponding to the non-base alignment feature is generated, and each feature element in the non-base alignment feature vector can correspond to the corresponding site The number of sequencing reads of the third gene.
  • the third gene sequencing read in the gene sequencing read can be determined, and then in the preset site interval For each site, determine the number of third gene sequencing reads whose base type of the third gene sequencing read is inconsistent with the base type of the reference genome, as the variation number of the third gene sequencing read, according to the third The number of mutations in gene sequencing reads determines the non-base arrangement characteristics of gene mutation candidate sites.
  • the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene sequencing read has the base type at the gene mutation candidate site and the gene mutation candidate site
  • the variant base types of the points are inconsistent, that is, the third gene sequence read is the remaining gene sequence read from the gene sequence read except the first gene sequence read and the second gene sequence read.
  • the remaining third-gene sequencing reads can be determined in the gene-sequencing reads, and for each site in the preset site interval, count the genetic mutations of the third-gene sequencing reads at that site. The amount of variation.
  • the number of mutations in the third gene sequencing read corresponding to each site can generate a non-base arrangement feature vector corresponding to the non-base arrangement feature vector, and each feature element in the non-base arrangement feature vector can be The variation number of the third gene sequencing read corresponding to the corresponding locus.
  • a third gene sequencing read excluding the first gene sequencing read and the second gene sequencing read can be selected from the gene sequencing reads of normal cells. Then for each site in the preset site interval, count the number of third gene sequencing reads corresponding to each site and the number of mutations at that site.
  • the gene sequencing reads derived from diseased cells in at least one gene sequencing read can be determined, and then based on the gene sequencing reads of the diseased cells Segment the non-base arrangement information of each site in the preset site interval to determine the non-base arrangement characteristics of gene mutation candidate sites. In this way, the non-base arrangement characteristics of gene mutation candidate sites can be determined based on the gene sequencing reads derived from diseased cells.
  • the process of determining the non-base arrangement characteristics of the gene mutation candidate site based on the gene sequencing reads of the diseased cells can refer to the process of determining the non-base arrangement characteristics of the gene sequencing reads of the normal cells.
  • the first gene sequencing reads, second gene sequencing reads, and third gene sequencing reads can be determined in the gene sequencing reads of diseased cells, and then targeted Set each site in the site interval, and count the number of first gene sequencing reads and the number of mutations corresponding to each site, the number of second gene sequencing reads and the number of mutations, and the number of third gene sequencing reads And the number of mutations, this information can correspond to the 11th to 16th dimensional features in the above feature matrix.
  • the non-base arrangement information related to the base arrangement of at least one gene sequencing read in the preset site interval can be determined to determine the non-base arrangement characteristics of the gene mutation candidate site, so that the gene mutation can be identified Considering the invariance of the base arrangement of gene data, making gene mutation identification easier and more accurate.
  • the following uses an example to illustrate the process of identifying gene mutations at gene mutation candidate sites.
  • Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 5, the above step 14 may include the following steps:
  • Step 141 Obtain a feature matrix of the gene mutation candidate site according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to the Base arrangement characteristics and non-base arrangement characteristics of gene mutation candidate sites, and the second dimension characteristics of the characteristic matrix correspond to the sites in the preset site interval;
  • Step 142 Identify the gene mutation of the gene mutation candidate site according to the feature matrix of the gene mutation candidate site.
  • the gene variation recognition model obtained based on neural network can be used to analyze the base arrangement feature and non-base arrangement feature.
  • the arrangement feature performs feature integration, and the base arrangement feature vector formed by the base arrangement feature and the non-base arrangement feature vector formed by the non-base arrangement feature are combined into a feature matrix.
  • the first dimensional feature of the feature matrix corresponds to base arrangement information and non-base arrangement information
  • the second dimensional feature corresponds to a site in the preset site interval.
  • the size of the feature matrix is the number of feature vectors ⁇ the size of the preset site interval.
  • the size of the feature matrix can be 16 ⁇ 150, where the first-dimensional feature corresponds to the 16-dimensional feature vector, and the first To 4 can correspond to the base arrangement feature, and the 5th to 16th dimension feature vectors can correspond to the non-base arrangement feature, and have the invariance of base arrangement.
  • the gene mutation identification model can be used to identify the gene mutation of the mutation candidate site according to the feature matrix.
  • the neural network model can be used to integrate base arrangement information and non-base arrangement information corresponding to gene mutation candidate sites, so that gene sequencing data can be analyzed more comprehensively, and gene mutation identification can be more accurate.
  • identifying the gene mutation at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site may include: according to the feature matrix of the gene mutation candidate site, Obtain the mutation value of the gene at the gene mutation candidate site, and if the mutation value is greater than or equal to a preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
  • the mutation value of the gene mutation may be used to characterize the possibility of true mutation at the candidate site of the gene mutation. For example, if the mutation value is greater, the possibility of true mutation at the candidate site of the gene mutation is greater.
  • the above-mentioned gene mutation recognition model can be used to process the obtained two-dimensional feature matrix to obtain the mutation value, and to determine whether the gene mutation at the gene mutation candidate site is a true mutation according to the mutation value.
  • the variation value can be between 0 and 1.
  • the preset threshold can be set according to the application scenario, for example, 0.3, 0.5. If the mutation value is greater than the preset threshold, the gene mutation at the candidate site of the gene mutation can be regarded as a true mutation, that is, the gene mutation caused by the disease; otherwise, The gene mutation at the candidate site of the gene mutation is a false mutation, that is, a gene abnormality caused by interference.
  • the embodiments of the present disclosure can use a gene mutation recognition model to identify gene mutations at candidate sites of gene mutations.
  • the base arrangement invariance of the gene data can be used to extract the gene mutation recognition model.
  • the feature matrix is subjected to matrix transformation, so that data enhancement can be performed during the model training process, so that the trained gene mutation recognition model has better robustness and reduces over-fitting problems.
  • Fig. 6 shows a flowchart of a process of obtaining a feature matrix of gene mutation candidate sites according to an embodiment of the present disclosure.
  • the data enhancement of base arrangement information can be applied in the training process of the gene mutation recognition model.
  • the feature matrix of gene mutation candidate sites is obtained, which may include:
  • Step 1411 Generate a feature vector of each first-dimensional feature in the preset site interval according to the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate sites;
  • Step 1412 Determine the base arrangement feature vector formed by the base arrangement feature in the feature vector
  • Step 1413 Randomly sort the base arrangement feature vector to obtain the feature matrix of the gene mutation candidate site.
  • the first dimension feature corresponds to the base arrangement information of the at least one gene sequencing read in the preset site interval
  • the feature vector of the first dimension feature may include a base arrangement feature vector formed by the base arrangement feature and Non-base arrangement feature vector formed by non-base arrangement feature. Since the non-base arrangement feature has base arrangement invariance, the non-base arrangement feature will not be affected after the arrangement order of the base arrangement feature vector is changed. Therefore, the base arrangement feature vector formed by the base arrangement feature in the feature vector can be randomly sorted to obtain the feature matrix of the gene mutation candidate site, realize the data enhancement processing of the base arrangement information, and make the gene mutation recognition obtained after training The model considers the nature of the invariance of the base arrangement and has better performance.
  • the first dimension feature corresponds to a 16-dimensional feature vector
  • the first to fourth dimension can correspond to the base arrangement feature
  • the fifth to 16th dimension feature vector can correspond to a non-base
  • the embodiments of the present disclosure extract the base arrangement characteristics and non-base arrangement characteristics of gene mutation candidate sites, so that the invariance of the base arrangement of the gene data can be considered when the gene mutation is identified, so that the recognition result of the gene mutation is more improved. Accurate, screen out germline gene mutations and interference caused by noise and errors, and improve the accuracy of gene mutation recognition.
  • the writing order of the steps does not mean a strict execution order but constitutes any limitation on the implementation process.
  • the specific execution order of each step should be based on its function and possibility.
  • the inner logic is determined.
  • Fig. 7 shows a block diagram of a gene mutation recognition device according to an embodiment of the present disclosure. As shown in Fig. 7, the gene mutation recognition device includes:
  • the first obtaining module 71 is configured to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
  • the second obtaining module 72 is configured to obtain the base arrangement characteristics of the candidate gene mutation sites
  • the determining module 73 is configured to determine the non-base arrangement characteristics of the gene mutation candidate site based on the non-base arrangement information of the at least one gene sequencing read in the preset site interval; wherein, the non-base arrangement The arrangement characteristics remain unchanged after the base arrangement sequence is changed;
  • the identification module 74 is configured to identify the gene mutation of the gene mutation candidate site based on the base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate site.
  • the second obtaining module 72 includes:
  • the first determining sub-module is used to determine the preset site interval where the candidate site of the gene mutation is located;
  • the second determining sub-module is used to obtain the base arrangement characteristics of the gene mutation candidate sites according to the base arrangement information of the reference genome in the preset site interval; wherein, the base arrangement characteristics are used for characterization The sequence of bases.
  • the determining module 73 includes:
  • the first acquisition sub-module is used to acquire the non-base arrangement information of each site of the at least one gene sequencing read in the preset site interval;
  • the third determining sub-module is used to determine the non-base arrangement characteristics of the gene mutation candidate sites based on the non-base arrangement information of each site in the preset site interval.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; according to the first gene sequencing read corresponding to each site in the preset site interval The number of reads of a gene sequence determines the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine the first gene sequencing read that has the same base type at the gene mutation candidate site and the reference genome; at each site in the preset site interval, determine The number of first gene sequencing reads in which the base type of the first gene sequencing read is inconsistent with the base type of the reference genome is used as the variation number of the first gene sequencing read; according to the first gene sequencing read Determine the non-base arrangement characteristics of the candidate site of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; according to each position in the preset site interval The number of second gene sequencing reads corresponding to the points determines the non-base arrangement characteristics of the gene mutation candidate sites.
  • the third determining submodule is specifically configured to:
  • the gene sequencing reads determine a second gene sequencing read that has the same mutation base type at the gene mutation candidate site and the gene mutation candidate site; in each of the preset site intervals Site, determine the number of second gene sequencing reads whose base type of the second gene sequencing read is inconsistent with the base type of the reference genome as the number of mutations of the second gene sequencing read; The number of mutations in the gene sequencing reads determines the non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the third determining submodule is specifically configured to:
  • the third gene sequencing read in the gene sequencing read; wherein the base type of the third gene sequencing read at the gene mutation candidate site is inconsistent with the base type of the reference genome, and the third gene The base type of the sequencing read at the gene mutation candidate site is inconsistent with the variant base type of the gene mutation candidate site; at each site in the preset site interval, the third gene sequencing read is determined The number of sequencing reads of the third gene whose base type is inconsistent with that of the reference genome is used as the number of variation of the third gene sequencing read; the number of variations of the third gene sequencing read is determined The non-base arrangement characteristics of the candidate sites of the gene mutation.
  • the third determining submodule is specifically configured to:
  • the third determining submodule is specifically configured to:
  • the identification module 74 includes:
  • the generation sub-module is used to obtain the feature matrix of the gene mutation candidate site according to the base arrangement feature and non-base arrangement feature of the gene mutation candidate site; wherein, the first dimension feature of the feature matrix corresponds to Base arrangement characteristics and non-base arrangement characteristics of the gene mutation candidate sites, and the second-dimensional characteristics of the characteristic matrix correspond to sites in the preset site interval;
  • the identification sub-module is used to identify the gene mutation of the gene mutation candidate site according to the feature matrix of the gene mutation candidate site.
  • the identification submodule is specifically used for:
  • the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
  • the generating submodule is specifically used for:
  • the base arrangement feature and non-base arrangement feature of the gene mutation candidate site generate a feature vector of each first dimension feature in the preset site interval; determine the formation of the base arrangement feature in the feature vector The base arrangement eigenvector of the base arrangement; random sorting is performed on the base arrangement eigenvector to obtain the characteristic matrix of the gene mutation candidate site.
  • the first obtaining module includes:
  • the second acquisition submodule is used to obtain the gene sequencing reads obtained by performing gene sequencing of somatic genes;
  • the comparison submodule is used to compare the base sequence of the gene sequencing reads with the base sequence of the reference genome , Obtain the comparison result;
  • the fourth determination sub-module used to determine the abnormal gene mutation candidate site of the gene of the somatic gene according to the comparison result;
  • the third acquisition sub-module used to obtain the gene mutation At least one gene sequencing read corresponding to the candidate site.
  • the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
  • the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
  • Fig. 8 is a block diagram showing a device 1900 for gene mutation identification according to an exemplary embodiment.
  • the device 1900 may be provided as a server.
  • the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions that can be executed by the processing component 1922, such as application programs.
  • the application program stored in the memory 1932 may include one or more modules each corresponding to a set of instructions.
  • the processing component 1922 is configured to execute instructions to perform the above-described methods.
  • the device 1900 may also include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to the network, and an input output (I/O) interface 1958.
  • the device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
  • a non-volatile computer-readable storage medium such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to complete the foregoing method.
  • the present disclosure may be a system, method, and/or computer program product.
  • the computer program product may include a computer-readable storage medium loaded with computer-readable program instructions for enabling a processor to implement various aspects of the present disclosure.
  • the computer-readable storage medium may be a tangible device that can hold and store instructions used by the instruction execution device.
  • the computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
  • Computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) Or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a printer with instructions stored thereon
  • RAM random access memory
  • ROM read-only memory
  • EPROM erasable programmable read-only memory
  • flash memory flash memory
  • SRAM static random access memory
  • CD-ROM compact disk read-only memory
  • DVD digital versatile disk
  • memory stick floppy disk
  • mechanical encoding device such as a printer with instructions stored thereon
  • the computer-readable storage medium used here is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (for example, light pulses through fiber optic cables), or through wires Transmission of electrical signals.
  • the computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing/processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and/or a wireless network.
  • the network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and/or edge servers.
  • the network adapter card or network interface in each computing/processing device receives computer-readable program instructions from the network, and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing/processing device .
  • the computer program instructions used to perform the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or in one or more programming languages.
  • Source code or object code written in any combination, the programming language includes object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" language or similar programming languages.
  • Computer-readable program instructions can be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly executed on a remote computer, or entirely on the remote computer or server carried out.
  • the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to access the Internet connection).
  • LAN local area network
  • WAN wide area network
  • an electronic circuit such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can be customized by using the status information of the computer-readable program instructions.
  • the computer-readable program instructions are executed to realize various aspects of the present disclosure.
  • These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device , A device that implements the functions/actions specified in one or more blocks in the flowchart and/or block diagram is produced. It is also possible to store these computer-readable program instructions in a computer-readable storage medium. These instructions make computers, programmable data processing apparatuses, and/or other devices work in a specific manner, so that the computer-readable medium storing instructions includes An article of manufacture, which includes instructions for implementing various aspects of the functions/actions specified in one or more blocks in the flowchart and/or block diagram.
  • each block in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and the module, program segment, or part of an instruction contains one or more functions for implementing the specified logical function.
  • Executable instructions may also occur in a different order from the order marked in the drawings. For example, two consecutive blocks can actually be executed in parallel, or they can sometimes be executed in the reverse order, depending on the functions involved.
  • each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or actions Or it can be realized by a combination of dedicated hardware and computer instructions.

Landscapes

  • Physics & Mathematics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Chemical & Material Sciences (AREA)
  • Analytical Chemistry (AREA)
  • Biophysics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Theoretical Computer Science (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Evolutionary Biology (AREA)
  • Genetics & Genomics (AREA)
  • Molecular Biology (AREA)
  • Epidemiology (AREA)
  • Public Health (AREA)
  • Primary Health Care (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

一种基因变异识别方法、装置和存储介质,其中,该方法包括:获取基因变异候选位点对应的至少一个基因测序读段(11);获取所述基因变异候选位点的碱基排列特征(12);基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变(13);基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别(14)。上述方法考虑非碱基排列特征不受碱基排列顺序制约的特点,更好地筛除由于胚系基因变异以及噪声、错误等干扰造成的伪基因变异,更好地对基因变异进行识别,提高基因变异识别的准确性。

Description

一种基因变异识别方法、装置和存储介质
本公开要求在2019年3月29日提交中国专利局、申请号为201910252747.9、申请名称为“一种基因变异识别方法、装置和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开涉及计算机技术领域,尤其涉及一种基因变异识别方法、装置和存储介质。
背景技术
随着生物技术的发展,通过基因测序技术可以测定人类基因的序列,碱基序列的分析可以作为进一步基因研究和改造的基础。目前,基因的二代测序技术相比于一代测试技术而言,极大地提高了基因测序的效率,降低了基因测序的成本,并且保持了基因测序的准确行性。第一代测试技术如果完成一个人类基因组的测序可能需要3年的时间,而使用二代测序技术则可以将时间缩短为仅仅1周。
发明内容
有鉴于此,本公开提出了一种基因变异识别技术方案。
根据本公开的一方面,提供了一种基因变异识别方法,所述方法包括:
获取基因变异候选位点对应的至少一个基因测序读段;获取所述基因变异候选位点的碱基排列特征;基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变;基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述获取所述基因变异候选位点的碱基排列特征,包括:
确定所述基因变异候选位点所在的预设位点区间;根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
在一种可能的实现方式中,所述基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致 的第二基因测序读段;根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别,包括:
根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位 点的基因变异进行识别,包括:
根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;
在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵,包括:
根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;确定所述特征向量中碱基排列特征形成的碱基排列特征向量;对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
在一种可能的实现方式中,获取基因变异候选位点对应的至少一个基因测序读段,包括:
获取由体细胞基因进行基因测序得到的基因测序读段;将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;获取所述基因变异候选位点对应的至少一个基因测序读段。
根据本公开的另一方面,提供了一种基因变异识别装置,所述装置包括:
第一获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;第二获取模块,用于获取所述基因变异候选位点的碱基排列特征;确定模块,用于基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变;识别模块,用于基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述第二获取模块,包括:
第一确定子模块,用于确定所述基因变异候选位点所在的预设位点区间;
第二确定子模块,用于根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
在一种可能的实现方式中,所述确定模块,包括:
第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;第三确定子模块,用于基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述识别模块,包括:
生成子模块,用于根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;识别子模块,用于根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述识别子模块,具体用于,
根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;
在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述生成子模块,具体用于,
根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;确定所述特征向量中碱基排列特征形成的碱基排列特征向量;对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
在一种可能的实现方式中,所述第一获取模块,包括:
第二获取子模块,用于获取由体细胞基因进行基因测序得到的基因测序读段;对比子模块,用于 将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;第四确定子模块,用于根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;第三获取子模块,用于获取所述基因变异候选位点对应的至少一个基因测序读段。
根据本公开的另一方面,提供了一种基因变异识别装置,包括:处理器;用于存储处理器可执行指令的存储器;其中,所述处理器被配置为执行上述方法。
根据本公开的另一方面,提供了一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其中,所述计算机程序指令被处理器执行时实现上述方法。
本公开实施例提供的基因变异识别方案,可以获取基因变异候选位点对应的至少一个基因测序读段,获取基因变异候选位点的碱基排列特征,基于至少一个基因测序读段在预设位点区间的碱基排列信息,确定基因变异候选位点的非碱基排列特征,从而可以基于基因变异候选位点的碱基排列特征和非碱基排列特征,对基因变异候选位点的基因变异进行识别。这里,非碱基排列特征在碱基排列顺序改变后保持不变,即可以认为非碱基排列特征具有碱基排列不变性的性质,因此,在对基因变异候选位点的基因变异进行识别时,可以考虑基因变异候选位点的基因变异不受碱基排列顺序制约的特点,更好地筛除由于胚系基因变异以及噪声、错误等干扰造成的伪基因变异,从而可以更好地对基因变异进行识别,提高基因变异识别的准确性,
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,而非限制本公开。
根据下面参考附图对示例性实施例的详细说明,本公开的其它特征及方面将变得清楚。
附图说明
包含在说明书中并且构成说明书的一部分的附图与说明书一起示出了本公开的示例性实施例、特征和方面,并且用于解释本公开的原理。
图1示出根据本公开一实施例的基因变异识别方法的流程图。
图2示出根据本公开一实施例的获取基因变异候选位点对应的至少一个基因测序读段的流程图。
图3示出根据本公开一实施例的基因变异候选位点的碱基排列特征过程的流程图。
图4示出根据本公开一实施例的基因变异候选位点的非碱基排列特征过程的流程图。
图5示出根据本公开一实施例的识别基因变异候选位点的基因变异过程的流程图。
图6示出根据本公开一实施例的得到基因变异候选位点的特征矩阵过程的流程图。
图7示出根据本公开一实施例的得到基因变异候选位点的特征矩阵过程的流程图。
图8示出根据本公开一实施例的得到基因变异候选位点的特征矩阵过程的流程图。
具体实施方式
以下将参考附图详细说明本公开的各种示例性实施例、特征和方面。附图中相同的附图标记表示功能相同或相似的元件。尽管在附图中示出了实施例的各种方面,但是除非特别指出,不必按比例绘制附图。
在这里专用的词“示例性”意为“用作例子、实施例或说明性”。这里作为“示例性”所说明的任何实施例不必解释为优于或好于其它实施例。
本文中术语“和/或”,仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。另外,本文中术语“至少一个”表示多种中的任意一个或多个中的至少两个的任意组合,例如,包括A、B、C中的至少一个,可以表示包括从A、B和C构成的集合中选择的任意一个或多个元素。
另外,为了更好地说明本公开,在下文的具体实施方式中给出了众多的具体细节。本领域技术人 员应当理解,没有某些具体细节,本公开同样可以实施。在一些实例中,对于本领域技术人员熟知的方法、手段、元件和电路未作详细描述,以便于凸显本公开的主旨。
本公开实施例提供的基因变异识别方案,可以获取基因变异候选位点对应的至少一个基因测序读段,从而可以利用至少一个基因测序读段对基因变异候选位点的基因变异进行识别。在基因变异识别过程中,可以确定基因变异候选位点的碱基排列特征,并根据至少一个基因测序读段在预设位点区间的碱基排列信息,确定基因变异候选位点的非碱基排列特征,然后可以通过碱基排列特征和非碱基排列特征对基因变异候选位点的基因变异进行识别。这里的非碱基排列特征在碱基排列顺序改变后保持不变,即可以认为基因变异候选位点的基因变异是否为真变异不受碱基排列顺序的影响,从而对基因变异候选位点的基因变异进行识别时,考虑基因数据的碱基排列不变性,提高基因变异识别的准确性。
在相关技术中,通常是利用支持向量机、随机森林等传统随机森林等传统机器学习方法进行基因变异识别,这种方式虽然实现简单,但基因变异识别的效果在基因数据量增加到一定程度之后会陷入瓶颈。还有一些相关技术采用深度学习方法,利用神经网络对基因变异进行识别。但是,神经网络提取的特征通常与碱基排列顺序相关,碱基排列顺序稍有不同就可能会得到不同的识别结果,造成神经网络过度拟合的问题。而本公开实施例提供的基因变异识别方案,考虑了基因数据的碱基排列不变性,利用基因变异识别模型提取基因变异候选位点的非碱基排列特征,使得到的识别结果不受碱基排列顺序的影响,提高基因变异识别模型的鲁棒性,缓解过度拟合的问题,减小基因变异识别模型训练的难度。下述实施例将会对基因变异识别过程作详细说明。
图1示出根据本公开一实施例的基因变异识别方法的流程图。该基因变异识别方法可以由基因变异识别装置或其它处理设备执行,其中,基因变异识别装置可以为用户设备(User Equipment,UE)、移动设备、用户终端、终端、蜂窝电话、无绳电话、个人数字处理(Personal Digital Assistant,PDA)、手持设备、计算设备、车载设备、可穿戴设备等,或者,基因变异识别装置可以为服务器。在一些可能的实现方式中,该基因变异识别方法可以通过处理器调用存储器中存储的计算机可读指令的方式来实现。如图1所示,该基因变异识别方法包括:
步骤11,获取基因变异候选位点对应的至少一个基因测序读段。
在本公开实施例中,基因变异识别装置可以获取由基因测序得到的基因测序读段,然后在基因测序得到的基因测序读段中,获取基因变异候选位点对应的至少一个基因测序读段。这里的基因测序读段可以理解为经过基因测序后标注有碱基类型的碱基序列,每个基因测序读段的长度可以相同也可以不同。在长度不同的情况下,每个基因测序读段的长度可以在预设长度范围内,从而可以保证每个基因测序读段的长度比较接近。碱基类型可以包括胞嘧啶(C)、鸟嘌呤(G)、腺嘌呤(A)、胸腺嘧啶(T),从而基因测序读段可以包括AGCT的碱基序列。这里的基因变异候选位点可以是碱基序列存在异常的位点。碱基序列的位点可以表示碱基序列的位置,针对每个位点,可以存在至少一个基因测序读段,即,在同一个位点可以存在由基因测序得到的至少一个基因测序读段。相应地,基因变异候选位点对应至少一个基因测序读段,其中,这至少一个基因测序读段都覆盖这一位点。基因变异候选位点可以为至少一个,每个基因变异候选位点可以对应至少一个基因测序读段。为了便于理解,本公开实施例以一个基因变异候选位点进行说明。
步骤12,获取所述基因变异候选位点的碱基排列特征。
在本公开实施例中,可以利用基因变异识别模型,根据基因变异候选位点的基因排列信息,提取基因变异候选位点的碱基排列特征。这里的碱基排列信息可以是与碱基排列顺序相关的信息,例如,某个基因测序读段在某个位点区间的碱基序列依次为A、C、G、T,则碱基排列信息可以为ACGT。 碱基排列信息可以包括预设位点区间内参考基因组的碱基类型、每种碱基类型的基因数量、每种碱基类型的缺失基因数量、每种碱基类型的插入基因数量等信息。由碱基排列信息得到的碱基排列特征与碱基排列顺序相关。
步骤13,基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变。
在本公开实施例中,在获取基因变异候选位点对应的至少一个基因测序读段之后,可以在预设位点区间,提取该基因变异候选位点对应的至少一个基因测序读段的碱基排列信息,并根据提取的碱基排列信息生成该基因变异候选位点的非碱基排列特征。非碱基排列信息可以是不受到碱基排列顺序限制的信息。从而可以根据至少一个基因测序读段在预设位点区间的非碱基排列信息,确定基因变异候选位点的非碱基排列特征。这里,非碱基排列信息可以包括位点处对应的基因测序读段的数量、在该位点发生变异的基因测序读段的数量等具有碱基排列不变性的信息。
这里,在提取非碱基排列信息时,可以随机选择该基因变异候选位点对应的若干个基因测序读段,提取随机选择的若干个基因测序读段的非碱基排列信息;还可以提取该基因变异候选位点对应的每个基因测序读段的非碱基排列信息。在提取至少一个基因测序读段在预设位点区间的非碱基排列信息时,可以提取至少一个基因测序读段在该预设位点区间内每个位点的非碱基排列信息,还可以随机选择该预设位点区间内若干个相邻位点,提取至少一个基因测序读段在若干个相邻位点的非碱基排列信息。在确定所述基因变异候选位点的非碱基排列特征时,可以利用基于神经网络训练得到的基因变异识别模型。
步骤14,基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
在本公开实施方式中,在确定碱基排列特征和非碱基排列特征之后,可以由碱基排列特征和非碱基排列特征得到基因变异候选位点的特征矩阵中,利用该特征矩阵对该基因变异候选位点的基因变异进行识别,例如,可以利用上述基因变异识别模型判断该基因变异候选位点的基因是否是由于病变引起的真变异,还是由于噪声等原因而导致的碱基序列异常的假变异。这里,得到的基因变异候选位点的特征矩阵可以为二维特征矩阵,特征矩阵的尺寸可以是特征向量的个数×预设位点区间的大小,其中的特征向量可以是基于碱基排列特征和非碱基排列特征生成的。由于变异候选位点的基因变异是否为病变引起的真基因变异不受碱基排列顺序的影响,更多地受到基因变异候选位点所在的基因环境的影响,例如,受到基因变异候选位点附近的其他位点存在变异基因等基因环境的影响,从而得到的特征矩阵中对应于碱基排列特征的特征向量的排列顺序可以不受限制,碱基排列特征的特征向量在特征矩阵中的排列顺序可以随机变动,提高基因变异识别的效率和准确率。
本公开实施例中可以根据基因变异候选位点的碱基排列特征和非碱基排列特征对基因变异候选位点的基因变异进行识别,从而可以考虑基因变异的碱基排列不变性,更好地对基因变异进行识别。在对基因变异候选位点的基因变异进行识别时,可以获取基因变异候选位点对应的至少一个基因测序读段。本公开实例还提供了一种获取基因变异候选位点对应的至少一个基因测序读段的过程。
图2示出根据本公开一实施例的获取基因变异候选位点对应的至少一个基因测序读段的流程图。在一种可能的实现方式中,获取基因变异候选位点对应的至少一个基因测序读段,可以包括以下步骤:
步骤111,获取由体细胞基因进行基因测序得到的基因测序读段。
这里,通过体细胞基因进行基因测序可以得到至少一个基因测序读段,基因测序读段可以是对体细胞基因进行碱基类型标注的序列。体细胞基因在进行基因测序之后,不仅可以得到基因测序读段中 每个基因的碱基类型,还可以得到基因测序读段中每个基因所在位点的基因位置信息。同一个位点可以对应至少一个基因测序读段。
在一种可能的实现方式中,通过体细胞基因进行基因测序可以得到至少一个基因测序读段,可以对基因测序得到的基因测序读段进行预处理,这里的预处理方式可以包括交叉污染筛选、测序质量筛选、比对质量筛选、读段长度异常筛选等。通过预处理,可以筛选掉交叉污染的基因测序读段,以及筛选掉测序质量和比对质量较低、读段长度异常的基因测序读段。
步骤112,将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果。
在本公开实施例中,在获取由体细胞基因进行基因测序得到的基因测序读段之后,可以将获取的基因测序读段的碱基序列与相同位点的参考基因组的碱基序列进行比对,得到对比结果。举例来说,可以将每个进行基因测序得到的基因测序读段与相同位点的参考基因组的碱基序列进行对比,确定基因测序读段的碱基序列与参考基因组的碱基序列不同的位点。还可以将具有相同位点的至少一个基因测序读段与相同位点的参考基因组的碱基序列进行对比,确定至少一个基因测序读段的碱基序列与参考基因组的碱基序列不同的位点。这里,参考基因组可以是标注有正确碱基序列的碱基序列。
步骤113,根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点。
在本公开实施例中,可以根据比对结果确定基因测序读段与参考基因组的碱基序列不同的位点,如果该位点对应的至少一个基因测序读段中,在该位点发送变异的基因测序读段的比例大于预设比例,则可以确定该位点为基因变异候选位点,否则,可以认为该位点不是基因变异候选位点。基因测序读段在该位点与参考基因组的碱基序列不同,可能是因为测序错误导致的不同,通过这种方式,可以减少由于基因测序失误引起的碱基序列异常现象。
步骤114,获取所述基因变异候选位点对应的至少一个基因测序读段。
在本公开实施例中,在确定基因变异候选位点之后,可以获取基因变异候选位点对应的至少一个基因测序读段。其中,每个基因变异候选位点对应的至少一个基因测序读段,在该基因变异候选位点的碱基序列与相同位点的参考基因组的碱基序列可以不同。这里的基因变异候选位点可以为至少一个。
通过上述获取基因变异候选位点对应的至少一个基因测序读段的过程,不仅可以较为准确地确定基因变异候选位点,还可以在基因测序得到的基因测序读段中确定基因变异候选位点对应的至少一个基因测序读段。
本公开实施例中可以根据基因变异候选位点对应的至少一个基因测序读段的碱基排列信息,确定该基因变异候选位点的碱基排列特征,从而在识别基因变异候选位点的基因变异时,可以根据该碱基排列特征对基因识别进行数据增强处理。下面通过一示例对确定基因变异候选位点的碱基排列特征的过程进行详细说明。
图3示出根据本公开一实施例的基因变异候选位点的碱基排列特征过程的流程图。如图3所示,上述步骤12可以包括以下步骤:
步骤121,确定所述基因变异候选位点所在的预设位点区间;
步骤122,根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
在本公开实施例的示例中,每一个基因变异候选位点可以存在至少一个基因测序读段。为了提高基因变异识别的准确度,不仅可以考虑该基因变异候选位点的碱基排列信息,还可以考虑该基因变异候选位点附近的位点的碱基排列信息。这里,碱基排列信息可以包括候选基因组的碱基排列信息,在碱基排列信息为候选基因组的碱基排列信息的情况下,可以认为每个基因测序读段的碱基排列信息相 同,均为候选基因组的碱基排列信息。从而可以根据基因变异候选位点的基因位置信息,确定该基因变异候选位点所在的预设位点区间,例如,可以将基因变异候选位点前后150个碱基形成的区间作为基因变异候选位点所在的预设位点区间。然后可以针对该预设位点区间内的每个位点,获取参考基因组在预设位点区间的碱基排列信息,由参考基因组在预设位点区间的碱基排列信息生成基因变异候选位点的碱基排列特征。碱基排列信息可以参考基因组在预设位点区间中每个位点的碱基序列组成,例如,预设位点区间包括4个碱基序列,分别为A、C、G、T,则碱基排列信息可以为ACGT的碱基排列顺序。碱基排列特征可以用碱基排列特征向量进行表示,可以是基因变异候选位点的特征矩阵的一部分,例如,如果表征碱基排列信息的碱基排列特征向量为4个,分别为a1、a2、a3和a4,则a1、a2、a3和a4可以为特征矩阵的前4维特征。
本公开实施例的示例中不仅在对基因变异候选位点的基因变异进行识别时,考虑了基因变异候选位点所对应的碱基排列特征,还考虑了基因变异候选位点具有碱基排列不变性的非碱基排列特征。下面通过一示例对确定基因变异候选位点的非碱基排列特征的过程进行详细说明。
图4示出根据本公开一实施例的基因变异候选位点的非碱基排列特征过程的流程图。如图4所示,上述步骤13可以包括以下步骤:
步骤131,获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;
步骤132,基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在本公开实施例的示例中,考虑到基因数据具有碱基排列不变性的性质,从而可以在基因变异识别过程中,获取至少一个基因测序读段在预设位点区间中每个位点的非碱基排列信息。这里,非碱基排列信息可以是具有碱基排列不变性的信息,例如,位点处对应的基因测序读段的数量、变异数量。非碱基排列信息可以为多种,相应地,每种非碱基排列信息生成的非碱基排列特征可以形成一个非碱基排列特征向量,非碱基排列特征向量可以为一个或多个。
本公开实施例提供的基因变异识别方案可以应用于已经确诊为患有癌症的病人,通过基因变异识别可以为病人指导用药。因此,基因测序读段中的一部分基因测序读段可以来源于正常细胞,正常细胞可以认为是没有发生病变的细胞。还有一部分基因测序读段可以来源于病变细胞。从而在确定基因变异候选位点的非碱基排列特征时,可以分别基于来源于正常细胞的基因测序读段和来源于病变细胞的基因测序读段,确定基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,确定基因变异候选位点的非碱基排列特征时,可以确定至少一个基因测序读段中来源于正常细胞的基因测序读段,然后基于正常细胞的基因测序读段在预设位点区间中每个位点的非碱基排列信息,确定基因变异候选位点的非碱基排列特征。这样,可以基于来源于正常细胞的基因测序读段确定基因变异候选位点的非碱基排列特征。
下面提供了基于正常细胞的基因测序读段确定基因变异候选位点的非碱基排列特征的几个示例。
在该公开实施例的一个示例中,在确定基因变异候选位点的非碱基排列特征时,可以在基因测序读段中,确定在基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段,然后根据预设位点区间中每个位点对应的第一基因测序读段的数量,确定基因变异候选位点的非碱基排列特征。
在该示例中,可以在基因测序读段中选择在基因变异候选位点未发生基因变异的第一基因测序读段,针对预设位点区间中的每个位点,可以统计第一基因测序读段在该位点的数量。换言之,可以统计有多少个第一基因测序读段包含该位点。其中,包含某一位点的第一基因测序读段可认为是该位点对应的第一基因测序读段。由于每个基因测序读段的长度可能不同,基因变异候选位点相对于每个基 因测序读段的位置不同,例如,基因变异候选位点可以位于基因测序读段的中间位置,还可以位于基因测序读段的边缘位置,从而预设位点区间中的每个位点所对应的基因测序读段的数量不同。由每个位点对应的第一基因测序读段的数量,可以生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱基排列特征向量中的每个特征元素可以对应相应位点的第一基因测序读段的数量。
在该公开实施例的另一个示例中,在确定基因变异候选位点的非碱基排列特征时,可以在基因测序读段中,确定在基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段,然后在预设位点区间中的每个位点,确定第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量,根据第一基因测序读段的变异数量,确定基因变异候选位点的非碱基排列特征。
在该示例中,可以在基因测序读段中选择在基因变异候选位点未发生基因变异的第一基因测序读段,针对预设位点区间中的每个位点,可以统计第一基因测序读段在该位点发生基因变异的变异数量。这里,虽然基因测序读段在基因变异候选位点未发生基因变异(即在基因变异候选位点与参考基因组的碱基类型一致),但是可能在基因变异候选位点之外的其他位点发生基因变异(即在其他位点与参考基因组的碱基类型不一致),从而可以针对预设位点区间的每个位点,统计在该位点的第一基因测序读段中发生变异的变异数量。换言之,针对每个位点,可以统计包含该位点的第一基因测序读段中,有多少个第一基因测序读段在该位点发生变异。由每个位点对应的第一基因测序读段中发生变异的变异数量,可以生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱基排列特征向量中的每个特征元素可以对应相应位点的第一基因测序读段的变异数量,换言之,包含该相应位点且在该相应位点发生变异的第一基因测序读段的数量。
举例来说,针对来源于正常细胞的基因测序读段,可以确定正常细胞的基因测序读段中在基因变异候选位点未发生变异的第一基因测序读段,然后针对预设位点区间中的每个位点,统计每个位点对应的第一基因测序读段的数量和在该位点发生变异的数量,这两个信息可以对应于上述特征矩阵中的第5维特征和第6维特征。
在该公开实施例的另一个示例中,在确定基因变异候选位点的非碱基排列特征时,可以在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段,然后根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。在该示例中,可以在基因测序读段中选择与基因变异候选位点变异一致的第二基因测序读段,针对预设位点区间中的每个位点,可以统计第二基因测序读段在该位点的数量。由每个位点对应的第二基因测序读段的数量,生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱基排列特征向量中的每个特征元素可以对应相应位点的第二基因测序读段的数量。
在该公开实施例的另一个示例中,在确定基因变异候选位点的非碱基排列特征时,可以在所述基因测序读段中,确定在基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段,然后在预设位点区间中的每个位点,确定第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量,根据第二基因测序读段的变异数量,确定基因变异候选位点的非碱基排列特征。在该示例中,可以在基因测序读段中选择与基因变异候选位点变异一致的第二基因测序读段(基因变异候选位点的变异碱基类型可通过基因测序得到),针对预设位点区间中的每个位点,统计第二基因测序读段在该位点发生基因变异的变异数量,换言之,统计包含该位点且在该位点发生变异的第二基因测序读段的数量。每个位点对应的第二基因测序读段中发生变异的变异数量,可以生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱 基排列特征向量中的每个特征元素可以对应相应位点的第二基因测序读段的变异数量。
举例来说,针对来源于正常细胞的基因测序读段,可以在正常细胞的基因测序读段中选择与基因变异候选位点变异一致的第二基因测序读段,然后针对预设位点区间中的每个位点,统计每个位点对应的第二基因测序读段的数量和在该位点发生变异的数量,这两个信息可以对应于上述特征矩阵中的第7维特征和第8维特征。
在该公开实施例的另一个示例中,在确定所述基因变异候选位点的非碱基排列特征时,可以确定基因测序读段中的第三基因测序读段,然后根据预设位点区间中每个位点对应的第三基因测序读段的数量,确定基因变异候选位点的非碱基排列特征。这里,第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致,即,第三基因序读段是基因序读段中除去第一基因序读段和第二基因序读段的剩余基因序读段。第三基因测序读段可以是在基因变异候选位点存在插入基因、缺失基因等情况的基因测序读段。在该示例中,可以在基因测序读段中确定剩余的第三基因测序读段,针对预设位点区间中的每个位点,可以统计第三基因测序读段在该位点的数量。由每个位点对应的第三基因测序读段的数量,生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱基排列特征向量中的每个特征元素可以对应相应位点的第三基因测序读段的数量。
在该公开实施例的另一个示例中,在确定基因变异候选位点的非碱基排列特征时,可以确定基因测序读段中的第三基因测序读段,然后在预设位点区间中的每个位点,确定第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量,根据第三基因测序读段的变异数量,确定基因变异候选位点的非碱基排列特征。这里,第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致,即,第三基因序读段是基因序读段中除去第一基因序读段和第二基因序读段的剩余基因序读段。在该示例中,可以在基因测序读段中确定剩余的第三基因测序读段,针对预设位点区间中的每个位点,统计第三基因测序读段在该位点发生基因变异的变异数量。每个位点对应的第三基因测序读段中发生变异的变异数量,可以生成非碱基排列特征对应的一个非碱基排列特征向量,该非碱基排列特征向量中的每个特征元素可以对应相应位点的第三基因测序读段的变异数量。
举例来说,针对来源于正常细胞的基因测序读段,可以在正常细胞的基因测序读段中选择除第一基因测序读段和第二基因测序读段之外的第三基因测序读段,然后针对预设位点区间中的每个位点,统计每个位点对应的第三基因测序读段的数量和在该位点发生变异的数量,这两个信息可以对应于上述特征矩阵中的第9维特征和第10维特征。
在一种可能的实现方式中,确定基因变异候选位点的非碱基排列特征时,可以确定至少一个基因测序读段中来源于病变细胞的基因测序读段,然后基于病变细胞的基因测序读段在预设位点区间中每个位点的非碱基排列信息,确定基因变异候选位点的非碱基排列特征。这样,可以基于来源于病变细胞的基因测序读段确定基因变异候选位点的非碱基排列特征。
在该实现方式中,基于病变细胞的基因测序读段确定基因变异候选位点的非碱基排列特征的过程,可以参见上述正常细胞的基因测序读段确定非碱基排列特征的过程。举例来说,针对来源于病变细胞的基因测序读段,可以在病变细胞的基因测序读段中确定第一基因测序读段、第二基因测序读段和第三基因测序读段,然后针对预设位点区间中的每个位点,统计每个位点对应的第一基因测序读段的数量和变异数量、第二基因测序读段的数量和变异数量和第三基因测序读段的数量和变异数量,这些信 息可以对应于上述特征矩阵中的第11至16维特征。
通过上述方式,可以针对至少一个基因测序读段在预设位点区间与碱基排列相关的非碱基排列信息,确定基因变异候选位点的非碱基排列特征,从而可以在基因变异识别时考虑基因数据的碱基排列不变性,使基因变异识别更加容易、准确。下面通过一示例对基因变异候选位点的基因变异进行识别的过程进行说明。
图5示出根据本公开一实施例的识别基因变异候选位点的基因变异过程的流程图。如图5所示,上述步骤14可以包括以下步骤:
步骤141,根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;
步骤142,根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
在本公开实施例的示例中,在确定基因变异候选位点的碱基排列特征和非碱基排列特征之后,可以利用基于神经网络得到的基因变异识别模型,对碱基排列特征和非碱基排列特征进行特征整合,将碱基排列特征形成的碱基排列特征向量与非碱基排列特征形成的非碱基排列特征向量合成一个特征矩阵。特征矩阵的第一维度特征对应于碱基排列信息和非碱基排列信息,第二维度特征对应于所述预设位点区间的位点。特征矩阵的尺寸是特征向量的个数×预设位点区间的大小。举例来说,若特征向量的个数为16,预设位点区间包括150个位点,则特征矩阵的尺寸可以为16×150,其中,第一维度特征对应于16维特征向量,第1至4为可以对应于碱基排列特征,第5至16维特征向量可以对应于非碱基排列特征,具有碱基排列不变性。然后可以利用上述基因变异识别模型根据该特征矩阵对变异候选位点的基因变异进行识别。通过这种方式,可以利用神经网络模型整合基因变异候选位点对应的碱基排列信息和非碱基排列信息,从而可以更加全面地对基因测序数据进行分析,使基因变异识别更加准确。
在一种可能的实现方式中,根据所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别,可以包括:根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值,在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。这里,基因发生变异的变异值可以是表征该基因变异候选位点发生真变异的可能性,举例来说,如果变异值越大,该基因变异候选位点发生真变异的可能性越大。可以利用上述基因变异识别模型对得到的二维的特征矩阵进行处理得到变异值,并根据变异值判断基因变异候选位点的基因变异是否为真变异。在一种可能的实现方式中,变异值可以在0至1之间。预设阈值可以根据应用场景进行设置,例如,0.3、0.5,如果变异值大于预设阈值,则可以认为该基因变异候选位点的基因变异为真变异,即为病变引起的基因变异;否则,可以为该基因变异候选位点的基因变异为假变异,即是干扰形成的基因异常。
本公开实施例可以利用基因变异识别模型对基因变异候选位点的基因变异进行识别,该基因变异识别模型在训练过程中,可以利用基因数据的碱基排列不变性,将基因变异识别模型提取的特征矩阵进行矩阵变换,从而可以在模型训练过程中进行数据增强处理,使训练的基因变异识别模型具有更好地鲁棒性,减少过度拟合等问题。
图6示出根据本公开一实施例的得到基因变异候选位点的特征矩阵过程的流程图。
在本公开实施例中,可以将碱基排列信息的数据增强应用在基因变异识别模型的训练过程中。如图6所示,根据基因变异候选位点的碱基排列特征和非碱基排列特征,得到基因变异候选位点的特征 矩阵,可以包括:
步骤1411,根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;
步骤1412,确定所述特征向量中碱基排列特征形成的碱基排列特征向量;
步骤1413,对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
这里,第一维度特征对应于所述至少一个基因测序读段在预设位点区间的碱基排列信息,第一维度特征的特征向量可以包括由碱基排列特征形成的碱基排列特征向量和由非碱基排列特征形成的非碱基排列特征向量。由于非碱基排列特征具有碱基排列不变性,从而在碱基排列特征向量的排列顺序改变之后,非碱基排列特征不会受到影响。因此,可以将特征向量中碱基排列特征形成的碱基排列特征向量进行随机排序,得到基因变异候选位点的特征矩阵,实现碱基排列信息的数据增强处理,使训练后得到的基因变异识别模型考虑碱基排列不变性的性质,具有更优越的性能。
举例来说,若特征向量的个数为16,第一维度特征对应于16维特征向量,第1至4为可以对应于碱基排列特征,第5至16维特征向量可以对应于非碱基排列特征,则可以将第1至4的特征向量进行随机排序,形成多个特征矩阵。
本公开实施例通过提取基因变异候选位点的碱基排列特征和非碱基排列特征,从而在基因变异进行识别时可以考虑基因数据的碱基排列不变性,使基因变异进行识别的识别结果更加准确,筛掉胚系基因变异以及由于噪声和错误带来的干扰,提高基因变异识别的准确率。
本领域技术人员可以理解,在具体实施方式的上述方法中,各步骤的撰写顺序并不意味着严格的执行顺序而对实施过程构成任何限定,各步骤的具体执行顺序应当以其功能和可能的内在逻辑确定。
图7示出根据本公开实施例的基因变异识别装置的框图,如图7所示,所述基因变异识别装置包括:
第一获取模块71,用于获取基因变异候选位点对应的至少一个基因测序读段;
第二获取模块72,用于获取所述基因变异候选位点的碱基排列特征;
确定模块73,用于基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变;
识别模块74,用于基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述第二获取模块72,包括:
第一确定子模块,用于确定所述基因变异候选位点所在的预设位点区间;
第二确定子模块,用于根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
在一种可能的实现方式中,所述确定模块73,包括:
第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;
第三确定子模块,用于基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述第三确定子模块,具体用于,
确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
在一种可能的实现方式中,所述识别模块74,包括:
生成子模块,用于根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;
识别子模块,用于根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异 进行识别。
在一种可能的实现方式中,所述识别子模块,具体用于,
根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;
在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述生成子模块,具体用于,
根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;确定所述特征向量中碱基排列特征形成的碱基排列特征向量;对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
在一种可能的实现方式中,所述第一获取模块,包括:
第二获取子模块,用于获取由体细胞基因进行基因测序得到的基因测序读段;对比子模块,用于将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;第四确定子模块,用于根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;第三获取子模块,用于获取所述基因变异候选位点对应的至少一个基因测序读段。
在一些实施例中,本公开实施例提供的装置具有的功能或包含的模块可以用于执行上文方法实施例描述的方法,其具体实现可以参照上文方法实施例的描述,为了简洁,这里不再赘述。
图8是根据一示例性实施例示出的一种用于基因变异识别装置1900的框图。例如,装置1900可以被提供为一服务器。参照图8,装置1900包括处理组件1922,其进一步包括一个或多个处理器,以及由存储器1932所代表的存储器资源,用于存储可由处理组件1922的执行的指令,例如应用程序。存储器1932中存储的应用程序可以包括一个或一个以上的每一个对应于一组指令的模块。此外,处理组件1922被配置为执行指令,以执行上述方法。
装置1900还可以包括一个电源组件1926被配置为执行装置1900的电源管理,一个有线或无线网络接口1950被配置为将装置1900连接到网络,和一个输入输出(I/O)接口1958。装置1900可以操作基于存储在存储器1932的操作系统,例如Windows ServerTM,Mac OS XTM,UnixTM,LinuxTM,FreeBSDTM或类似。
在示例性实施例中,还提供了一种非易失性计算机可读存储介质,例如包括计算机程序指令的存储器1932,上述计算机程序指令可由装置1900的处理组件1922执行以完成上述方法。
本公开可以是系统、方法和/或计算机程序产品。计算机程序产品可以包括计算机可读存储介质,其上载有用于使处理器实现本公开的各个方面的计算机可读程序指令。
计算机可读存储介质可以是可以保持和存储由指令执行设备使用的指令的有形设备。计算机可读存储介质例如可以是――但不限于――电存储设备、磁存储设备、光存储设备、电磁存储设备、半导体存储设备或者上述的任意合适的组合。计算机可读存储介质的更具体的例子(非穷举的列表)包括:便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、静态随机存取存储器(SRAM)、便携式压缩盘只读存储器(CD-ROM)、数字多功能盘(DVD)、记忆棒、软盘、机械编码设备、例如其上存储有指令的打孔卡或凹槽内凸起结构、以及上述的任意合适的组合。这里所使用的计算机可读存储介质不被解释为瞬时信号本身,诸如无线电波或者其他自由传播的电磁波、通过波导或其他传输媒介传播的电磁波(例如,通过光纤电缆的光脉冲)、或者通过电线传输的电信号。
这里所描述的计算机可读程序指令可以从计算机可读存储介质下载到各个计算/处理设备,或者通过网络、例如因特网、局域网、广域网和/或无线网下载到外部计算机或外部存储设备。网络可以 包括铜传输电缆、光纤传输、无线传输、路由器、防火墙、交换机、网关计算机和/或边缘服务器。每个计算/处理设备中的网络适配卡或者网络接口从网络接收计算机可读程序指令,并转发该计算机可读程序指令,以供存储在各个计算/处理设备中的计算机可读存储介质中。
用于执行本公开操作的计算机程序指令可以是汇编指令、指令集架构(ISA)指令、机器指令、机器相关指令、微代码、固件指令、状态设置数据、或者以一种或多种编程语言的任意组合编写的源代码或目标代码,所述编程语言包括面向对象的编程语言—诸如Smalltalk、C++等,以及常规的过程式编程语言—诸如“C”语言或类似的编程语言。计算机可读程序指令可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络—包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。在一些实施例中,通过利用计算机可读程序指令的状态信息来个性化定制电子电路,例如可编程逻辑电路、现场可编程门阵列(FPGA)或可编程逻辑阵列(PLA),该电子电路可以执行计算机可读程序指令,从而实现本公开的各个方面。
这里参照根据本公开实施例的方法、装置(系统)和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其它可编程数据处理装置的处理器,从而生产出一种机器,使得这些指令在通过计算机或其它可编程数据处理装置的处理器执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
也可以把计算机可读程序指令加载到计算机、其它可编程数据处理装置、或其它设备上,使得在计算机、其它可编程数据处理装置或其它设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其它可编程数据处理装置、或其它设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,所述模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实施例,上述说明是示例性的,并非穷尽性的,并且也不限于所披露的各实施例。在不偏离所说明的各实施例的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实施例的原理、实际应用或对市场中技术的技术改进,或者使本技术领域的其它普通技术人员能理解本文披露的各实施例。

Claims (32)

  1. 一种基因变异识别方法,其特征在于,所述方法包括:
    获取基因变异候选位点对应的至少一个基因测序读段;
    获取所述基因变异候选位点的碱基排列特征;
    基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征碱基排列顺序改变后保持不变;
    基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
  2. 根据权利要求1所述的方法,其特征在于,所述获取所述基因变异候选位点的碱基排列特征,包括:
    确定所述基因变异候选位点所在的预设位点区间;
    根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
  3. 根据权利要求1或2所述的方法,其特征在于,所述基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;
    基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  4. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;
    根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  5. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;
    在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;
    根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  6. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;
    根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  7. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致 的第二基因测序读段;
    在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;
    根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  8. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;
    根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  9. 根据权利要求3所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;
    在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;
    根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  10. 根据权利要求3至9中任意一项所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;
    基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  11. 根据权利要求3至9中任意一项所述的方法,其特征在于,所述基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征,包括:
    确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;
    基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  12. 根据权利要求1至11中任意一项所述的方法,其特征在于,所述基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别,包括:
    根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;
    根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
  13. 根据权利要求12所述的方法,其特征在于,所述根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别,包括:
    根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;
    在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
  14. 根据权利要求12所述的方法,其特征在于,所述根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵,包括:
    根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;
    确定所述特征向量中碱基排列特征形成的碱基排列特征向量;
    对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
  15. 根据权利要求1至14中任意一项所述的方法,其特征在于,获取基因变异候选位点对应的至少一个基因测序读段,包括:
    获取由体细胞基因进行基因测序得到的基因测序读段;
    将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;
    根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
    获取所述基因变异候选位点对应的至少一个基因测序读段。
  16. 一种基因变异识别装置,其特征在于,所述装置包括:
    第一获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;
    第二获取模块,用于获取所述基因变异候选位点的碱基排列特征;
    确定模块,用于基于所述至少一个基因测序读段在预设位点区间的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征;其中,所述非碱基排列特征在碱基排列顺序改变后保持不变;
    识别模块,用于基于所述基因变异候选位点的碱基排列特征和非碱基排列特征,对所述基因变异候选位点的基因变异进行识别。
  17. 根据权利要求16所述的装置,其特征在于,所述第二获取模块,包括:
    第一确定子模块,用于确定所述基因变异候选位点所在的预设位点区间;
    第二确定子模块,用于根据参考基因组在所述预设位点区间的碱基排列信息,获取所述基因变异候选位点的碱基排列特征;其中,所述碱基排列特征用于表征碱基排列顺序。
  18. 根据权利要求16或17所述的装置,其特征在于,所述确定模块,包括:
    第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的非碱基排列信息;
    第三确定子模块,用于基于所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  19. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;
    根据所述预设位点区间中每个位点对应的第一基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  20. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    在所述基因测序读段中,确定在所述基因变异候选位点与参考基因组的碱基类型一致的第一基因测序读段;
    在所述预设位点区间中的每个位点,确定所述第一基因测序读段的碱基类型与参考基因组的碱基类型不一致的第一基因测序读段的数量,作为第一基因测序读段的变异数量;
    根据所述第一基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  21. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;
    根据所述预设位点区间中每个位点对应的第二基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  22. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    在所述基因测序读段中,确定在所述基因变异候选位点与基因变异候选位点的变异碱基类型一致的第二基因测序读段;
    在所述预设位点区间中的每个位点,确定所述第二基因测序读段的碱基类型与参考基因组的碱基类型不一致的第二基因测序读段的数量,作为第二基因测序读段的变异数量;
    根据所述第二基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  23. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;
    根据所述预设位点区间中每个位点对应的第三基因测序读段的数量,确定所述基因变异候选位点的非碱基排列特征。
  24. 根据权利要求18所述的装置,其特征在于,所述第三确定子模块,具体用于,
    确定所述基因测序读段中的第三基因测序读段;其中,所述第三基因测序读段在基因变异候选位点的碱基类型与参考基因组的碱基类型不一致,并且,第三基因测序读段在基因变异候选位点的碱基类型与基因变异候选位点的变异碱基类型不一致;
    在所述预设位点区间中的每个位点,确定所述第三基因测序读段的碱基类型与参考基因组的碱基类型不一致的第三基因测序读段的数量,作为所述第三基因测序读段的变异数量;
    根据所述第三基因测序读段的变异数量,确定所述基因变异候选位点的非碱基排列特征。
  25. 根据权利要求18至24中任意一项所述的装置,其特征在于,所述第三确定子模块,具体用于,
    确定所述至少一个基因测序读段中来源于正常细胞的基因测序读段;
    基于所述正常细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  26. 根据权利要求18至24中任意一项所述的装置,其特征在于,所述第三确定子模块,具体用于,
    确定所述至少一个基因测序读段中来源于病变细胞的基因测序读段;
    基于所述病变细胞的基因测序读段在所述预设位点区间中每个位点的非碱基排列信息,确定所述基因变异候选位点的非碱基排列特征。
  27. 根据权利要求16至26中任意一项所述的装置,其特征在于,所述识别模块,包括:
    生成子模块,用于根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,得到所述基因变异候选位点的特征矩阵;其中,所述特征矩阵的第一维度特征对应于所述基因变异候选位点的碱基排列特征和非碱基排列特征,所述特征矩阵的第二维度特征对应于所述预设位点区间的位点;
    识别子模块,用于根据所述基因变异候选位点的特征矩阵,对所述基因变异候选位点的基因变异进行识别。
  28. 根据权利要求27所述的装置,其特征在于,所述识别子模块,具体用于,
    根据所述基因变异候选位点的特征矩阵,得到所述基因变异候选位点的基因发生变异的变异值;
    在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
  29. 根据权利要求27所述的装置,其特征在于,所述生成子模块,具体用于,
    根据所述基因变异候选位点的碱基排列特征和非碱基排列特征,生成所述预设位点区间的每个第一维度特征的特征向量;
    确定所述特征向量中碱基排列特征形成的碱基排列特征向量;
    对所述碱基排列特征向量进行随机排序,得到所述基因变异候选位点的特征矩阵。
  30. 根据权利要求16至29中任意一项所述的装置,其特征在于,所述第一获取模块,包括:
    第二获取子模块,用于获取由体细胞基因进行基因测序得到的基因测序读段;
    对比子模块,用于将所述基因测序读段的碱基序列与参考基因组的碱基序列进行比对,得到比对结果;
    第四确定子模块,用于根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
    第三获取子模块,用于获取所述基因变异候选位点对应的至少一个基因测序读段。
  31. 一种基因变异识别装置,其特征在于,包括:
    处理器;
    用于存储处理器可执行指令的存储器;
    其中,所述处理器被配置为:执行权利要求1至15中任意一项所述的方法。
  32. 一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其特征在于,所述计算机程序指令被处理器执行时实现权利要求1至15中任意一项所述的方法。
PCT/CN2019/089504 2019-03-29 2019-05-31 一种基因变异识别方法、装置和存储介质 Ceased WO2020199337A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
SG11202101410WA SG11202101410WA (en) 2019-03-29 2019-05-31 Genovariation identification method and device, and storage medium
JP2021517044A JP7064655B2 (ja) 2019-03-29 2019-05-31 遺伝子変異認識方法、装置および記憶媒体
US17/162,465 US20210151124A1 (en) 2019-03-29 2021-01-29 Genetic variation identification method, genetic variation identification apparatuses, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910252747.9A CN109979531B (zh) 2019-03-29 2019-03-29 一种基因变异识别方法、装置和存储介质
CN201910252747.9 2019-03-29

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/162,465 Continuation US20210151124A1 (en) 2019-03-29 2021-01-29 Genetic variation identification method, genetic variation identification apparatuses, and storage medium

Publications (1)

Publication Number Publication Date
WO2020199337A1 true WO2020199337A1 (zh) 2020-10-08

Family

ID=67081906

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/089504 Ceased WO2020199337A1 (zh) 2019-03-29 2019-05-31 一种基因变异识别方法、装置和存储介质

Country Status (6)

Country Link
US (1) US20210151124A1 (zh)
JP (1) JP7064655B2 (zh)
CN (1) CN109979531B (zh)
SG (1) SG11202101410WA (zh)
TW (1) TWI740262B (zh)
WO (1) WO2020199337A1 (zh)

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111081313A (zh) * 2019-12-13 2020-04-28 北京市商汤科技开发有限公司 基因变异的识别方法及装置、电子设备和存储介质
CN111091873B (zh) * 2019-12-13 2023-07-18 北京市商汤科技开发有限公司 基因变异的识别方法及装置、电子设备和存储介质
CN111899790A (zh) * 2020-08-17 2020-11-06 天津诺禾医学检验所有限公司 测序数据的处理方法及装置
CN113539357B (zh) * 2021-06-10 2024-04-30 阿里巴巴达摩院(杭州)科技有限公司 基因检测方法、模型训练方法、装置、设备及系统
CN115458052B (zh) * 2022-08-16 2023-06-30 珠海横琴铂华医学检验有限公司 基于一代测序的基因突变分析方法、设备和存储介质

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2015062184A1 (en) * 2013-11-01 2015-05-07 Accurascience, Llc Method and apparatus for calling single-nucleotide variations and other variations
CN106611106A (zh) * 2016-12-06 2017-05-03 北京荣之联科技股份有限公司 基因变异检测方法及装置
CN108595912A (zh) * 2018-05-07 2018-09-28 深圳市瀚海基因生物科技有限公司 检测染色体非整倍性的方法、装置及系统

Family Cites Families (17)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002046462A2 (en) * 2000-12-07 2002-06-13 Isis Innovation Limited Functional genetic variants of matrix metalloproteinases (nmps)
WO2012149042A2 (en) * 2011-04-25 2012-11-01 Bio-Rad Laboratories, Inc. Methods and compositions for nucleic acid analysis
AU2013236175A1 (en) * 2012-03-22 2014-10-16 Kanagawa Prefectural Hospital Organization Method for identification and detection of mutant gene using intercalator
CN104603609B (zh) * 2013-07-31 2016-08-24 株式会社日立制作所 基因分析装置、基因分析系统以及基因分析方法
WO2015200869A1 (en) * 2014-06-26 2015-12-30 10X Genomics, Inc. Analysis of nucleic acid sequences
SG11201703638XA (en) * 2014-11-10 2017-06-29 Alnylam Pharmaceuticals Inc Hepatitis b virus (hbv) irna compositions and methods of use thereof
JP6675164B2 (ja) 2015-07-28 2020-04-01 株式会社理研ジェネシス 変異判定方法、変異判定プログラムおよび記録媒体
JP6679065B2 (ja) 2015-10-07 2020-04-15 国立研究開発法人国立がん研究センター 稀少突然変異の検出方法、検出装置及びコンピュータプログラム
US9988624B2 (en) * 2015-12-07 2018-06-05 Zymergen Inc. Microbial strain improvement by a HTP genomic engineering platform
US20200199612A1 (en) * 2016-03-18 2020-06-25 Monsanto Technology Llc Transgenic plants with enhanced traits
CN106529211A (zh) * 2016-11-04 2017-03-22 成都鑫云解码科技有限公司 变异位点的获取方法及装置
CN106503489A (zh) * 2016-11-04 2017-03-15 成都鑫云解码科技有限公司 心血管系统对应的基因的突变位点的获取方法及装置
CN106407747A (zh) * 2016-11-04 2017-02-15 成都鑫云解码科技有限公司 肿瘤对应的基因的突变位点的获取方法及装置
KR101936933B1 (ko) * 2016-11-29 2019-01-09 연세대학교 산학협력단 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스
JP7350659B2 (ja) * 2017-06-06 2023-09-26 ザイマージェン インコーポレイテッド Saccharopolyspora spinosaの改良のためのハイスループット(HTP)ゲノム操作プラットフォーム
CN109033751B (zh) * 2018-07-20 2021-07-27 东南大学 一种非编码区单核苷酸基因组变异的功能预测方法
CN109411016B (zh) * 2018-11-14 2020-12-01 钟祥博谦信息科技有限公司 基因变异位点检测方法、装置、设备及存储介质

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2015062184A1 (en) * 2013-11-01 2015-05-07 Accurascience, Llc Method and apparatus for calling single-nucleotide variations and other variations
CN106611106A (zh) * 2016-12-06 2017-05-03 北京荣之联科技股份有限公司 基因变异检测方法及装置
CN108595912A (zh) * 2018-05-07 2018-09-28 深圳市瀚海基因生物科技有限公司 检测染色体非整倍性的方法、装置及系统

Also Published As

Publication number Publication date
US20210151124A1 (en) 2021-05-20
TWI740262B (zh) 2021-09-21
CN109979531B (zh) 2021-08-31
SG11202101410WA (en) 2021-03-30
CN109979531A (zh) 2019-07-05
JP2022502766A (ja) 2022-01-11
TW202036584A (zh) 2020-10-01
JP7064655B2 (ja) 2022-05-10

Similar Documents

Publication Publication Date Title
CN109994155B (zh) 一种基因变异识别方法、装置和存储介质
TWI740262B (zh) 一種基因變異識別方法、裝置和儲存介質
Kaminski et al. pLM-BLAST: distant homology detection based on direct comparison of sequence representations from protein language models
Graham et al. BinSanity: unsupervised clustering of environmental microbial assemblies using coverage and affinity propagation
Schrider et al. Soft sweeps are the dominant mode of adaptation in the human genome
Peltzer et al. EAGER: efficient ancient genome reconstruction
Lee et al. DUDE-Seq: fast, flexible, and robust denoising for targeted amplicon sequencing
Nguyen et al. TIPP: taxonomic identification and phylogenetic profiling
Yaveroğlu et al. Proper evaluation of alignment-free network comparison methods
CN111292802A (zh) 用于检测突变的方法、电子设备和计算机存储介质
Sarmashghi et al. Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT
WO2023143016A1 (zh) 特征提取模型的生成方法、图像特征提取方法和装置
Qian et al. MetaCon: unsupervised clustering of metagenomic contigs with probabilistic k-mers statistics and coverage
CN109979530B (zh) 一种基因变异识别方法、装置和存储介质
Alganmi et al. Evaluation of an optimized germline exomes pipeline using BWA-MEM2 and Dragen-GATK tools
CN111933214B (zh) 用于检测rna水平体细胞基因变异的方法、计算设备
CN118655989A (zh) 提示词生成方法及文本处理方法
Qian et al. TEtrimmer: a tool to automate the manual curation of transposable elements
CN106326904A (zh) 获取特征排序模型的装置和方法以及特征排序方法
Huang et al. Reveel: large-scale population genotyping using low-coverage sequencing data
Qi et al. CREATE: a novel attention-based framework for efficient classification of transposable elements
Bonham-Carter et al. Cellular proliferation biases clonal lineage tracing and trajectory inference
Charitakis et al. Comparative analysis of packages and algorithms for the analysis of spatially resolved transcriptomics data
Nyström-Persson et al. Precise and scalable metagenomic profiling with sample-tailored minimizer libraries
CN110570908B (zh) 测序序列多态识别方法及装置、存储介质、电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19923039

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021517044

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 04/02/2022)

122 Ep: pct application non-entry in european phase

Ref document number: 19923039

Country of ref document: EP

Kind code of ref document: A1