WO2020199336A1 - 一种基因变异识别方法、装置和存储介质 - Google Patents
一种基因变异识别方法、装置和存储介质 Download PDFInfo
- Publication number
- WO2020199336A1 WO2020199336A1 PCT/CN2019/089499 CN2019089499W WO2020199336A1 WO 2020199336 A1 WO2020199336 A1 WO 2020199336A1 CN 2019089499 W CN2019089499 W CN 2019089499W WO 2020199336 A1 WO2020199336 A1 WO 2020199336A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- gene
- sequence
- site
- sequencing read
- attribute information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/10—Processes for the isolation, preparation or purification of DNA or RNA
- C12N15/1034—Isolating an individual clone by screening libraries
- C12N15/1093—General methods of preparing gene libraries, not provided for in other subgroups
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/10—Processes for the isolation, preparation or purification of DNA or RNA
- C12N15/1034—Isolating an individual clone by screening libraries
- C12N15/1082—Preparation or screening gene libraries by chromosomal integration of polynucleotide sequences, HR-, site-specific-recombination, transposons, viral vectors
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
-
- C—CHEMISTRY; METALLURGY
- C40—COMBINATORIAL TECHNOLOGY
- C40B—COMBINATORIAL CHEMISTRY; LIBRARIES, e.g. CHEMICAL LIBRARIES
- C40B40/00—Libraries per se, e.g. arrays, mixtures
- C40B40/04—Libraries containing only organic compounds
- C40B40/06—Libraries containing nucleotides or polynucleotides, or derivatives thereof
- C40B40/08—Libraries containing RNA or DNA which encodes proteins, e.g. gene libraries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/12—Computing arrangements based on biological models using genetic models
- G06N3/126—Evolutionary algorithms, e.g. genetic algorithms or genetic programming
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B5/00—ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks
Definitions
- the present disclosure relates to the field of computer technology, and in particular to a method, device and storage medium for identifying gene mutations.
- the sequence of human genes can be determined through gene sequencing technology, and the analysis of gene sequence can be used as the basis for further genetic research and modification.
- the gene sequencing technology greatly improves the efficiency of gene sequencing, reduces the cost of gene sequencing, and maintains the accuracy of gene sequencing. If the first-generation sequencing technology completes the sequencing of a human genome, it may take three years, while the second-generation sequencing technology can shorten the time to only one week.
- the present disclosure proposes a gene mutation identification scheme.
- a method for identifying gene mutations comprising:
- sequence feature is a feature related to the location of the site
- the gene variation at the gene variation candidate site is identified.
- the attribute information includes sequence attribute information; determining the sequence characteristics of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read includes:
- the gene location information of the gene mutation candidate site determine the preset site interval where the gene mutation candidate site is located
- sequence attribute information is information characterizing gene attributes related to the position of the site
- the sequence feature of the gene mutation candidate site is generated.
- the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
- the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
- the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
- the number of inserted genes of each gene type at each site of the at least one gene sequencing read is counted.
- the sequence attribute information includes at least one of the following information:
- the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
- the attribute information includes non-sequence attribute information; and determining the non-sequence characteristics of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read includes:
- non-sequence attribute information of the at least one gene sequencing read wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the site;
- the non-sequence feature of the gene mutation candidate site is determined.
- the non-sequence information includes at least one of the following information:
- Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
- the determining the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read includes:
- the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
- the determining the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read includes:
- the non-sequence features corresponding to the gene mutation candidate sites are determined.
- the identifying the gene mutation at the gene mutation candidate site based on the sequence feature and the non-sequence feature includes:
- the gene mutation of the gene mutation candidate site is identified.
- the identifying the gene mutation at the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site includes:
- the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
- the obtaining at least one gene sequencing read corresponding to the gene mutation candidate site includes:
- a gene mutation identification device comprising:
- the obtaining module is used to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
- the determining module is configured to determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site;
- the recognition module is used for recognizing the gene mutation of the gene mutation candidate site based on the sequence feature and the non-sequence feature.
- the attribute information includes sequence attribute information; the determining module includes:
- the first determining sub-module is used to determine the preset site interval where the gene mutation candidate site is located according to the gene location information of the gene mutation candidate site;
- the first acquiring submodule is used to acquire the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is related to the position of the site Information that characterizes genetic attributes;
- the first generation sub-module is used to generate the sequence characteristics of the gene mutation candidate sites according to the sequence attribute information of each site in the preset site interval.
- the first acquisition submodule is specifically configured to determine the gene type of the at least one gene sequencing read at the each site; and count each site corresponding to each site. The number of genes for each gene type.
- the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the missing gene at each locus; count the number of missing genes of each gene type at the at least one gene sequencing read at each locus.
- the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the inserted gene at each site; count the number of inserted genes of each gene type at each site of the at least one gene sequencing read.
- the sequence attribute information includes at least one of the following information:
- the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
- the attribute information includes non-sequence attribute information; the determining module includes:
- the second acquisition submodule is used to acquire the non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
- the second determining sub-module is used to determine the non-sequence characteristics of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
- the non-sequence information includes at least one of the following information:
- Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
- the second determining submodule is specifically configured to determine the comparison quality of each gene sequencing read according to the comparison quality of each site in each gene sequencing read;
- the comparison quality is used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing reads; according to the comparison quality of each gene sequencing read, the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
- the second determining submodule is specifically configured to determine the positive and negative chain information of the gene chain to which each gene sequencing read belongs, according to the positive and negative chain information of the gene chain to which the at least one gene sequencing read belongs. Negative strand ratio; according to the positive-negative strand ratio, the non-sequence features corresponding to the gene mutation candidate sites are determined.
- the identification module includes:
- the integration sub-module is specifically used to perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
- the recognition sub-module is used to identify the gene mutation of the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site.
- the recognition submodule is specifically configured to obtain the mutation value of the gene at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site; When the value is greater than or equal to the preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
- the acquisition module is specifically used for:
- a gene mutation identification device including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above method.
- a non-volatile computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor.
- the embodiments of the present disclosure provide for obtaining at least one gene sequencing read corresponding to a gene mutation candidate site, and the sequence feature and non-sequence feature of the gene mutation candidate site can be determined based on the attribute information of the at least one gene sequencing read.
- the sequence feature and non-sequence feature of the gene to identify the gene mutation at the candidate site of the gene mutation.
- sequence feature can be the feature related to the position of the locus
- non-sequence feature can be the feature not related to the position of the locus, so that in the process of gene variation identification, the sequence feature of the gene can be combined with the non-sequence feature Analyze the characteristics of gene mutation sites more comprehensively, screen out germline gene mutations and interference caused by noise and errors, better identify gene mutations, and enhance the accuracy of gene mutation identification.
- Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
- Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
- Fig. 3 shows a flowchart of the sequence characterization process of gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 4 shows a flowchart of the non-sequence feature process of gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure.
- Fig. 6 shows a block diagram of a neural network model according to an embodiment of the present disclosure.
- Fig. 7 shows a block diagram of a gene mutation identification device according to an embodiment of the present disclosure.
- Fig. 8 shows a block diagram of a device for gene mutation identification according to an exemplary embodiment of the present disclosure.
- the gene mutation identification scheme can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, so that the gene mutation of the gene mutation candidate site can be identified based on the at least one gene sequencing read.
- sequence features can be generated based on the sequence attribute information of at least one gene sequencing read, and non-sequence features can be generated based on the non-sequence attribute information of at least one gene sequencing read, and then the sequence feature and non-sequence feature can be paired
- the gene mutation at the gene mutation candidate site is identified, so that the sequence attribute information and non-sequence attribute information of at least one gene sequencing read can be integrated, and the sequence attribute information of the gene sequencing read can be used more comprehensively.
- the neural network model integrated by multimodal information can be used to extract the sequence features and non-sequence features of gene mutation candidate sites in the process of gene mutation identification, so that the sequence attribute information and non-sequence features of the gene sequence can be integrated.
- Sequence attribute information analyze gene data more comprehensively, screen out germline gene mutations and interference caused by noise and errors, and better identify gene mutations. The following examples will illustrate the process of gene mutation identification in detail.
- Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
- the gene mutation identification method can be executed by a gene mutation identification device or other processing equipment, where the gene mutation identification device can be User Equipment (UE), mobile equipment, user terminal, terminal, cellular phone, cordless phone, personal digital Processing (Personal Digital Assistant, PDA), handheld devices, computing devices, vehicle-mounted devices, wearable devices, etc., or the gene mutation recognition device may be a server.
- the gene mutation identification method may be implemented by a processor calling computer-readable instructions stored in a memory.
- the gene mutation identification method includes:
- Step 11 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
- the gene mutation recognition device can obtain the gene sequencing reads obtained by gene sequencing, and then obtain at least one gene sequencing read corresponding to the gene mutation candidate site from the gene sequencing reads obtained by the gene sequencing.
- the gene sequencing reads here can be understood as gene sequences marked with gene types after gene sequencing, and the length of each gene sequencing read can be the same or different. In the case of different lengths, the length of each gene sequencing read segment can be within a preset length range, thereby ensuring that the length of each gene sequencing read segment is relatively close.
- the gene type can be understood as the base type, and the gene type can include cytosine (C), guanine (G), adenine (A), and thymine (T), so that the gene sequencing read can be a gene sequence including AGCT.
- the candidate gene mutation site here can be a site where the gene sequence is abnormal.
- the locus of the gene sequence may indicate the position of the gene sequence.
- the gene mutation candidate site corresponds to at least one gene sequencing read, wherein the at least one gene sequencing read is abnormal at this site.
- each gene mutation candidate site may correspond to at least one gene sequencing read.
- the embodiment of the present disclosure uses a gene mutation candidate site for description.
- Step 12 Determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site.
- the attribute information of the at least one gene sequencing read corresponding to the gene mutation candidate site can be extracted, and based on the extracted attribute information Generate the sequence feature and non-sequence feature of the gene mutation candidate site.
- the attribute information may include sequence attribute information and non-sequence attribute information.
- the sequence attribute information may be information related to the position of the locus that characterizes the gene attribute of the gene sequencing read.
- Non-sequence attribute information can be information that is not restricted by the position of the site and can characterize gene attributes.
- the sequence attribute information of at least one gene sequencing read at the candidate gene mutation site can be extracted, and the sequence of at least one gene sequencing read at a location near the gene mutation candidate site can also be extracted.
- Property information when determining the sequence characteristics of gene mutation candidate sites, a neural network model with a convolutional layer and a pooling layer can be used to extract gene mutation candidate sites from at least one gene sequencing read corresponding to the gene mutation candidate site Sequence characteristics.
- the neural network model can include two branch structures. One branch can extract sequence features of gene sequencing reads, and the branch can include a convolutional layer and a pooling layer; the other branch can extract non-sequence features of gene sequencing reads.
- the neural network model can thus integrate multiple modal information (sequence attribute information and non-sequence attribute information) to identify gene mutations at gene mutation candidate sites.
- the aforementioned neural network model can be used to extract the non-sequence features of at least one gene sequencing read from another branch of the neural network model.
- the branch structure can include a fully connected layer, The fully connected layer can be used to extract non-sequential features that are not restricted by location.
- Step 13 based on the sequence feature and the non-sequence feature, identify the gene variation of the gene variation candidate site.
- the sequence feature and non-sequence feature can be fused to identify the gene mutation of the gene mutation candidate site, for example,
- the aforementioned neural network model is used to determine whether the gene at the gene mutation candidate site is mutated, or whether the gene at the gene mutation candidate site is abnormal in gene sequence due to noise or other reasons.
- the gene mutation of the gene mutation candidate site can be identified according to the sequence characteristics and non-sequence characteristics of the gene mutation candidate site, so that the gene sequencing data can be analyzed more comprehensively.
- it is first necessary to obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
- the example of the present disclosure also provides a process for obtaining at least one gene sequencing read corresponding to the gene mutation candidate site.
- Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
- obtaining at least one gene sequencing read corresponding to the gene mutation candidate site may include the following steps:
- Step 111 Obtain gene sequencing reads obtained by gene sequencing of somatic genes.
- At least one gene sequencing read segment can be obtained through gene sequencing of the somatic cell gene, and the gene sequencing read segment can be a sequence that annotates the gene type of the somatic gene.
- the gene sequencing read segment can be a sequence that annotates the gene type of the somatic gene.
- At least one gene sequencing read can be obtained through gene sequencing of somatic genes, and the gene sequencing reads obtained by gene sequencing can be preprocessed.
- the preprocessing methods here can include cross contamination screening, Sequencing quality screening, comparison quality screening, abnormal read length screening, etc. Through preprocessing, cross-contaminated gene sequencing reads can be screened out, and gene sequencing reads with low sequencing quality and comparison quality and abnormal read length can be screened out.
- Step 112 Compare the gene sequence of the gene sequencing read with the gene sequence of the reference genome to obtain the comparison result.
- the gene sequence of the obtained gene sequencing reads can be compared with the gene sequence of the reference genome at the same site, Get the comparison result.
- each gene sequencing read obtained by performing gene sequencing can be compared with the gene sequence of the reference genome at the same site to determine the sites where the gene sequence of the gene sequencing read is different from the gene sequence of the reference genome. It is also possible to compare at least one gene sequencing read with the same site with the gene sequence of the reference genome at the same site to determine the site where the gene sequence of the at least one gene sequencing read is different from the gene sequence of the reference genome.
- Step 113 According to the comparison result, it is determined that the gene of the somatic gene has an abnormal gene mutation candidate site.
- a site that differs from the gene sequence of the gene sequencing read from the reference genome can be determined according to the comparison result. If at least one gene sequencing read corresponding to the site is in at least one gene sequencing read, the mutated gene is sent at that site If the ratio of sequencing reads is greater than the preset ratio, it can be determined that the site is a candidate site for gene mutation; otherwise, it can be considered that the site is not a candidate site for gene mutation.
- the gene sequence of the gene sequencing read at this position is different from the gene sequence of the reference genome, which may be caused by sequencing errors. In this way, the abnormalities of gene sequence caused by gene sequencing errors can be reduced.
- Step 114 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
- each gene mutation candidate site corresponds to at least one gene sequencing read, and the gene sequence at the gene mutation candidate site may be different from the gene sequence of the reference genome at the same site. There may be at least one gene mutation candidate site.
- the sequence characteristics of the gene mutation candidate site can be determined according to the sequence attribute information of at least one gene sequencing read corresponding to the gene mutation candidate site, so that when the gene mutation of the gene mutation candidate site is identified , You can consider the sequence attributes of at least one gene sequencing read corresponding to the gene mutation candidate site.
- the process of determining the sequence characteristics of gene mutation candidate sites will be described in detail below through an example.
- Fig. 3 shows a flowchart of the sequence characterization process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 3, the above step 12 may include the following steps:
- Step 121a according to the gene location information of the gene mutation candidate site, determine the preset site interval where the gene mutation candidate site is located;
- Step 122a Obtain the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is information characterizing gene attributes related to the position of the site ;
- Step 123a According to the sequence attribute information of each site in the preset site interval, the sequence feature of the gene mutation candidate site is generated.
- the preset site interval in which the gene mutation candidate site is located can be determined according to the gene location information of the gene mutation candidate site. For example, 150 before and after the gene mutation candidate site can be determined. The interval of 1 base pair is used as the preset site interval where the candidate site of gene mutation is located.
- Sequence features can be represented by sequence feature vectors.
- At least one sequence feature vector corresponding to at least one site in the preset site interval where the gene mutation candidate site is located can form a sequence feature matrix of the gene mutation candidate site. For example, if the preset site interval where the gene mutation candidate site is located includes 3 sites b1, b2, b3, the sequence feature vectors corresponding to the 3 sites are a1, a2, and a3, respectively.
- the sequence feature matrix is [a1a2a3], where the sequence features of a1, a2, and a3 correspond to the sequence attribute information of b1, b2, and b3, respectively.
- the sequence attribute information may include, but is not limited to: the gene type of the reference genome; the number of genes of each gene type; the number of missing genes of each gene type; the number of inserted genes of each gene type.
- the gene type of the reference genome may be the gene type of the reference genome at the gene mutation candidate site.
- the number of genes of each gene type can be the number of genes of each gene type at the candidate site of at least one gene sequencing read of the gene mutation.
- the candidate site of the gene mutation corresponds to 5 gene sequencing reads, and each gene is sequenced.
- the gene types of the reads at the gene mutation candidate sites are: A, C, C, G, G, the number of genes for each gene type are: A is 1; C is 2; G is 2 .
- the number of missing genes of each gene type can be the number of missing genes of each gene type at the candidate site of the gene mutation in at least one gene sequencing read, for example, the number of genes missing in each gene sequencing read at the candidate site of the gene mutation
- the types are: A, C, C, G, G, and the number of missing genes for each gene type is: A is 1; C is 2; G is 2.
- the number of inserted genes of each gene type can be the number of inserted genes of each gene type at the candidate site of the gene mutation of at least one gene sequencing read, for example, the gene inserted at the candidate site of the gene mutation of each gene sequencing read.
- the types are: A, C, C, G, G, and the number of inserted genes for each gene type is: A is 1; C is 2; G is 2.
- the sequence attribute information of at least one gene sequencing read at each site in the preset site interval it may be determined for each site in the preset site interval
- the gene type of at least one gene sequencing read at the locus, and the number of genes of each gene type corresponding to the locus is counted, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined. Click the number of genes for each gene type.
- the gene sequence of each gene sequencing read may be compared with the gene of the reference genome. Based on the comparison result of sequence comparison, for each site in the preset site interval, determine the gene type of the missing gene of each gene sequencing read at that site, and count at least one gene sequencing read at The number of missing genes of each gene type at this locus, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined, and the number of missing genes of each gene type at that locus can be determined.
- the gene sequence of each gene sequencing read may be compared with the gene of the reference genome. Based on the comparison result of sequence comparison, for each site in the preset site interval, determine the gene type of the missing gene of each gene sequencing read at that site, and count at least one gene sequencing read at The number of inserted genes of each gene type at the locus, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined, and the number of inserted genes of each gene type at the locus can be determined.
- the sequence attribute information includes the gene type of the reference genome, the number of genes of each gene type, the number of missing genes of each gene type, the number of inserted genes of each gene type, and the sequence in determining the candidate site of gene mutation
- the above four information of at least one gene sequencing read corresponding to the gene mutation candidate site can be extracted at that site, for example, For the 5 gene sequencing reads corresponding to the candidate gene mutation site, for a certain site in the preset site interval, the gene type of the reference genome at that site can be determined respectively, and the 5 gene sequencing reads can be determined at that location.
- the sequence characteristics of the gene mutation candidate sites may include the sequence characteristics of each site in the preset site interval.
- Fig. 4 shows a flowchart of the non-sequence feature process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 4, the above step 12 may include the following steps:
- Step 121b acquiring non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
- Step 122b Generate the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
- the non-sequence information may include at least one of the following information: comparison quality; positive and negative chain preference; gene sequencing read length; edge preference.
- the non-sequence attribute information of at least one gene attribute sequence read can be obtained, and then the non-sequence feature of the gene mutation candidate site can be generated from the obtained non-sequence attribute information.
- each gene sequencing read may be The comparison quality of the locus is to determine the comparison quality of each gene sequencing read, and then according to the comparison quality of each gene sequencing read, determine the non-sequence feature corresponding to the gene mutation candidate locus.
- the comparison quality can be used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing read. If the comparison quality of a certain gene sequence is lower than the preset value, it can be considered that the gene sequence is a gene obtained by gene sequencing.
- the quality of comparison can be used as a reference factor for judging whether the gene at the gene mutation candidate site has mutation.
- the comparison quality of each gene sequencing read can be determined according to the comparison quality of each gene sequence. Taking a gene sequencing read as an example, you can The average or median value of the comparison quality of the gene sequences included in the gene sequencing reads is used as the comparison quality of the gene sequencing reads.
- At least one gene sequence can also be randomly selected in the gene sequencing reads, and the selected at least one gene The average or intermediate value of the sequence comparison quality is used as the comparison quality of the sequence reads of the gene.
- the comparison quality corresponding to the gene mutation candidate site is obtained. For example, the average or mean value of the comparison quality of at least one gene sequencing read corresponding to the gene mutation candidate site is calculated to obtain the The comparison quality corresponding to the gene mutation candidate site, so that the non-sequence features corresponding to the gene mutation candidate site can be determined according to the comparison quality corresponding to the gene mutation candidate site.
- the positive and negative chain of the gene chain to which each gene sequencing read belongs can be determined.
- Information determine the positive-negative chain ratio of the gene chain to which at least one gene sequencing read belongs, and then determine the non-sequence feature corresponding to the gene mutation candidate site according to the determined positive-negative chain ratio.
- the positive-negative strand preference can be the ratio of the positive strand and the negative strand in the gene strand to which the gene sequencing read belongs.
- the gene strand can include the positive strand and the negative strand, where the positive strand can be the same base sequence as the ribonucleic acid (RNA)
- RNA ribonucleic acid
- a single strand of deoxyribonucleic acid (DNA) the negative strand can be a single strand of deoxyribonucleic acid (DNA) complementary to the base sequence of ribonucleic acid (RNA).
- gene mutation candidate sites correspond to 5 gene sequencing reads, among which, 3 gene sequencing reads correspond to the positive strand of the gene chain, and 2 gene sequencing reads correspond to the negative strand of the gene chain, and the positive and negative strands are preferred It can be 3:2.
- the length of the gene sequencing read of each gene sequencing read can be used, Determine the non-sequence characteristics of candidate sites for gene mutations.
- the length of a gene sequencing read can be the length of the base sequence of each gene sequencing read. For example, if a gene sequencing read includes 4 base sequences, the length of the gene sequencing read is 4, which can be determined by The length of each gene sequencing read determines the non-sequence feature of the gene mutation candidate site, and the non-sequence feature of the gene mutation candidate site can also be determined by the median or average of the length of at least one gene sequencing read.
- the gene mutation when determining the non-sequence characteristics of the candidate gene mutation site based on the non-sequence attribute information of at least one gene sequencing read, the gene mutation can be determined according to the edge preference of each gene sequencing read Non-sequence features of candidate sites.
- the marginal preference can be the ratio of the marginal position to the middle position of a certain site in the gene sequencing reads.
- gene sequencing reads can be divided into 3 evenly, where the two segments at both ends of the gene sequencing read can be used as edge positions, and the middle segment of the gene sequencing read can be used as the middle position, and the gene mutation candidate site corresponds to
- the edge preference of the gene mutation candidate site can be 3: 2.
- the non-sequence characteristics of the gene mutation candidate site can be determined from the marginal preference of the gene mutation candidate site in each gene sequencing read, and the median value of the marginal preference corresponding to at least one gene sequencing read or The average value is used to determine the non-sequence characteristics of candidate sites of gene mutation.
- the non-sequence feature of the gene mutation candidate site can be generated based on the non-sequence attribute information of at least one gene sequencing read at the gene mutation candidate site, so that the non-sequence of the gene mutation candidate site can be considered when gene mutation identification Characteristic features make gene mutation identification more accurate.
- the non-sequence feature of at least one gene sequencing read may be generated from a combination of any at least one information in the non-sequence attribute information.
- the following uses an example to illustrate the process of identifying gene mutations at gene mutation candidate sites.
- Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 5, the above step 13 may include the following steps:
- Step 131 Perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
- Step 132 based on the integration characteristics of the gene mutation candidate site, identify the gene mutation of the gene mutation candidate site.
- the neural network model can be used to perform feature integration on the sequence feature and the non-sequence feature, and the sequence feature matrix formed by the sequence feature can be combined with the non-sequence feature.
- the non-sequence feature matrix formed by the sequence feature is synthesized into a feature matrix, and the integrated feature matrix formed by the integrated feature is obtained, and then the neural network model is used to identify the gene mutation of the mutation candidate site according to the integrated feature matrix.
- the neural network model can be used to integrate sequence attribute information and non-sequence attribute information corresponding to gene mutation candidate sites, so that gene sequencing data can be analyzed more comprehensively, and gene mutation identification can be more accurate.
- you can select gene sequencing reads with Single Nucleotide Polymorphism (SNP), and gene sequencing reads with Insertion/Deletion (InDel) as training samples to train The gene variation recognition model obtained afterwards can effectively identify SNP and InDel gene variation.
- SNP Single Nucleotide Polymorphism
- InDel Insertion/Deletion
- identifying the genetic variation of the gene mutation candidate site according to the integration feature of the gene mutation candidate site may include: according to the integration feature of the gene mutation candidate site, Obtain the mutation value of the gene at the gene mutation candidate site; if the mutation value is greater than or equal to a preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
- the mutation value of the gene mutation may be indicative of the possibility of mutation of the candidate site of the gene mutation. For example, the greater the mutation value, the greater the possibility of mutation of the candidate site of the gene mutation.
- the above-mentioned neural network can be used to process the two-dimensional feature to obtain the mutation value, and to determine whether the gene at the gene mutation candidate site has mutation according to the mutation value.
- the variation value can be between 0 and 1.
- the preset threshold can be set according to the application scenario, for example, 0.3, 0.5. If the mutation value is greater than the preset threshold, it can be considered that the gene at the candidate site of gene mutation has been mutated, otherwise, it can be the gene at the candidate site of gene mutation. No mutation has occurred.
- a neural network model can be used to identify gene mutations at gene mutation candidate sites, and the neural network model can extract sequence features and non-sequence features of gene mutation candidate sites.
- the embodiment of the present disclosure also provides a structure of a neural network model.
- Fig. 6 shows a block diagram of a neural network model according to an embodiment of the present disclosure.
- the neural network model can include two branch structures, the first branch and the second branch.
- the first branch may be used to extract the sequence features of at least one gene sequencing read corresponding to the gene mutation candidate site, and the first branch may include a convolutional layer and a pooling layer.
- the second branch may be used to extract the non-sequence features of at least one gene sequencing read corresponding to the gene mutation candidate site, and the second branch may include a fully connected layer.
- the sequence features and non-sequence features can be integrated, for example, the sequence feature matrix of sequence features and the non-sequence feature matrix of non-sequence features can be spliced together.
- the integrated feature matrix of the integrated feature is obtained, and then the mutation value of the gene mutation candidate site can be obtained through the fully connected layer.
- the sequence attribute information and non-sequence attribute information of at least one gene sequencing read corresponding to the gene mutation candidate site are extracted, and the gene mutation is identified by the integration feature integrating the sequence attribute information and the non-sequence attribute information, thereby Comprehensively consider the sequence attribute information and non-sequence attribute information corresponding to gene mutation candidate sites, analyze gene sequencing information more comprehensively, better identify gene mutations at gene candidate sites, and screen out germline gene mutations and noise and The interference caused by errors improves the accuracy of gene mutation identification.
- the writing order of the steps does not mean a strict execution order but constitutes any limitation on the implementation process.
- the specific execution order of each step should be based on its function and possibility.
- the inner logic is determined.
- Fig. 7 shows a block diagram of a gene mutation recognition device according to an embodiment of the present disclosure. As shown in Fig. 7, the gene mutation recognition device includes:
- the obtaining module 71 is used to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
- the determining module 72 is configured to determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site ;
- the identification module 73 is configured to identify the gene mutation at the gene mutation candidate site based on the sequence feature and the non-sequence feature.
- the attribute information includes sequence attribute information; the determining module 72 includes:
- the first determining sub-module is used to determine the preset site interval where the gene mutation candidate site is located according to the gene location information of the gene mutation candidate site;
- the first acquiring submodule is used to acquire the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is related to the position of the site Information that characterizes genetic attributes;
- the first generation sub-module is used to generate the sequence characteristics of the gene mutation candidate sites according to the sequence attribute information of each site in the preset site interval.
- the first acquisition submodule is specifically configured to determine the gene type of the at least one gene sequencing read at the each site; and count each site corresponding to each site. The number of genes for each gene type.
- the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the missing gene at each locus; count the number of missing genes of each gene type at the at least one gene sequencing read at each locus.
- the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the inserted gene at each site; count the number of inserted genes of each gene type at each site of the at least one gene sequencing read.
- the sequence attribute information includes at least one of the following information:
- the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
- the attribute information includes non-sequence attribute information; the determining module includes:
- the second acquisition submodule is used to acquire the non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
- the second determining sub-module is used to determine the non-sequence characteristics of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
- the non-sequence information includes at least one of the following information:
- Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
- the second determining submodule is specifically configured to determine the comparison quality of each gene sequencing read according to the comparison quality of each site in each gene sequencing read;
- the comparison quality is used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing reads; according to the comparison quality of each gene sequencing read, the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
- the second determining submodule is specifically configured to determine the positive and negative chain information of the gene chain to which each gene sequencing read belongs, according to the positive and negative chain information of the gene chain to which the at least one gene sequencing read belongs. Negative strand ratio; according to the positive-negative strand ratio, the non-sequence features corresponding to the gene mutation candidate sites are determined.
- the identification module 73 includes:
- the integration sub-module is specifically used to perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
- the recognition sub-module is used to identify the gene mutation of the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site.
- the recognition submodule is specifically configured to obtain the mutation value of the gene at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site; When the value is greater than or equal to the preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
- the obtaining module 71 is specifically configured to:
- the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
- the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
- Fig. 8 is a block diagram showing a device 1900 for gene mutation identification according to an exemplary embodiment.
- the device 1900 may be provided as a server.
- the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions that can be executed by the processing component 1922, such as application programs.
- the application program stored in the memory 1932 may include one or more modules each corresponding to a set of instructions.
- the processing component 1922 is configured to execute instructions to perform the above-described methods.
- the device 1900 may also include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to the network, and an input output (I/O) interface 1958.
- the device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
- a non-volatile computer-readable storage medium such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to complete the foregoing method.
- the present disclosure may be a system, method, and/or computer program product.
- the computer program product may include a computer-readable storage medium loaded with computer-readable program instructions for enabling a processor to implement various aspects of the present disclosure.
- the computer-readable storage medium may be a tangible device that can hold and store instructions used by the instruction execution device.
- the computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
- Computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) Or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a printer with instructions stored thereon
- RAM random access memory
- ROM read-only memory
- EPROM erasable programmable read-only memory
- flash memory flash memory
- SRAM static random access memory
- CD-ROM compact disk read-only memory
- DVD digital versatile disk
- memory stick floppy disk
- mechanical encoding device such as a printer with instructions stored thereon
- the computer-readable storage medium used here is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (for example, light pulses through fiber optic cables), or through wires Transmission of electrical signals.
- the computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing/processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and/or a wireless network.
- the network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and/or edge servers.
- the network adapter card or network interface in each computing/processing device receives computer-readable program instructions from the network, and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing/processing device .
- the computer program instructions used to perform the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or in one or more programming languages.
- Source code or object code written in any combination, the programming language includes object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" language or similar programming languages.
- Computer-readable program instructions can be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly executed on a remote computer, or entirely on the remote computer or server carried out.
- the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to access the Internet connection).
- LAN local area network
- WAN wide area network
- an electronic circuit such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can be customized by using the status information of the computer-readable program instructions.
- the computer-readable program instructions are executed to realize various aspects of the present disclosure.
- These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device , A device that implements the functions/actions specified in one or more blocks in the flowchart and/or block diagram is produced. It is also possible to store these computer-readable program instructions in a computer-readable storage medium. These instructions make computers, programmable data processing apparatuses, and/or other devices work in a specific manner, so that the computer-readable medium storing instructions includes An article of manufacture, which includes instructions for implementing various aspects of the functions/actions specified in one or more blocks in the flowchart and/or block diagram.
- each block in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and the module, program segment, or part of an instruction contains one or more functions for implementing the specified logical function.
- Executable instructions may also occur in a different order from the order marked in the drawings. For example, two consecutive blocks can actually be executed in parallel, or they can sometimes be executed in the reverse order, depending on the functions involved.
- each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or actions Or it can be realized by a combination of dedicated hardware and computer instructions.
Landscapes
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Chemical & Material Sciences (AREA)
- Biophysics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Theoretical Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Genetics & Genomics (AREA)
- Bioinformatics & Computational Biology (AREA)
- General Engineering & Computer Science (AREA)
- Organic Chemistry (AREA)
- Evolutionary Biology (AREA)
- Medical Informatics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Data Mining & Analysis (AREA)
- Biomedical Technology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Analytical Chemistry (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Evolutionary Computation (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Computational Linguistics (AREA)
- Computing Systems (AREA)
- Mathematical Physics (AREA)
- General Physics & Mathematics (AREA)
- Biochemistry (AREA)
- Microbiology (AREA)
- Physiology (AREA)
- Databases & Information Systems (AREA)
- Immunology (AREA)
- Public Health (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Epidemiology (AREA)
- Bioethics (AREA)
Abstract
Description
Claims (28)
- 一种基因变异识别方法,其特征在于,所述方法包括:获取基因变异候选位点对应的至少一个基因测序读段;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求1所述的方法,其特征在于,所述属性信息包括序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征,包括:根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
- 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:确定所述至少一个基因测序读段在所述每个位点的基因类型;统计所述每个位点对应的每种基因类型的基因数量。
- 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
- 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
- 根据权利要求1至5任意一项所述的方法,其特征在于,所述序列属性信息包括以下至少一种信息:参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
- 根据权利要求1至6任意一项所述的方法,其特征在于,所述属性信息包括非序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的非序列特征,包括:获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
- 根据权利要求7所述的方法,其特征在于,所述非序列信息包括以下至少一种信息:对比质量;正负链偏好;基因测序读段长度;边缘偏好。
- 根据权利要求8所述的方法,其特征在于,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
- 根据权利要求8所述的方法,其特征在于,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
- 根据权利要求1-10任意一项所述的方法,其特征在于,所述基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别,包括:将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求11所述的方法,其特征在于,所述基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别,包括:根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
- 根据权利要求1至12任意一项所述的方法,其特征在于,所述获取基因变异候选位点对应的至少一个基因测序读段,包括:获取由体细胞基因进行基因测序得到的基因测序读段;将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;获取所述基因变异候选位点对应的至少一个基因测序读段。
- 一种基因变异识别装置,其特征在于,所述装置包括:获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;确定模块,用于根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;识别模块,用于基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求14所述的装置,其特征在于,所述属性信息包括序列属性信息;所述确定模块,包括:第一确定子模块,用于根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;第一生成子模块,用于根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
- 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于确定所述至少一个基因测序读段在所述每个位点的基因类型;统计所述每个位点对应的每种基因类型的基因数量。
- 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
- 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
- 根据权利要求14至18任意一项所述的装置,其特征在于,所述序列属性信息包括以下至少一种信息:参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
- 根据权利要求14至19任意一项所述的装置,其特征在于,所述属性信息包括非序列属性信息;所述确定模块,包括:第二获取子模块,用于获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;第二确定子模块,用于根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
- 根据权利要求20任意一项所述的装置,其特征在于,所述非序列信息包括以下至少一种信息:对比质量;正负链偏好;基因测序读段长度;边缘偏好。
- 根据权利要求21所述的装置,其特征在于,所述第二确定子模块,具体用于根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
- 根据权利要求21所述的装置,其特征在于,所述第二确定子模块,具体用于根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
- 根据权利要求14-23任意一项所述的装置,其特征在于,所述识别模块,包括:整合子模块,具体用于将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;识别子模块,用于基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
- 根据权利要求24所述的装置,其特征在于,所述识别子模块,具体用于根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
- 根据权利要求14至25任意一项所述的装置,其特征在于,所述获取模块,具体用于,获取由体细胞基因进行基因测序得到的基因测序读段;将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;获取所述基因变异候选位点对应的至少一个基因测序读段。
- 一种基因变异识别装置,其特征在于,包括:处理器;用于存储处理器可执行指令的存储器;其中,其中,所述处理器通过调用所述可执行指令实现如权利要求1至13中任意一项所述的方法。
- 一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其特征在于,所述计算机程序指令被处理器执行时实现权利要求1至13中任意一项所述的方法。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2021514554A JP7064654B2 (ja) | 2019-03-29 | 2019-05-31 | 遺伝子変異認識方法、装置および記憶媒体 |
| SG11202011523VA SG11202011523VA (en) | 2019-03-29 | 2019-05-31 | Gene mutation identification method and apparatus, and storage medium |
| KR1020217020204A KR20210116454A (ko) | 2019-03-29 | 2019-05-31 | 유전자 변이 인식 방법 및 장치 및 기억 매체 |
| US17/102,136 US20210082539A1 (en) | 2019-03-29 | 2020-11-23 | Gene mutation identification method and apparatus, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910251891.0A CN109994155B (zh) | 2019-03-29 | 2019-03-29 | 一种基因变异识别方法、装置和存储介质 |
| CN201910251891.0 | 2019-03-29 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/102,136 Continuation US20210082539A1 (en) | 2019-03-29 | 2020-11-23 | Gene mutation identification method and apparatus, and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020199336A1 true WO2020199336A1 (zh) | 2020-10-08 |
Family
ID=67131990
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/089499 Ceased WO2020199336A1 (zh) | 2019-03-29 | 2019-05-31 | 一种基因变异识别方法、装置和存储介质 |
Country Status (7)
| Country | Link |
|---|---|
| US (1) | US20210082539A1 (zh) |
| JP (1) | JP7064654B2 (zh) |
| KR (1) | KR20210116454A (zh) |
| CN (1) | CN109994155B (zh) |
| SG (1) | SG11202011523VA (zh) |
| TW (1) | TWI748263B (zh) |
| WO (1) | WO2020199336A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113628683A (zh) * | 2021-08-24 | 2021-11-09 | 慧算医疗科技(上海)有限公司 | 一种高通量测序突变检测方法、设备、装置及可读存储介质 |
| US20240221954A1 (en) * | 2021-10-28 | 2024-07-04 | Chengdu Boe Optoelectronics Technology Co., Ltd. | Disease prediction methods and devices, electronic devices, and computer readable storage media |
Families Citing this family (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111081318B (zh) * | 2019-12-06 | 2023-06-06 | 人和未来生物科技(长沙)有限公司 | 一种融合基因检测方法、系统和介质 |
| CN111081314A (zh) * | 2019-12-13 | 2020-04-28 | 北京市商汤科技开发有限公司 | 基因变异的识别方法及装置、电子设备和存储介质 |
| CN111091873B (zh) * | 2019-12-13 | 2023-07-18 | 北京市商汤科技开发有限公司 | 基因变异的识别方法及装置、电子设备和存储介质 |
| CN111081313A (zh) * | 2019-12-13 | 2020-04-28 | 北京市商汤科技开发有限公司 | 基因变异的识别方法及装置、电子设备和存储介质 |
| CN111091867B (zh) * | 2019-12-18 | 2021-11-09 | 中国科学院大学 | 基因变异位点筛选方法及系统 |
| CN111304308B (zh) * | 2020-03-02 | 2025-09-16 | 北京泛生子基因科技有限公司 | 一种审核高通量测序基因变异检测结果的方法 |
| CN113517022B (zh) * | 2021-06-10 | 2024-06-25 | 阿里巴巴达摩院(杭州)科技有限公司 | 基因检测方法、特征提取方法、装置、设备及系统 |
| CN113539357B (zh) * | 2021-06-10 | 2024-04-30 | 阿里巴巴达摩院(杭州)科技有限公司 | 基因检测方法、模型训练方法、装置、设备及系统 |
| CN113299344A (zh) * | 2021-06-23 | 2021-08-24 | 深圳华大医学检验实验室 | 基因测序分析方法、装置、存储介质和计算机设备 |
| CN115458052B (zh) * | 2022-08-16 | 2023-06-30 | 珠海横琴铂华医学检验有限公司 | 基于一代测序的基因突变分析方法、设备和存储介质 |
| CN115620802B (zh) * | 2022-09-02 | 2023-12-05 | 蔓之研(上海)生物科技有限公司 | 一种基因数据的处理方法及系统 |
| CN120319304B (zh) * | 2025-04-15 | 2026-02-03 | 中普康睿河北生物科技有限公司 | 用于基因多态性快速分型诊断的智能识别方法及系统 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014149134A2 (en) * | 2013-03-15 | 2014-09-25 | Guardant Health Inc. | Systems and methods to detect rare mutations and copy number variation |
| CN105989246A (zh) * | 2015-01-28 | 2016-10-05 | 深圳华大基因研究院 | 一种基于基因组组装的变异检测方法和装置 |
| WO2016179049A1 (en) * | 2015-05-01 | 2016-11-10 | Guardant Health, Inc | Diagnostic methods |
| CN107944228A (zh) * | 2017-12-08 | 2018-04-20 | 广州漫瑞生物信息技术有限公司 | 一种基因测序变异位点的可视化方法 |
Family Cites Families (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| ES2704126T3 (es) * | 2011-08-23 | 2019-03-14 | Found Medicine Inc | Moléculas de fusión KIF5B-RET y usos de las mismas |
| EP2959011A1 (en) * | 2013-02-19 | 2015-12-30 | Cergentis B.V. | Sequencing strategies for genomic regions of interest |
| KR20160010277A (ko) * | 2014-07-18 | 2016-01-27 | 에스케이텔레콤 주식회사 | 산모의 무세포 dna의 차세대 서열분석을 통한 태아의 단일유전자 유전변이의 예측방법 |
| CN104293940B (zh) * | 2014-09-30 | 2017-07-28 | 天津华大基因科技有限公司 | 构建测序文库的方法及其应用 |
| CN104462869B (zh) * | 2014-11-28 | 2017-12-26 | 天津诺禾致源生物信息科技有限公司 | 检测体细胞单核苷酸突变的方法和装置 |
| JP6675164B2 (ja) | 2015-07-28 | 2020-04-01 | 株式会社理研ジェネシス | 変異判定方法、変異判定プログラムおよび記録媒体 |
| JP6679065B2 (ja) | 2015-10-07 | 2020-04-15 | 国立研究開発法人国立がん研究センター | 稀少突然変異の検出方法、検出装置及びコンピュータプログラム |
| CN105574361B (zh) * | 2015-11-05 | 2018-11-02 | 上海序康医疗科技有限公司 | 一种检测基因组拷贝数变异的方法 |
| CN106529211A (zh) * | 2016-11-04 | 2017-03-22 | 成都鑫云解码科技有限公司 | 变异位点的获取方法及装置 |
| KR101936933B1 (ko) * | 2016-11-29 | 2019-01-09 | 연세대학교 산학협력단 | 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스 |
| CN106611106B (zh) | 2016-12-06 | 2019-05-03 | 北京荣之联科技股份有限公司 | 基因变异检测方法及装置 |
| CN106683081B (zh) * | 2016-12-17 | 2020-10-30 | 复旦大学 | 基于影像组学的脑胶质瘤分子标记物无损预测方法和预测系统 |
| KR102035615B1 (ko) | 2017-08-07 | 2019-10-23 | 연세대학교 산학협력단 | 유전자 패널에 기초한 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스 |
| CN108021788B (zh) * | 2017-12-06 | 2022-08-05 | 北京新合睿恩生物医疗科技有限公司 | 基于细胞游离dna的深度测序数据提取生物标记物的方法和装置 |
| EP3587586A1 (en) * | 2018-06-22 | 2020-01-01 | Julius-Maximilians-Universität Würzburg | Method for statistically determining a quantification of old and new rna |
| CN109326316B (zh) * | 2018-09-18 | 2020-10-09 | 哈尔滨工业大学(深圳) | 一种癌症相关SNP、基因、miRNA和蛋白质相互作用的多层网络模型构建方法和应用 |
-
2019
- 2019-03-29 CN CN201910251891.0A patent/CN109994155B/zh active Active
- 2019-05-31 KR KR1020217020204A patent/KR20210116454A/ko not_active Withdrawn
- 2019-05-31 SG SG11202011523VA patent/SG11202011523VA/en unknown
- 2019-05-31 WO PCT/CN2019/089499 patent/WO2020199336A1/zh not_active Ceased
- 2019-05-31 JP JP2021514554A patent/JP7064654B2/ja not_active Expired - Fee Related
- 2019-10-16 TW TW108137265A patent/TWI748263B/zh not_active IP Right Cessation
-
2020
- 2020-11-23 US US17/102,136 patent/US20210082539A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2014149134A2 (en) * | 2013-03-15 | 2014-09-25 | Guardant Health Inc. | Systems and methods to detect rare mutations and copy number variation |
| CN105989246A (zh) * | 2015-01-28 | 2016-10-05 | 深圳华大基因研究院 | 一种基于基因组组装的变异检测方法和装置 |
| WO2016179049A1 (en) * | 2015-05-01 | 2016-11-10 | Guardant Health, Inc | Diagnostic methods |
| CN107944228A (zh) * | 2017-12-08 | 2018-04-20 | 广州漫瑞生物信息技术有限公司 | 一种基因测序变异位点的可视化方法 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113628683A (zh) * | 2021-08-24 | 2021-11-09 | 慧算医疗科技(上海)有限公司 | 一种高通量测序突变检测方法、设备、装置及可读存储介质 |
| CN113628683B (zh) * | 2021-08-24 | 2024-04-09 | 慧算医疗科技(上海)有限公司 | 一种高通量测序突变检测方法、设备、装置及可读存储介质 |
| US20240221954A1 (en) * | 2021-10-28 | 2024-07-04 | Chengdu Boe Optoelectronics Technology Co., Ltd. | Disease prediction methods and devices, electronic devices, and computer readable storage media |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109994155B (zh) | 2021-08-20 |
| US20210082539A1 (en) | 2021-03-18 |
| CN109994155A (zh) | 2019-07-09 |
| KR20210116454A (ko) | 2021-09-27 |
| TW202036582A (zh) | 2020-10-01 |
| SG11202011523VA (en) | 2020-12-30 |
| JP7064654B2 (ja) | 2022-05-10 |
| JP2022500773A (ja) | 2022-01-04 |
| TWI748263B (zh) | 2021-12-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| TWI748263B (zh) | 一種基因變異辨識方法、裝置和儲存介質 | |
| Ono et al. | PBSIM2: a simulator for long-read sequencers with a novel generative model of quality scores | |
| Sessegolo et al. | Transcriptome profiling of mouse samples using nanopore sequencing of cDNA and RNA molecules | |
| Hu et al. | Bayesian detection of convergent rate changes of conserved noncoding elements on phylogenetic trees | |
| Ayling et al. | New approaches for metagenome assembly with short reads | |
| Rakocevic et al. | Fast and accurate genomic analyses using genome graphs | |
| Girgis | Red: an intelligent, rapid, accurate tool for detecting repeats de-novo on the genomic scale | |
| Modolo et al. | UrQt: an efficient software for the Unsupervised Quality trimming of NGS data | |
| TWI740262B (zh) | 一種基因變異識別方法、裝置和儲存介質 | |
| CN111402951B (zh) | 拷贝数变异预测方法、装置、计算机设备和存储介质 | |
| CN111292802A (zh) | 用于检测突变的方法、电子设备和计算机存储介质 | |
| Sarmashghi et al. | Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT | |
| Wang et al. | Tool evaluation for the detection of variably sized indels from next generation whole genome and targeted sequencing data | |
| CN109979530B (zh) | 一种基因变异识别方法、装置和存储介质 | |
| Ferrario et al. | Transferring entropy to the realm of GxG interactions | |
| CN111933214A (zh) | 用于检测rna水平体细胞基因变异的方法、计算设备 | |
| Huang et al. | Reveel: large-scale population genotyping using low-coverage sequencing data | |
| Bonham-Carter et al. | Cellular proliferation biases clonal lineage tracing and trajectory inference | |
| Fenton et al. | Distinguishing multiple-merger from Kingman coalescence using two-site frequency spectra | |
| US20160026756A1 (en) | Method and apparatus for separating quality levels in sequence data and sequencing longer reads | |
| CN110570908B (zh) | 测序序列多态识别方法及装置、存储介质、电子设备 | |
| HK40007439B (zh) | 一种基因变异识别方法、装置和存储介质 | |
| HK40007439A (zh) | 一种基因变异识别方法、装置和存储介质 | |
| Zheng et al. | CIGenotyper: A machine learning approach for genotyping complex indel calls | |
| HK40006878B (zh) | 一种基因变异识别方法、装置和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19923296 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021514554 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 04.02.2022) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19923296 Country of ref document: EP Kind code of ref document: A1 |