WO2020199336A1 - 一种基因变异识别方法、装置和存储介质 - Google Patents

一种基因变异识别方法、装置和存储介质 Download PDF

Info

Publication number
WO2020199336A1
WO2020199336A1 PCT/CN2019/089499 CN2019089499W WO2020199336A1 WO 2020199336 A1 WO2020199336 A1 WO 2020199336A1 CN 2019089499 W CN2019089499 W CN 2019089499W WO 2020199336 A1 WO2020199336 A1 WO 2020199336A1
Authority
WO
WIPO (PCT)
Prior art keywords
gene
sequence
site
sequencing read
attribute information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/089499
Other languages
English (en)
French (fr)
Inventor
胡志强
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to JP2021514554A priority Critical patent/JP7064654B2/ja
Priority to SG11202011523VA priority patent/SG11202011523VA/en
Priority to KR1020217020204A priority patent/KR20210116454A/ko
Publication of WO2020199336A1 publication Critical patent/WO2020199336A1/zh
Priority to US17/102,136 priority patent/US20210082539A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1093General methods of preparing gene libraries, not provided for in other subgroups
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1082Preparation or screening gene libraries by chromosomal integration of polynucleotide sequences, HR-, site-specific-recombination, transposons, viral vectors
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • CCHEMISTRY; METALLURGY
    • C40COMBINATORIAL TECHNOLOGY
    • C40BCOMBINATORIAL CHEMISTRY; LIBRARIES, e.g. CHEMICAL LIBRARIES
    • C40B40/00Libraries per se, e.g. arrays, mixtures
    • C40B40/04Libraries containing only organic compounds
    • C40B40/06Libraries containing nucleotides or polynucleotides, or derivatives thereof
    • C40B40/08Libraries containing RNA or DNA which encodes proteins, e.g. gene libraries
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/12Computing arrangements based on biological models using genetic models
    • G06N3/126Evolutionary algorithms, e.g. genetic algorithms or genetic programming
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • G16B40/20Supervised data analysis
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B5/00ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks

Definitions

  • the present disclosure relates to the field of computer technology, and in particular to a method, device and storage medium for identifying gene mutations.
  • the sequence of human genes can be determined through gene sequencing technology, and the analysis of gene sequence can be used as the basis for further genetic research and modification.
  • the gene sequencing technology greatly improves the efficiency of gene sequencing, reduces the cost of gene sequencing, and maintains the accuracy of gene sequencing. If the first-generation sequencing technology completes the sequencing of a human genome, it may take three years, while the second-generation sequencing technology can shorten the time to only one week.
  • the present disclosure proposes a gene mutation identification scheme.
  • a method for identifying gene mutations comprising:
  • sequence feature is a feature related to the location of the site
  • the gene variation at the gene variation candidate site is identified.
  • the attribute information includes sequence attribute information; determining the sequence characteristics of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read includes:
  • the gene location information of the gene mutation candidate site determine the preset site interval where the gene mutation candidate site is located
  • sequence attribute information is information characterizing gene attributes related to the position of the site
  • the sequence feature of the gene mutation candidate site is generated.
  • the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
  • the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
  • the obtaining the sequence attribute information of each site of the at least one gene sequencing read in the preset site interval includes:
  • the number of inserted genes of each gene type at each site of the at least one gene sequencing read is counted.
  • the sequence attribute information includes at least one of the following information:
  • the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
  • the attribute information includes non-sequence attribute information; and determining the non-sequence characteristics of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read includes:
  • non-sequence attribute information of the at least one gene sequencing read wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the site;
  • the non-sequence feature of the gene mutation candidate site is determined.
  • the non-sequence information includes at least one of the following information:
  • Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
  • the determining the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read includes:
  • the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
  • the determining the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read includes:
  • the non-sequence features corresponding to the gene mutation candidate sites are determined.
  • the identifying the gene mutation at the gene mutation candidate site based on the sequence feature and the non-sequence feature includes:
  • the gene mutation of the gene mutation candidate site is identified.
  • the identifying the gene mutation at the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site includes:
  • the variation value is greater than or equal to a preset threshold, it is determined that the gene at the gene variation candidate site has a variation.
  • the obtaining at least one gene sequencing read corresponding to the gene mutation candidate site includes:
  • a gene mutation identification device comprising:
  • the obtaining module is used to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
  • the determining module is configured to determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site;
  • the recognition module is used for recognizing the gene mutation of the gene mutation candidate site based on the sequence feature and the non-sequence feature.
  • the attribute information includes sequence attribute information; the determining module includes:
  • the first determining sub-module is used to determine the preset site interval where the gene mutation candidate site is located according to the gene location information of the gene mutation candidate site;
  • the first acquiring submodule is used to acquire the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is related to the position of the site Information that characterizes genetic attributes;
  • the first generation sub-module is used to generate the sequence characteristics of the gene mutation candidate sites according to the sequence attribute information of each site in the preset site interval.
  • the first acquisition submodule is specifically configured to determine the gene type of the at least one gene sequencing read at the each site; and count each site corresponding to each site. The number of genes for each gene type.
  • the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the missing gene at each locus; count the number of missing genes of each gene type at the at least one gene sequencing read at each locus.
  • the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the inserted gene at each site; count the number of inserted genes of each gene type at each site of the at least one gene sequencing read.
  • the sequence attribute information includes at least one of the following information:
  • the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
  • the attribute information includes non-sequence attribute information; the determining module includes:
  • the second acquisition submodule is used to acquire the non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
  • the second determining sub-module is used to determine the non-sequence characteristics of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
  • the non-sequence information includes at least one of the following information:
  • Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
  • the second determining submodule is specifically configured to determine the comparison quality of each gene sequencing read according to the comparison quality of each site in each gene sequencing read;
  • the comparison quality is used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing reads; according to the comparison quality of each gene sequencing read, the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
  • the second determining submodule is specifically configured to determine the positive and negative chain information of the gene chain to which each gene sequencing read belongs, according to the positive and negative chain information of the gene chain to which the at least one gene sequencing read belongs. Negative strand ratio; according to the positive-negative strand ratio, the non-sequence features corresponding to the gene mutation candidate sites are determined.
  • the identification module includes:
  • the integration sub-module is specifically used to perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
  • the recognition sub-module is used to identify the gene mutation of the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site.
  • the recognition submodule is specifically configured to obtain the mutation value of the gene at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site; When the value is greater than or equal to the preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
  • the acquisition module is specifically used for:
  • a gene mutation identification device including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the above method.
  • a non-volatile computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor.
  • the embodiments of the present disclosure provide for obtaining at least one gene sequencing read corresponding to a gene mutation candidate site, and the sequence feature and non-sequence feature of the gene mutation candidate site can be determined based on the attribute information of the at least one gene sequencing read.
  • the sequence feature and non-sequence feature of the gene to identify the gene mutation at the candidate site of the gene mutation.
  • sequence feature can be the feature related to the position of the locus
  • non-sequence feature can be the feature not related to the position of the locus, so that in the process of gene variation identification, the sequence feature of the gene can be combined with the non-sequence feature Analyze the characteristics of gene mutation sites more comprehensively, screen out germline gene mutations and interference caused by noise and errors, better identify gene mutations, and enhance the accuracy of gene mutation identification.
  • Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
  • Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
  • Fig. 3 shows a flowchart of the sequence characterization process of gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 4 shows a flowchart of the non-sequence feature process of gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure.
  • Fig. 6 shows a block diagram of a neural network model according to an embodiment of the present disclosure.
  • Fig. 7 shows a block diagram of a gene mutation identification device according to an embodiment of the present disclosure.
  • Fig. 8 shows a block diagram of a device for gene mutation identification according to an exemplary embodiment of the present disclosure.
  • the gene mutation identification scheme can obtain at least one gene sequencing read corresponding to the gene mutation candidate site, so that the gene mutation of the gene mutation candidate site can be identified based on the at least one gene sequencing read.
  • sequence features can be generated based on the sequence attribute information of at least one gene sequencing read, and non-sequence features can be generated based on the non-sequence attribute information of at least one gene sequencing read, and then the sequence feature and non-sequence feature can be paired
  • the gene mutation at the gene mutation candidate site is identified, so that the sequence attribute information and non-sequence attribute information of at least one gene sequencing read can be integrated, and the sequence attribute information of the gene sequencing read can be used more comprehensively.
  • the neural network model integrated by multimodal information can be used to extract the sequence features and non-sequence features of gene mutation candidate sites in the process of gene mutation identification, so that the sequence attribute information and non-sequence features of the gene sequence can be integrated.
  • Sequence attribute information analyze gene data more comprehensively, screen out germline gene mutations and interference caused by noise and errors, and better identify gene mutations. The following examples will illustrate the process of gene mutation identification in detail.
  • Fig. 1 shows a flowchart of a method for identifying gene mutations according to an embodiment of the present disclosure.
  • the gene mutation identification method can be executed by a gene mutation identification device or other processing equipment, where the gene mutation identification device can be User Equipment (UE), mobile equipment, user terminal, terminal, cellular phone, cordless phone, personal digital Processing (Personal Digital Assistant, PDA), handheld devices, computing devices, vehicle-mounted devices, wearable devices, etc., or the gene mutation recognition device may be a server.
  • the gene mutation identification method may be implemented by a processor calling computer-readable instructions stored in a memory.
  • the gene mutation identification method includes:
  • Step 11 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
  • the gene mutation recognition device can obtain the gene sequencing reads obtained by gene sequencing, and then obtain at least one gene sequencing read corresponding to the gene mutation candidate site from the gene sequencing reads obtained by the gene sequencing.
  • the gene sequencing reads here can be understood as gene sequences marked with gene types after gene sequencing, and the length of each gene sequencing read can be the same or different. In the case of different lengths, the length of each gene sequencing read segment can be within a preset length range, thereby ensuring that the length of each gene sequencing read segment is relatively close.
  • the gene type can be understood as the base type, and the gene type can include cytosine (C), guanine (G), adenine (A), and thymine (T), so that the gene sequencing read can be a gene sequence including AGCT.
  • the candidate gene mutation site here can be a site where the gene sequence is abnormal.
  • the locus of the gene sequence may indicate the position of the gene sequence.
  • the gene mutation candidate site corresponds to at least one gene sequencing read, wherein the at least one gene sequencing read is abnormal at this site.
  • each gene mutation candidate site may correspond to at least one gene sequencing read.
  • the embodiment of the present disclosure uses a gene mutation candidate site for description.
  • Step 12 Determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site.
  • the attribute information of the at least one gene sequencing read corresponding to the gene mutation candidate site can be extracted, and based on the extracted attribute information Generate the sequence feature and non-sequence feature of the gene mutation candidate site.
  • the attribute information may include sequence attribute information and non-sequence attribute information.
  • the sequence attribute information may be information related to the position of the locus that characterizes the gene attribute of the gene sequencing read.
  • Non-sequence attribute information can be information that is not restricted by the position of the site and can characterize gene attributes.
  • the sequence attribute information of at least one gene sequencing read at the candidate gene mutation site can be extracted, and the sequence of at least one gene sequencing read at a location near the gene mutation candidate site can also be extracted.
  • Property information when determining the sequence characteristics of gene mutation candidate sites, a neural network model with a convolutional layer and a pooling layer can be used to extract gene mutation candidate sites from at least one gene sequencing read corresponding to the gene mutation candidate site Sequence characteristics.
  • the neural network model can include two branch structures. One branch can extract sequence features of gene sequencing reads, and the branch can include a convolutional layer and a pooling layer; the other branch can extract non-sequence features of gene sequencing reads.
  • the neural network model can thus integrate multiple modal information (sequence attribute information and non-sequence attribute information) to identify gene mutations at gene mutation candidate sites.
  • the aforementioned neural network model can be used to extract the non-sequence features of at least one gene sequencing read from another branch of the neural network model.
  • the branch structure can include a fully connected layer, The fully connected layer can be used to extract non-sequential features that are not restricted by location.
  • Step 13 based on the sequence feature and the non-sequence feature, identify the gene variation of the gene variation candidate site.
  • the sequence feature and non-sequence feature can be fused to identify the gene mutation of the gene mutation candidate site, for example,
  • the aforementioned neural network model is used to determine whether the gene at the gene mutation candidate site is mutated, or whether the gene at the gene mutation candidate site is abnormal in gene sequence due to noise or other reasons.
  • the gene mutation of the gene mutation candidate site can be identified according to the sequence characteristics and non-sequence characteristics of the gene mutation candidate site, so that the gene sequencing data can be analyzed more comprehensively.
  • it is first necessary to obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
  • the example of the present disclosure also provides a process for obtaining at least one gene sequencing read corresponding to the gene mutation candidate site.
  • Fig. 2 shows a flow chart of obtaining at least one gene sequencing read corresponding to a gene mutation candidate site according to an embodiment of the present disclosure.
  • obtaining at least one gene sequencing read corresponding to the gene mutation candidate site may include the following steps:
  • Step 111 Obtain gene sequencing reads obtained by gene sequencing of somatic genes.
  • At least one gene sequencing read segment can be obtained through gene sequencing of the somatic cell gene, and the gene sequencing read segment can be a sequence that annotates the gene type of the somatic gene.
  • the gene sequencing read segment can be a sequence that annotates the gene type of the somatic gene.
  • At least one gene sequencing read can be obtained through gene sequencing of somatic genes, and the gene sequencing reads obtained by gene sequencing can be preprocessed.
  • the preprocessing methods here can include cross contamination screening, Sequencing quality screening, comparison quality screening, abnormal read length screening, etc. Through preprocessing, cross-contaminated gene sequencing reads can be screened out, and gene sequencing reads with low sequencing quality and comparison quality and abnormal read length can be screened out.
  • Step 112 Compare the gene sequence of the gene sequencing read with the gene sequence of the reference genome to obtain the comparison result.
  • the gene sequence of the obtained gene sequencing reads can be compared with the gene sequence of the reference genome at the same site, Get the comparison result.
  • each gene sequencing read obtained by performing gene sequencing can be compared with the gene sequence of the reference genome at the same site to determine the sites where the gene sequence of the gene sequencing read is different from the gene sequence of the reference genome. It is also possible to compare at least one gene sequencing read with the same site with the gene sequence of the reference genome at the same site to determine the site where the gene sequence of the at least one gene sequencing read is different from the gene sequence of the reference genome.
  • Step 113 According to the comparison result, it is determined that the gene of the somatic gene has an abnormal gene mutation candidate site.
  • a site that differs from the gene sequence of the gene sequencing read from the reference genome can be determined according to the comparison result. If at least one gene sequencing read corresponding to the site is in at least one gene sequencing read, the mutated gene is sent at that site If the ratio of sequencing reads is greater than the preset ratio, it can be determined that the site is a candidate site for gene mutation; otherwise, it can be considered that the site is not a candidate site for gene mutation.
  • the gene sequence of the gene sequencing read at this position is different from the gene sequence of the reference genome, which may be caused by sequencing errors. In this way, the abnormalities of gene sequence caused by gene sequencing errors can be reduced.
  • Step 114 Obtain at least one gene sequencing read corresponding to the gene mutation candidate site.
  • each gene mutation candidate site corresponds to at least one gene sequencing read, and the gene sequence at the gene mutation candidate site may be different from the gene sequence of the reference genome at the same site. There may be at least one gene mutation candidate site.
  • the sequence characteristics of the gene mutation candidate site can be determined according to the sequence attribute information of at least one gene sequencing read corresponding to the gene mutation candidate site, so that when the gene mutation of the gene mutation candidate site is identified , You can consider the sequence attributes of at least one gene sequencing read corresponding to the gene mutation candidate site.
  • the process of determining the sequence characteristics of gene mutation candidate sites will be described in detail below through an example.
  • Fig. 3 shows a flowchart of the sequence characterization process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 3, the above step 12 may include the following steps:
  • Step 121a according to the gene location information of the gene mutation candidate site, determine the preset site interval where the gene mutation candidate site is located;
  • Step 122a Obtain the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is information characterizing gene attributes related to the position of the site ;
  • Step 123a According to the sequence attribute information of each site in the preset site interval, the sequence feature of the gene mutation candidate site is generated.
  • the preset site interval in which the gene mutation candidate site is located can be determined according to the gene location information of the gene mutation candidate site. For example, 150 before and after the gene mutation candidate site can be determined. The interval of 1 base pair is used as the preset site interval where the candidate site of gene mutation is located.
  • Sequence features can be represented by sequence feature vectors.
  • At least one sequence feature vector corresponding to at least one site in the preset site interval where the gene mutation candidate site is located can form a sequence feature matrix of the gene mutation candidate site. For example, if the preset site interval where the gene mutation candidate site is located includes 3 sites b1, b2, b3, the sequence feature vectors corresponding to the 3 sites are a1, a2, and a3, respectively.
  • the sequence feature matrix is [a1a2a3], where the sequence features of a1, a2, and a3 correspond to the sequence attribute information of b1, b2, and b3, respectively.
  • the sequence attribute information may include, but is not limited to: the gene type of the reference genome; the number of genes of each gene type; the number of missing genes of each gene type; the number of inserted genes of each gene type.
  • the gene type of the reference genome may be the gene type of the reference genome at the gene mutation candidate site.
  • the number of genes of each gene type can be the number of genes of each gene type at the candidate site of at least one gene sequencing read of the gene mutation.
  • the candidate site of the gene mutation corresponds to 5 gene sequencing reads, and each gene is sequenced.
  • the gene types of the reads at the gene mutation candidate sites are: A, C, C, G, G, the number of genes for each gene type are: A is 1; C is 2; G is 2 .
  • the number of missing genes of each gene type can be the number of missing genes of each gene type at the candidate site of the gene mutation in at least one gene sequencing read, for example, the number of genes missing in each gene sequencing read at the candidate site of the gene mutation
  • the types are: A, C, C, G, G, and the number of missing genes for each gene type is: A is 1; C is 2; G is 2.
  • the number of inserted genes of each gene type can be the number of inserted genes of each gene type at the candidate site of the gene mutation of at least one gene sequencing read, for example, the gene inserted at the candidate site of the gene mutation of each gene sequencing read.
  • the types are: A, C, C, G, G, and the number of inserted genes for each gene type is: A is 1; C is 2; G is 2.
  • the sequence attribute information of at least one gene sequencing read at each site in the preset site interval it may be determined for each site in the preset site interval
  • the gene type of at least one gene sequencing read at the locus, and the number of genes of each gene type corresponding to the locus is counted, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined. Click the number of genes for each gene type.
  • the gene sequence of each gene sequencing read may be compared with the gene of the reference genome. Based on the comparison result of sequence comparison, for each site in the preset site interval, determine the gene type of the missing gene of each gene sequencing read at that site, and count at least one gene sequencing read at The number of missing genes of each gene type at this locus, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined, and the number of missing genes of each gene type at that locus can be determined.
  • the gene sequence of each gene sequencing read may be compared with the gene of the reference genome. Based on the comparison result of sequence comparison, for each site in the preset site interval, determine the gene type of the missing gene of each gene sequencing read at that site, and count at least one gene sequencing read at The number of inserted genes of each gene type at the locus, so that at least one gene sequencing read corresponding to the gene mutation candidate locus can be determined, and the number of inserted genes of each gene type at the locus can be determined.
  • the sequence attribute information includes the gene type of the reference genome, the number of genes of each gene type, the number of missing genes of each gene type, the number of inserted genes of each gene type, and the sequence in determining the candidate site of gene mutation
  • the above four information of at least one gene sequencing read corresponding to the gene mutation candidate site can be extracted at that site, for example, For the 5 gene sequencing reads corresponding to the candidate gene mutation site, for a certain site in the preset site interval, the gene type of the reference genome at that site can be determined respectively, and the 5 gene sequencing reads can be determined at that location.
  • the sequence characteristics of the gene mutation candidate sites may include the sequence characteristics of each site in the preset site interval.
  • Fig. 4 shows a flowchart of the non-sequence feature process of gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 4, the above step 12 may include the following steps:
  • Step 121b acquiring non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
  • Step 122b Generate the non-sequence feature of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
  • the non-sequence information may include at least one of the following information: comparison quality; positive and negative chain preference; gene sequencing read length; edge preference.
  • the non-sequence attribute information of at least one gene attribute sequence read can be obtained, and then the non-sequence feature of the gene mutation candidate site can be generated from the obtained non-sequence attribute information.
  • each gene sequencing read may be The comparison quality of the locus is to determine the comparison quality of each gene sequencing read, and then according to the comparison quality of each gene sequencing read, determine the non-sequence feature corresponding to the gene mutation candidate locus.
  • the comparison quality can be used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing read. If the comparison quality of a certain gene sequence is lower than the preset value, it can be considered that the gene sequence is a gene obtained by gene sequencing.
  • the quality of comparison can be used as a reference factor for judging whether the gene at the gene mutation candidate site has mutation.
  • the comparison quality of each gene sequencing read can be determined according to the comparison quality of each gene sequence. Taking a gene sequencing read as an example, you can The average or median value of the comparison quality of the gene sequences included in the gene sequencing reads is used as the comparison quality of the gene sequencing reads.
  • At least one gene sequence can also be randomly selected in the gene sequencing reads, and the selected at least one gene The average or intermediate value of the sequence comparison quality is used as the comparison quality of the sequence reads of the gene.
  • the comparison quality corresponding to the gene mutation candidate site is obtained. For example, the average or mean value of the comparison quality of at least one gene sequencing read corresponding to the gene mutation candidate site is calculated to obtain the The comparison quality corresponding to the gene mutation candidate site, so that the non-sequence features corresponding to the gene mutation candidate site can be determined according to the comparison quality corresponding to the gene mutation candidate site.
  • the positive and negative chain of the gene chain to which each gene sequencing read belongs can be determined.
  • Information determine the positive-negative chain ratio of the gene chain to which at least one gene sequencing read belongs, and then determine the non-sequence feature corresponding to the gene mutation candidate site according to the determined positive-negative chain ratio.
  • the positive-negative strand preference can be the ratio of the positive strand and the negative strand in the gene strand to which the gene sequencing read belongs.
  • the gene strand can include the positive strand and the negative strand, where the positive strand can be the same base sequence as the ribonucleic acid (RNA)
  • RNA ribonucleic acid
  • a single strand of deoxyribonucleic acid (DNA) the negative strand can be a single strand of deoxyribonucleic acid (DNA) complementary to the base sequence of ribonucleic acid (RNA).
  • gene mutation candidate sites correspond to 5 gene sequencing reads, among which, 3 gene sequencing reads correspond to the positive strand of the gene chain, and 2 gene sequencing reads correspond to the negative strand of the gene chain, and the positive and negative strands are preferred It can be 3:2.
  • the length of the gene sequencing read of each gene sequencing read can be used, Determine the non-sequence characteristics of candidate sites for gene mutations.
  • the length of a gene sequencing read can be the length of the base sequence of each gene sequencing read. For example, if a gene sequencing read includes 4 base sequences, the length of the gene sequencing read is 4, which can be determined by The length of each gene sequencing read determines the non-sequence feature of the gene mutation candidate site, and the non-sequence feature of the gene mutation candidate site can also be determined by the median or average of the length of at least one gene sequencing read.
  • the gene mutation when determining the non-sequence characteristics of the candidate gene mutation site based on the non-sequence attribute information of at least one gene sequencing read, the gene mutation can be determined according to the edge preference of each gene sequencing read Non-sequence features of candidate sites.
  • the marginal preference can be the ratio of the marginal position to the middle position of a certain site in the gene sequencing reads.
  • gene sequencing reads can be divided into 3 evenly, where the two segments at both ends of the gene sequencing read can be used as edge positions, and the middle segment of the gene sequencing read can be used as the middle position, and the gene mutation candidate site corresponds to
  • the edge preference of the gene mutation candidate site can be 3: 2.
  • the non-sequence characteristics of the gene mutation candidate site can be determined from the marginal preference of the gene mutation candidate site in each gene sequencing read, and the median value of the marginal preference corresponding to at least one gene sequencing read or The average value is used to determine the non-sequence characteristics of candidate sites of gene mutation.
  • the non-sequence feature of the gene mutation candidate site can be generated based on the non-sequence attribute information of at least one gene sequencing read at the gene mutation candidate site, so that the non-sequence of the gene mutation candidate site can be considered when gene mutation identification Characteristic features make gene mutation identification more accurate.
  • the non-sequence feature of at least one gene sequencing read may be generated from a combination of any at least one information in the non-sequence attribute information.
  • the following uses an example to illustrate the process of identifying gene mutations at gene mutation candidate sites.
  • Fig. 5 shows a flow chart of the gene mutation process of identifying gene mutation candidate sites according to an embodiment of the present disclosure. As shown in Figure 5, the above step 13 may include the following steps:
  • Step 131 Perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
  • Step 132 based on the integration characteristics of the gene mutation candidate site, identify the gene mutation of the gene mutation candidate site.
  • the neural network model can be used to perform feature integration on the sequence feature and the non-sequence feature, and the sequence feature matrix formed by the sequence feature can be combined with the non-sequence feature.
  • the non-sequence feature matrix formed by the sequence feature is synthesized into a feature matrix, and the integrated feature matrix formed by the integrated feature is obtained, and then the neural network model is used to identify the gene mutation of the mutation candidate site according to the integrated feature matrix.
  • the neural network model can be used to integrate sequence attribute information and non-sequence attribute information corresponding to gene mutation candidate sites, so that gene sequencing data can be analyzed more comprehensively, and gene mutation identification can be more accurate.
  • you can select gene sequencing reads with Single Nucleotide Polymorphism (SNP), and gene sequencing reads with Insertion/Deletion (InDel) as training samples to train The gene variation recognition model obtained afterwards can effectively identify SNP and InDel gene variation.
  • SNP Single Nucleotide Polymorphism
  • InDel Insertion/Deletion
  • identifying the genetic variation of the gene mutation candidate site according to the integration feature of the gene mutation candidate site may include: according to the integration feature of the gene mutation candidate site, Obtain the mutation value of the gene at the gene mutation candidate site; if the mutation value is greater than or equal to a preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
  • the mutation value of the gene mutation may be indicative of the possibility of mutation of the candidate site of the gene mutation. For example, the greater the mutation value, the greater the possibility of mutation of the candidate site of the gene mutation.
  • the above-mentioned neural network can be used to process the two-dimensional feature to obtain the mutation value, and to determine whether the gene at the gene mutation candidate site has mutation according to the mutation value.
  • the variation value can be between 0 and 1.
  • the preset threshold can be set according to the application scenario, for example, 0.3, 0.5. If the mutation value is greater than the preset threshold, it can be considered that the gene at the candidate site of gene mutation has been mutated, otherwise, it can be the gene at the candidate site of gene mutation. No mutation has occurred.
  • a neural network model can be used to identify gene mutations at gene mutation candidate sites, and the neural network model can extract sequence features and non-sequence features of gene mutation candidate sites.
  • the embodiment of the present disclosure also provides a structure of a neural network model.
  • Fig. 6 shows a block diagram of a neural network model according to an embodiment of the present disclosure.
  • the neural network model can include two branch structures, the first branch and the second branch.
  • the first branch may be used to extract the sequence features of at least one gene sequencing read corresponding to the gene mutation candidate site, and the first branch may include a convolutional layer and a pooling layer.
  • the second branch may be used to extract the non-sequence features of at least one gene sequencing read corresponding to the gene mutation candidate site, and the second branch may include a fully connected layer.
  • the sequence features and non-sequence features can be integrated, for example, the sequence feature matrix of sequence features and the non-sequence feature matrix of non-sequence features can be spliced together.
  • the integrated feature matrix of the integrated feature is obtained, and then the mutation value of the gene mutation candidate site can be obtained through the fully connected layer.
  • the sequence attribute information and non-sequence attribute information of at least one gene sequencing read corresponding to the gene mutation candidate site are extracted, and the gene mutation is identified by the integration feature integrating the sequence attribute information and the non-sequence attribute information, thereby Comprehensively consider the sequence attribute information and non-sequence attribute information corresponding to gene mutation candidate sites, analyze gene sequencing information more comprehensively, better identify gene mutations at gene candidate sites, and screen out germline gene mutations and noise and The interference caused by errors improves the accuracy of gene mutation identification.
  • the writing order of the steps does not mean a strict execution order but constitutes any limitation on the implementation process.
  • the specific execution order of each step should be based on its function and possibility.
  • the inner logic is determined.
  • Fig. 7 shows a block diagram of a gene mutation recognition device according to an embodiment of the present disclosure. As shown in Fig. 7, the gene mutation recognition device includes:
  • the obtaining module 71 is used to obtain at least one gene sequencing read corresponding to the gene mutation candidate site;
  • the determining module 72 is configured to determine the sequence feature and non-sequence feature of the gene mutation candidate site according to the attribute information of the at least one gene sequencing read, wherein the sequence feature is a feature related to the location of the site ;
  • the identification module 73 is configured to identify the gene mutation at the gene mutation candidate site based on the sequence feature and the non-sequence feature.
  • the attribute information includes sequence attribute information; the determining module 72 includes:
  • the first determining sub-module is used to determine the preset site interval where the gene mutation candidate site is located according to the gene location information of the gene mutation candidate site;
  • the first acquiring submodule is used to acquire the sequence attribute information of each site in the preset site interval of the at least one gene sequencing read; wherein, the sequence attribute information is related to the position of the site Information that characterizes genetic attributes;
  • the first generation sub-module is used to generate the sequence characteristics of the gene mutation candidate sites according to the sequence attribute information of each site in the preset site interval.
  • the first acquisition submodule is specifically configured to determine the gene type of the at least one gene sequencing read at the each site; and count each site corresponding to each site. The number of genes for each gene type.
  • the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the missing gene at each locus; count the number of missing genes of each gene type at the at least one gene sequencing read at each locus.
  • the first acquisition submodule is specifically configured to determine each gene sequencing read based on the comparison result of the gene sequence of each gene sequencing read and the gene sequence of the reference genome. Segment the gene type of the inserted gene at each site; count the number of inserted genes of each gene type at each site of the at least one gene sequencing read.
  • the sequence attribute information includes at least one of the following information:
  • the gene type of the reference gene the number of genes for each gene type; the number of missing genes for each gene type; the number of inserted genes for each gene type.
  • the attribute information includes non-sequence attribute information; the determining module includes:
  • the second acquisition submodule is used to acquire the non-sequence attribute information of the at least one gene sequencing read; wherein the non-sequence attribute information is information that characterizes gene attributes that is not related to the position of the locus;
  • the second determining sub-module is used to determine the non-sequence characteristics of the gene mutation candidate site according to the non-sequence attribute information of the at least one gene sequencing read.
  • the non-sequence information includes at least one of the following information:
  • Contrast quality positive and negative chain preference; gene sequencing read length; edge preference.
  • the second determining submodule is specifically configured to determine the comparison quality of each gene sequencing read according to the comparison quality of each site in each gene sequencing read;
  • the comparison quality is used to characterize the accuracy of the gene sequencing of each gene sequence in the gene sequencing reads; according to the comparison quality of each gene sequencing read, the non-sequence characteristics corresponding to the gene mutation candidate sites are determined.
  • the second determining submodule is specifically configured to determine the positive and negative chain information of the gene chain to which each gene sequencing read belongs, according to the positive and negative chain information of the gene chain to which the at least one gene sequencing read belongs. Negative strand ratio; according to the positive-negative strand ratio, the non-sequence features corresponding to the gene mutation candidate sites are determined.
  • the identification module 73 includes:
  • the integration sub-module is specifically used to perform feature integration of the sequence feature and the non-sequence feature to obtain the integration feature of the gene mutation candidate site;
  • the recognition sub-module is used to identify the gene mutation of the gene mutation candidate site based on the integration characteristics of the gene mutation candidate site.
  • the recognition submodule is specifically configured to obtain the mutation value of the gene at the gene mutation candidate site according to the integration characteristics of the gene mutation candidate site; When the value is greater than or equal to the preset threshold, it is determined that the gene at the gene mutation candidate site has mutation.
  • the obtaining module 71 is specifically configured to:
  • the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
  • the functions or modules contained in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments.
  • Fig. 8 is a block diagram showing a device 1900 for gene mutation identification according to an exemplary embodiment.
  • the device 1900 may be provided as a server.
  • the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions that can be executed by the processing component 1922, such as application programs.
  • the application program stored in the memory 1932 may include one or more modules each corresponding to a set of instructions.
  • the processing component 1922 is configured to execute instructions to perform the above-described methods.
  • the device 1900 may also include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to the network, and an input output (I/O) interface 1958.
  • the device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
  • a non-volatile computer-readable storage medium such as the memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to complete the foregoing method.
  • the present disclosure may be a system, method, and/or computer program product.
  • the computer program product may include a computer-readable storage medium loaded with computer-readable program instructions for enabling a processor to implement various aspects of the present disclosure.
  • the computer-readable storage medium may be a tangible device that can hold and store instructions used by the instruction execution device.
  • the computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
  • Computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) Or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a printer with instructions stored thereon
  • RAM random access memory
  • ROM read-only memory
  • EPROM erasable programmable read-only memory
  • flash memory flash memory
  • SRAM static random access memory
  • CD-ROM compact disk read-only memory
  • DVD digital versatile disk
  • memory stick floppy disk
  • mechanical encoding device such as a printer with instructions stored thereon
  • the computer-readable storage medium used here is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (for example, light pulses through fiber optic cables), or through wires Transmission of electrical signals.
  • the computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing/processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and/or a wireless network.
  • the network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and/or edge servers.
  • the network adapter card or network interface in each computing/processing device receives computer-readable program instructions from the network, and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing/processing device .
  • the computer program instructions used to perform the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, status setting data, or in one or more programming languages.
  • Source code or object code written in any combination, the programming language includes object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" language or similar programming languages.
  • Computer-readable program instructions can be executed entirely on the user's computer, partly on the user's computer, executed as a stand-alone software package, partly on the user's computer and partly executed on a remote computer, or entirely on the remote computer or server carried out.
  • the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to access the Internet connection).
  • LAN local area network
  • WAN wide area network
  • an electronic circuit such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can be customized by using the status information of the computer-readable program instructions.
  • the computer-readable program instructions are executed to realize various aspects of the present disclosure.
  • These computer-readable program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device , A device that implements the functions/actions specified in one or more blocks in the flowchart and/or block diagram is produced. It is also possible to store these computer-readable program instructions in a computer-readable storage medium. These instructions make computers, programmable data processing apparatuses, and/or other devices work in a specific manner, so that the computer-readable medium storing instructions includes An article of manufacture, which includes instructions for implementing various aspects of the functions/actions specified in one or more blocks in the flowchart and/or block diagram.
  • each block in the flowchart or block diagram may represent a module, program segment, or part of an instruction, and the module, program segment, or part of an instruction contains one or more functions for implementing the specified logical function.
  • Executable instructions may also occur in a different order from the order marked in the drawings. For example, two consecutive blocks can actually be executed in parallel, or they can sometimes be executed in the reverse order, depending on the functions involved.
  • each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or actions Or it can be realized by a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Chemical & Material Sciences (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Theoretical Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Biotechnology (AREA)
  • Genetics & Genomics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Organic Chemistry (AREA)
  • Evolutionary Biology (AREA)
  • Medical Informatics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Data Mining & Analysis (AREA)
  • Biomedical Technology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Analytical Chemistry (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Evolutionary Computation (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Computational Linguistics (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • General Physics & Mathematics (AREA)
  • Biochemistry (AREA)
  • Microbiology (AREA)
  • Physiology (AREA)
  • Databases & Information Systems (AREA)
  • Immunology (AREA)
  • Public Health (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Epidemiology (AREA)
  • Bioethics (AREA)

Abstract

本公开提供了一种基因变异识别方法、装置和存储介质,其中,该方法包括:获取基因变异候选位点对应的至少一个基因测序读段;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。本公开实施例的可以将基因的序列特征和非序列特征相结合,分析基因变异位点的特征。

Description

一种基因变异识别方法、装置和存储介质
本公开要求在2019年3月29日提交中国专利局、申请号为201910251891.0、申请名称为“一种基因变异识别方法、装置和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开涉及计算机技术领域,尤其涉及一种基因变异识别方法、装置和存储介质。
背景技术
随着生物技术的发展,通过基因测序技术可以测定人类基因的序列,基因序列的分析可以作为进一步基因研究和改造的基础。目前,基因的二代测序技术相比于一代测序技术而言,极大地提高了基因测序的效率,降低了基因测序的成本,并且保持了基因测序的准确行性。第一代测序技术如果完成一个人类基因组的测序可能需要3年的时间,而使用二代测序技术则可以将时间缩短为仅仅1周。
发明内容
有鉴于此,本公开提出了一种基因变异识别方案。
根据本公开的一方面,提供了一种基因变异识别方法,所述方法包括:
获取基因变异候选位点对应的至少一个基因测序读段;
根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;
基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述属性信息包括序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征,包括:
根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
在一种可能的实现方式中,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
确定所述至少一个基因测序读段在所述每个位点的基因类型;
统计所述每个位点对应的每种基因类型的基因数量。
在一种可能的实现方式中,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;
统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
在一种可能的实现方式中,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;
统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
在一种可能的实现方式中,所述序列属性信息包括以下至少一种信息:
参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
在一种可能的实现方式中,所述属性信息包括非序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的非序列特征,包括:
获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
在一种可能的实现方式中,所述非序列信息包括以下至少一种信息:
对比质量;正负链偏好;基因测序读段长度;边缘偏好。
在一种可能的实现方式中,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:
根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;
根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:
根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;
根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别,包括:
将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别,包括:
根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;
在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述获取基因变异候选位点对应的至少一个基因测序读段,包括:
获取由体细胞基因进行基因测序得到的基因测序读段;
将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;
根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
获取所述基因变异候选位点对应的至少一个基因测序读段。
根据本公开的另一方面,提供了一种基因变异识别装置,所述装置包括:
获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;
确定模块,用于根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;
识别模块,用于基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述属性信息包括序列属性信息;所述确定模块,包括:
第一确定子模块,用于根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
第一生成子模块,用于根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
在一种可能的实现方式中,所述第一获取子模块,具体用于确定所述至少一个基因测序读段在所述每个位点的基因类型;统计所述每个位点对应的每种基因类型的基因数量。
在一种可能的实现方式中,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
在一种可能的实现方式中,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
在一种可能的实现方式中,所述序列属性信息包括以下至少一种信息:
参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
在一种可能的实现方式中,所述属性信息包括非序列属性信息;所述确定模块,包括:
第二获取子模块,用于获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
第二确定子模块,用于根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
在一种可能的实现方式中,所述非序列信息包括以下至少一种信息:
对比质量;正负链偏好;基因测序读段长度;边缘偏好。
在一种可能的实现方式中,所述第二确定子模块,具体用于根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述第二确定子模块,具体用于根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述识别模块,包括:
整合子模块,具体用于将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
识别子模块,用于基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述识别子模块,具体用于根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述获取模块,具体用于,
获取由体细胞基因进行基因测序得到的基因测序读段;
将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;
根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
获取所述基因变异候选位点对应的至少一个基因测序读段。
根据本公开的另一方面,提供了一种基因变异识别装置,包括:处理器;用于存储处理器可执行指令的存储器;其中,所述处理器被配置为执行上述方法。
根据本公开的另一方面,提供了一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其中,所述计算机程序指令被处理器执行时实现上述方法。
本公开实施例提供获取基因变异候选位点对应的至少一个基因测序读段,可以根据至少一个基因测序读段的属性信息,确定基因变异候选位点的序列特征和非序列特征,从而可以基于确定的序列特征和非序列特征对基因变异候选位点的基因变异进行识别。这里,序列特征可以是与位点的位置相关的特征,非序列特征可以是与位点的位置不相关的特征,从而在基因变异识别过程中,可以将基因的序列特征和非序列特征相结合,更加全面地分析基因变异位点的特征,筛掉胚系基因变异以及由于噪声和错误带来的干扰,更好地对基因变异进行识别,增强基因变异识别的准确性。
根据下面参考附图对示例性实施例的详细说明,本公开的其它特征及方面将变得清楚。
附图说明
包含在说明书中并且构成说明书的一部分的附图与说明书一起示出了本公开的示例性实施例、特征和方面,并且用于解释本公开的原理。
图1示出根据本公开一实施例的基因变异识别方法的流程图。
图2示出根据本公开一实施例的获取基因变异候选位点对应的至少一个基因测序读段的流程图。
图3示出根据本公开一实施例的基因变异候选位点的序列特征过程的流程图。
图4示出根据本公开一实施例的基因变异候选位点的非序列特征过程的流程图。
图5示出根据本公开一实施例的识别基因变异候选位点的基因变异过程的流程图。
图6示出根据本公开一实施例的神经网络模型的框图。
图7示出根据本公开一实施例的基因变异识别装置的框图。
图8示出根据本公开一示例性实施例示出的一种用于基因变异识别的装置的框图。
具体实施方式
以下将参考附图详细说明本公开的各种示例性实施例、特征和方面。附图中相同的附图标记表示功能相同或相似的元件。尽管在附图中示出了实施例的各种方面,但是除非特别指出,不必按比例绘制附图。
在这里专用的词“示例性”意为“用作例子、实施例或说明性”。这里作为“示例性”所说明的任何实施例不必解释为优于或好于其它实施例。
另外,为了更好的说明本公开,在下文的具体实施方式中给出了众多的具体细节。本领域技术人员应当理解,没有某些具体细节,本公开同样可以实施。在一些实例中,对于本领域技术人员熟知的方法、手段、元件和电路未作详细描述,以便于凸显本公开的主旨。
本公开实施例提供的基因变异识别方案,可以获取基因变异候选位点对应的至少一个基因测序读段,从而可以根据至少一个基因测序读段对基因变异候选位点的基因变异进行识别。在基因变异识别过程中,可以根据至少一个基因测序读段的序列属性信息生成序列特征,根据至少一个基因测序读段的非序列属性信息生成非序列特征,然后可以通过序列特征和非序列特征对基因变异候选位点的基因变异进行识别,从而可以整合至少一个基因测序读段的序列属性信息和非序列属性信息,更加全面地利用基因测序读段的序列属性信息。
在相关技术中,通常是利用支持向量机、随机森林等传统随机森林等传统机器学习方法进行基因变异识别,这种方式虽然实现简单,但难以利用基因变异候选位点附近基因序列的序列属性信息,基因变异识别的效果在基因数据量增加到一定程度之后会陷入瓶颈。还有一些相关技术采用深度学习方法,利用神经网络对基因变异进行识别。但是,神经网络难以整合基因序列的非序列信息,无法对基因数据进行更加全面地分析。在本公开实施例中,在基因变异识别过程中可以利用由多模态信息整合的神经网络模型提取基因变异候选位点的序列特征和非序列特征,从而可以综合基因序列的序列属性信息和非序列属性信息,更加全面地对基因数据进行分析,筛掉胚系基因变异以及由于噪声和错误带来的干扰,更好地对基因变异进行识别。下述实施例将会对基因变异识别过程作详细说明。
图1示出根据本公开一实施例的基因变异识别方法的流程图。该基因变异识别方法可以由基因变异识别装置或其它处理设备执行,其中,基因变异识别装置可以为用户设备(User Equipment,UE)、移动设备、用户终端、终端、蜂窝电话、无绳电话、个人数字处理(Personal Digital Assistant,PDA)、手持设备、计算设备、车载设备、可穿戴设备等,或者,基因变异识别装置可以为服务器。在一些可能的实现方式中,该基因变异识别方法可以通过处理器调用存储器中存储的计算机可读指令的方式来实现。
如图1所示,该基因变异识别方法包括:
步骤11,获取基因变异候选位点对应的至少一个基因测序读段。
在本公开实施例中,基因变异识别装置可以获取由基因测序得到的基因测序读段,然后在基因测序得到的基因测序读段中,获取基因变异候选位点对应的至少一个基因测序读段。这里的基因测序读段可以理解为经过基因测序后标注有基因类型的基因序列,每个基因测序读段的长度可以相同也可以不同。在长度不同的情况下,每个基因测序读段的长度可以在预设长度范围内,从而可以保证每个基因测序读段的长度比较接近。基因类型可以理解为碱基类型,基因类型可以包括胞嘧啶(C)、鸟嘌呤(G)、腺嘌呤(A)、胸腺嘧啶(T),从而基因测序读段可以是包括AGCT的基因序列。这里的基因 变异候选位点可以是基因序列存在异常的位点。基因序列的位点可以表示基因序列的位置,针对每个位点,可以存在至少一个基因测序读段,即,在同一个位点可以存在由基因测序得到的至少一个基因测序读段。相应地,基因变异候选位点对应至少一个基因测序读段,其中,这至少一个基因测序读段都在这一位点上出现异常。基因变异候选位点可以为至少一个,每个基因变异候选位点可以对应至少一个基因测序读段。为了便于理解,本公开实施例以一个基因变异候选位点进行说明。
步骤12,根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征。
在本公开实施例中,在获取基因变异候选位点对应的至少一个基因测序读段之后,可以提取该基因变异候选位点对应的至少一个基因测序读段的属性信息,并根据提取的属性信息生成该基因变异候选位点的序列特征和非序列特征。属性信息可以包括序列属性信息和非序列属性信息。序列属性信息可以是与位点的位置相关的表征基因测序读段的基因属性的信息。非序列属性信息可以是不受到位点的位置限制并且可以表征基因属性的信息。在提取属性信息时,可以随机选择该基因候选位点对应的若干个基因测序读段,提取随机选择的若干个基因测序读段的属性信息;还可以提取该基因候选位点对应的每个基因测序读段的属性信息。
这里,在提取序列属性信息时,可以提取至少一个基因测序读段在该基因变异候选位点的序列属性信息,还可以提取至少一个基因测序读段在该基因变异候选位点附近位点的序列属性信息。这里,在确定基因变异候选位点的序列特征时,可以利用带有卷积层和池化层的神经网络模型,对基因变异候选位点对应的至少一个基因测序读段提取基因变异候选位点的序列特征。该神经网络模型可以包括两个分支结构,其中一个分支可以提取基因测序读段的序列特征,该分支可以包括卷积层和池化层;另一个分支可以提取基因测序读段的非序列特征。该神经网络模型从而可以整合多种模态信息(序列属性信息和非序列属性信息),对基因变异候选位点的基因变异进行识别。在确定基因变异候选位点的非序列特征时,可以利用上述神经网络模型,由该神经网络模型的另一个分支提取至少一个基因测序读段的非序列特征,该分支结构可以包括全连接层,全连接层可以用于提取不受位置限制的非序列特征。
步骤13,基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
在本公开实施方式中,在确定基因变异候选位点的序列特征和非序列特征之后,可将序列特征和非序列特征进行融合,对该基因变异候选位点的基因变异进行识别,例如,可以利用上述神经网络模型判断该基因变异候选位点的基因是否变异,或者,该基因变异候选位点的基因是否是由于噪声等原因而导致的基因序列异常。
本公开实施例中可以根据基因变异候选位点的序列特征和非序列特征对基因变异候选位点的基因变异进行识别,从而可以更加全面地对基因测序数据进行分析。在对基因变异候选位点的基因变异进行识别时,首先需要获取基因变异候选位点对应的至少一个基因测序读段。本公开实例还提供了一种获取基因变异候选位点对应的至少一个基因测序读段的过程。
图2示出根据本公开一实施例的获取基因变异候选位点对应的至少一个基因测序读段的流程图。在一种可能的实现方式中,获取基因变异候选位点对应的至少一个基因测序读段,可以包括以下步骤:
步骤111,获取由体细胞基因进行基因测序得到的基因测序读段。
这里,通过体细胞基因进行基因测序可以得到至少一个基因测序读段,基因测序读段可以是对体细胞基因进行基因类型标注的序列。体细胞基因在进行基因测序之后,不仅可以得到基因测序读段中每个基因的基因类型,还可以得到基因测序读段中每个基因所在位点的基因位置信息。同一个位点可以对应至少一个基因测序读段。
在一种可能的实现方式中,通过体细胞基因进行基因测序可以得到至少一个基因测序读段,可以对基因测序得到的基因测序读段进行预处理,这里的预处理方式可以包括交叉污染筛选、测序质量筛选、比对质量筛选、读段长度异常筛选等。通过预处理,可以筛选掉交叉污染的基因测序读段,以及筛选掉测序质量和比对质量较低、读段长度异常的基因测序读段。
步骤112,将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果。
在本公开实施例中,在获取由体细胞基因进行基因测序得到的基因测序读段之后,可以将获取的基因测序读段的基因序列与相同位点的参考基因组的基因序列的进行比对,得到对比结果。举例来说,可以将每个进行基因测序得到的基因测序读段与相同位点的参考基因组的基因序列进行对比,确定基因测序读段的基因序列与参考基因组的基因序列不同的位点。还可以将具有相同位点的至少一个基因测序读段与相同位点的参考基因组的基因序列进行对比,确定至少一个基因测序读段的基因序列与参考基因组的基因序列不同的位点。
步骤113,根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点。
在本公开实施例中,可以根据比对结果确定基因测序读段与参考基因组的基因序列不同的位点,如果该位点对应的至少一个基因测序读段中,在该位点发送变异的基因测序读段的比例大于预设比例,则可以确定该位点为基因变异候选位点,否则,可以认为该位点不是基因变异候选位点。基因测序读段在该位点与参考基因组的基因序列不同,可能是因为测序错误导致的不同,通过这种方式,可以减少由于基因测序失误引起的基因序列异常现象。
步骤114,获取所述基因变异候选位点对应的至少一个基因测序读段。
在本公开实施例中,在确定基因变异候选位点之后,可以获取基因变异候选位点对应的至少一个基因测序读段。其中,每个基因变异候选位点对应的至少一个基因测序读段,在该基因变异候选位点的基因序列与相同位点的参考基因组的基因序列可以不同。这里的基因变异候选位点可以为至少一个。
通过上述获取基因变异候选位点对应的至少一个基因测序读段的过程,不仅可以较为准确地确定基因变异候选位点,还可以在基因测序得到的基因测序读段中确定基因变异候选位点对应的至少一个基因测序读段。
本公开实施例中可以根据基因变异候选位点对应的至少一个基因测序读段的序列属性信息,确定该基因变异候选位点的序列特征,从而在对基因变异候选位点的基因变异进行识别时,可以考虑基因变异候选位点所对应的至少一个基因测序读段的序列属性。下面通过一示例对确定基因变异候选位点的序列特征的过程进行详细说明。
图3示出根据本公开一实施例的基因变异候选位点的序列特征过程的流程图。如图3所示,上述步骤12可以包括以下步骤:
步骤121a,根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
步骤122a,获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
步骤123a,根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
在本公开实施例的示例中,对于每一个基因变异候选位点可以存在至少一个基因测序读段。为了提高基因变异识别的准确度,不仅可以考虑该基因变异候选位点的序列属性信息,还可以考虑该基因变异候选位点附近的位点的序列属性信息。在确定基因变异候选位点的序列特征时,可以根据基因变异候选位点的基因位置信息,确定该基因变异候选位点所在的预设位点区间,例如,可以将基因变异候选位点前后150个碱基对的区间作为基因变异候选位点所在的预设位点区间。然后可以针对该预设位点区间内的每个位点,获取至少一个基因测序读段在该位点的序列属性信息,由该位点的序列属性信息可以生成该位点对应序列特征。序列特征可以用序列特征向量进行表示。由基因变异候选位点所在预设位点区间中至少一个位点对应的至少一个序列特征向量,可以形成基因变异候选位点的序列特征矩阵。举例来说,若基因变异候选位点所在预设位点区间包括3个位点b1、b2、b3,3个位点对应的序列特征向量分别为a1、a2、a3,基因变异候选位点的序列特征矩阵为[a1a2a3],其中,a1、a2、a3的序列特征分别对应b1、b2、b3的序列属性信息。
这里,序列属性信息可以包括但不限于:参考基因组的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。参考基因组的基因类型可以是参考基因组在基因变异候选位点的基因类型。每种基因类型的基因数量可以是至少一个基因测序读段在该基因变异候选位点每种基因类型的基因数量,例如,该基因变异候选位点对应5个基因测序读段,每个基因测序读段在该基因变异候选位点的基因类型分别为:A、C、C、G、G,则每种基因类型的基因数量分别为:A为1个;C为2个;G为2个。每种基因类型的缺失基因数量可以是至少一个基因测序读段在该基因变异候选位点每种基因类型的缺失基因数量,例如,每个基因测序读段在该基因变异候选位点缺失的基因类型分别为:A、C、C、G、G,则每种基因类型的缺失基因数量分别为:A为1个;C为2个;G为2个。每种基因类型的插入基因数量可以是至少一个基因测序读段在该基因变异候选位点每种基因类型的插入基因数量,例如,每个基因测序读段在该基因变异候选位点插入的基因类型分别为:A、C、C、G、G,则每种基因类型的插入基因数量分别为:A为1个;C为2个;G为2个。
在一种可能的实现方式中,在获取至少一个基因测序读段在预设位点区间中每个位点的序列属性信息时,可以针对该预设位点区间中的每个位点,确定至少一个基因测序读段在该位点的基因类型,并统计该位点所对应的每种基因类型的基因数量,从而可以确定基因变异候选位点对应的至少一个基因测序读段,在该位点每种基因类型的基因数量。
在一种可能的实现方式中,在获取至少一个基因测序读段在预设位点区间中每个位点的序列属性信息时,可以根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,针对该预设位点区间中的每个位点,确定每个基因测序读段在该位点的缺失基因的基因类型,并统计至少一个基因测序读段在该位点上每种基因类型的缺失基因数量,从而可以确定基因变异候选位点对应的至少一个基因测序读段,在该位点每种基因类型的缺失基因数量。
在一种可能的实现方式中,在获取至少一个基因测序读段在预设位点区间中每个位点的序列属性 信息时,可以根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,针对该预设位点区间中的每个位点,确定每个基因测序读段在该位点的缺失基因的基因类型,并统计至少一个基因测序读段在该位点上每种基因类型的插入基因数量,从而可以确定基因变异候选位点对应的至少一个基因测序读段,在该位点每种基因类型的插入基因数量。
举例来说,假设序列属性信息包括参考基因组的基因类型、每种基因类型的基因数量、每种基因类型的缺失基因数量、每种基因类型的插入基因数量,在确定基因变异候选位点的序列特征时,可以针对基因变异候选位点所在的预设位点区间中的每一个位点,提取基因变异候选位点对应的至少一个基因测序读段在该位点的上述四个信息,例如,基因变异候选位点对应的5个基因测序读段,针对预预设位点区间中的某一位点,可以分别确定参考基因组在该位点的基因类型、5个基因测序读段在该位点各基因类型的基因数量、5个基因测序读段在该位点各基因类型的缺失基因数量和5个基因测序读段在该位点各基因类型的插入基因数量。然后综合该位点对应的至少一个序列属性信息,可以得到该位点的序列特征。基因变异候选位点的序列特征可以包括预设位点区间中每个位点的序列特征。
本公开实施例的示例中不仅在对基因变异候选位点的基因变异进行识别时,考虑了基因变异候选位点所对应的至少一个基因测序读段的序列属性,还考虑了至少一个基因测序读段的非序列属性。下面通过一示例对确定基因变异候选位点的非序列特征的过程进行详细说明。
图4示出根据本公开一实施例的基因变异候选位点的非序列特征过程的流程图。如图4所示,上述步骤12可以包括以下步骤:
步骤121b,获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
步骤122b,根据所述至少一个基因测序读段的非序列属性信息,生成所述基因变异候选位点的非序列特征。
在本公开实施例的示例中,为了提高基因变异识别的准确度,不仅可以考虑至少一个基因测序读段的序列属性信息,还可以考虑至少一个基因测序读段的非序列属性信息。这里,非序列信息可以包括以下至少一种信息:对比质量;正负链偏好;基因测序读段长度;边缘偏好。在确定基因变异候选位点的非序列特征时,可以获取至少一个基因属性序列读段的非序列属性信息,然后由获取的非序列属性信息生成基因变异候选位点的非序列特征。
在一种可能的实现方式中,在根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征时,可以根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量,然后根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。这里,对比质量可以用于表征基因测序读段中每个基因序列的基因测序的准确性,如果某个基因序列的对比质量低于预设值,则可以认为该基因序列由基因测序得到的基因类型不准确,从而可以将对比质量作为判断基因变异候选位点的基因是否发生变异的一个参考因素。举例来说,基因变异候选位点对应至少一个基因测序读段,则可以根据每个基因序列的对比质量,确定每个基因测序读段的对比质量,以一个基因测序读段举例,可以将该基因测序读段所包括的基因序列的对比质量的平均值或者中间值,作为该基因测序读段的对比质量,还可以在该基因测序读段随机选择至少一个基因序列,将选择的至少一个基因序列对比质量的平均值或者中间值作为该基因测序读段的对比质量。然后 由每个基因测序读段的对比质量得到该基因变异候选位点对应的对比质量,例如,计算该基因变异候选位点对应的至少一个基因测序读段对比质量的平均值或者均值,得到该基因变异候选位点对应的对比质量,从而可以根据该基因变异候选位点对应的对比质量确定基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,在根据至少一个基因测序读段的非序列属性信息,确定基因变异候选位点的非序列特征时,可以根据每个基因测序读段所属基因链的正负链信息,确定至少一个基因测序读段所属基因链的正负链比例,然后根据确定的正负链比例,确定基因变异候选位点对应的非序列特征。这里,正负链偏好可以是基因测序读段所属基因链中正链和负链的比例,基因链可以包括正链和负链,其中,正链可以是与核糖核酸(RNA)的碱基序列相同的脱氧核糖核酸(DNA)单链,负链可以是与核糖核酸(RNA)的碱基序列互补的脱氧核糖核酸(DNA)单链。举例来说,基因变异候选位点对应5个基因测序读段,其中,3个基因测序读段对应基因链的正链,2个基因测序读段对应基因链的负链,则正负链偏好可以是3:2。
在一种可能的实现方式中,在根据至少一个基因测序读段的非序列属性信息,确定基因变异候选位点的非序列特征时,可以根据每个基因测序读段的基因测序读段长度,确定基因变异候选位点的非序列特征。基因测序读段长度可以是每个基因测序读段所具有碱基序列的长度,举例来说,一个基因测序读段包括4个碱基序列,则该基因测序读段的长度为4,可以由每个基因测序读段长度确定基因变异候选位点的非序列特征,还可以由至少一个基因测序读段长度的中间值或者平均值确定基因变异候选位点的非序列特征。
在一种可能的实现方式中,在根据至少一个基因测序读段的非序列属性信息,确定基因变异候选位点的非序列特征时,可以根据每个基因测序读段的边缘偏好,确定基因变异候选位点的非序列特征。这里,边缘偏好可以是某一位点在基因测序读段中位于边缘位置与中间位置的比例。举例来说,可以将基因测序读段平均分为3段,其中,基因测序读段两端的2段可以作为边缘位置,基因测序读段中间的1段可以作为中间位置,基因变异候选位点对应5个基因测序读段,基因变异候选位点如果位于其中3个基因测序读段的边缘位置,位于其中2个基因测序读段的中间位置,该基因变异候选位点的边缘偏好可以为3:2。相应地,可以由基因变异候选位点在每个基因测序读段的边缘偏好,确定基因变异候选位点的非序列特征,还可以由至少一个基因测序读段所对应的边缘偏好的中间值或者平均值,确定基因变异候选位点的非序列特征。
通过上述方式,可以针对至少一个基因测序读段在基因变异候选位点的非序列属性信息生成基因变异候选位点的非序列特征,从而可以在基因变异识别时考虑基因变异候选位点的非序列特征度特征,使基因变异识别更加准确。在确定非序列特征时,可以是由非序列属性信息中任意至少一个信息的组合生成至少一个基因测序读段的非序列特征。
下面通过一示例对基因变异候选位点的基因变异进行识别的过程进行说明。
图5示出根据本公开一实施例的识别基因变异候选位点的基因变异过程的流程图。如图5所示,上述步骤13可以包括以下步骤:
步骤131,将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
步骤132,基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识 别。
在本公开实施例中,在确定基因变异候选位点的序列特征和非序列维度特征之后,可以利用神经网络模型对序列特征和非序列特征进行特征整合,将序列特征形成的序列特征矩阵与非序列特征形成的非序列特征矩阵合成为一个特征矩阵,得到由整合特征形成的整合特征矩阵,然后利用神经网络模型根据该整合特征矩阵对变异候选位点的基因变异进行识别。通过这种方式,可以利用神经网络模型整合基因变异候选位点对应的序列属性信息和非序列属性信息,从而可以更加全面地对基因测序数据进行分析,使基因变异识别更加准确。在训练过程中,可以选取存在单核苷酸多态性(Single Nucleotide Polymorphism,SNP)的基因测序读段、存在插入/缺失(Insertion/Deletion,InDel)的基因测序读段作为训练样本,从而训练后得到的基因变异识别模型可以有效地对SNP、InDel的基因变异进行识别。
在一种可能的实现方式中,根据所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别,可以包括:根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。这里,基因发生变异的变异值可以是表征该基因变异候选位点发生变异的可能性,例如,变异值越大,该基因变异候选位点发生变异的可能性越大。可以利用上述神经网络对二维特征进行处理得到变异值,并根据变异值判断基因变异候选位点的基因是否存在变异。在一种可能的实现方式中,变异值可以在0至1之间。预设阈值可以根据应用场景进行设置,例如,0.3、0.5,如果变异值大于预设阈值,则可以认为该基因变异候选位点的基因发生变异,否则,可以为该基因变异候选位点的基因未发生变异。
本公开实施例中可以利用神经网络模型对基因变异候选位点的基因变异进行识别,该神经网络模型可以提取基因变异候选位点的序列特征和非序列特征。本公开实施例还提供了一种神经网络模型的结构。
图6示出根据本公开一实施例的神经网络模型的框图。如图6所示,神经网络模型可以包括两个分支结构,第一分支和第二分支。第一分支可以用于提取基因变异候选位点对应的至少一个基因测序读段的序列特征,第一分支可以包括卷积层和池化层。第二分支可以用于提取基因变异候选位点对应的至少一个基因测序读段的非序列特征,第二分支可以包括全连接层。神经网络模型提取基因变异候选位点的序列特征和非序列特征之后,可以将序列特征和非序列特征进行整合,例如,将序列特征的序列特征矩阵与非序列特征的非序列特征矩阵进行拼接,得到整合特征的整合特征矩阵,然后再经过全连接层可以得到基因变异候选位点的变异值。
本公开实施例通过提取基因变异候选位点对应的至少一个基因测序读段的序列属性信息和非序列属性信息,利用对序列属性信息和非序列属性信息整合的整合特征对基因变异进行识别,从而综合考虑基因变异候选位点对应的序列属性信息和非序列属性信息,更加全面地分析基因测序信息,更好地对基因候选位点的基因变异进行识别,筛掉胚系基因变异以及由于噪声和错误带来的干扰,提高基因变异识别的准确率。
本领域技术人员可以理解,在具体实施方式的上述方法中,各步骤的撰写顺序并不意味着严格的执行顺序而对实施过程构成任何限定,各步骤的具体执行顺序应当以其功能和可能的内在逻辑确定。
图7示出根据本公开实施例的基因变异识别装置的框图,如图7所示,所述基因变异识别装置包括:
获取模块71,用于获取基因变异候选位点对应的至少一个基因测序读段;
确定模块72,用于根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;
识别模块73,用于基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述属性信息包括序列属性信息;所述确定模块72,包括:
第一确定子模块,用于根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
第一生成子模块,用于根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
在一种可能的实现方式中,所述第一获取子模块,具体用于确定所述至少一个基因测序读段在所述每个位点的基因类型;统计所述每个位点对应的每种基因类型的基因数量。
在一种可能的实现方式中,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
在一种可能的实现方式中,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
在一种可能的实现方式中,所述序列属性信息包括以下至少一种信息:
参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
在一种可能的实现方式中,所述属性信息包括非序列属性信息;所述确定模块,包括:
第二获取子模块,用于获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
第二确定子模块,用于根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
在一种可能的实现方式中,所述非序列信息包括以下至少一种信息:
对比质量;正负链偏好;基因测序读段长度;边缘偏好。
在一种可能的实现方式中,所述第二确定子模块,具体用于根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述第二确定子模块,具体用于根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;根据所述正负链比例,确定 所述基因变异候选位点对应的非序列特征。
在一种可能的实现方式中,所述识别模块73,包括:
整合子模块,具体用于将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
识别子模块,用于基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
在一种可能的实现方式中,所述识别子模块,具体用于根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
在一种可能的实现方式中,所述获取模块71,具体用于,
获取由体细胞基因进行基因测序得到的基因测序读段;
将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;
根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
获取所述基因变异候选位点对应的至少一个基因测序读段。
在一些实施例中,本公开实施例提供的装置具有的功能或包含的模块可以用于执行上文方法实施例描述的方法,其具体实现可以参照上文方法实施例的描述,为了简洁,这里不再赘述。
图8是根据一示例性实施例示出的一种用于基因变异识别的装置1900的框图。例如,装置1900可以被提供为一服务器。参照图8,装置1900包括处理组件1922,其进一步包括一个或多个处理器,以及由存储器1932所代表的存储器资源,用于存储可由处理组件1922的执行的指令,例如应用程序。存储器1932中存储的应用程序可以包括一个或一个以上的每一个对应于一组指令的模块。此外,处理组件1922被配置为执行指令,以执行上述方法。
装置1900还可以包括一个电源组件1926被配置为执行装置1900的电源管理,一个有线或无线网络接口1950被配置为将装置1900连接到网络,和一个输入输出(I/O)接口1958。装置1900可以操作基于存储在存储器1932的操作系统,例如Windows ServerTM,Mac OS XTM,UnixTM,LinuxTM,FreeBSDTM或类似。
在示例性实施例中,还提供了一种非易失性计算机可读存储介质,例如包括计算机程序指令的存储器1932,上述计算机程序指令可由装置1900的处理组件1922执行以完成上述方法。
本公开可以是系统、方法和/或计算机程序产品。计算机程序产品可以包括计算机可读存储介质,其上载有用于使处理器实现本公开的各个方面的计算机可读程序指令。
计算机可读存储介质可以是可以保持和存储由指令执行设备使用的指令的有形设备。计算机可读存储介质例如可以是――但不限于――电存储设备、磁存储设备、光存储设备、电磁存储设备、半导体存储设备或者上述的任意合适的组合。计算机可读存储介质的更具体的例子(非穷举的列表)包括:便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、静态随机存取存储器(SRAM)、便携式压缩盘只读存储器(CD-ROM)、数字多功能盘(DVD)、记忆棒、软盘、机械编码设备、例如其上存储有指令的打孔卡或凹槽内凸起结构、以及上述的任意合适的组合。这里所使用的计算机可读存储介质不被解释为瞬时信号本身,诸如无线 电波或者其他自由传播的电磁波、通过波导或其他传输媒介传播的电磁波(例如,通过光纤电缆的光脉冲)、或者通过电线传输的电信号。
这里所描述的计算机可读程序指令可以从计算机可读存储介质下载到各个计算/处理设备,或者通过网络、例如因特网、局域网、广域网和/或无线网下载到外部计算机或外部存储设备。网络可以包括铜传输电缆、光纤传输、无线传输、路由器、防火墙、交换机、网关计算机和/或边缘服务器。每个计算/处理设备中的网络适配卡或者网络接口从网络接收计算机可读程序指令,并转发该计算机可读程序指令,以供存储在各个计算/处理设备中的计算机可读存储介质中。
用于执行本公开操作的计算机程序指令可以是汇编指令、指令集架构(ISA)指令、机器指令、机器相关指令、微代码、固件指令、状态设置数据、或者以一种或多种编程语言的任意组合编写的源代码或目标代码,所述编程语言包括面向对象的编程语言—诸如Smalltalk、C++等,以及常规的过程式编程语言—诸如“C”语言或类似的编程语言。计算机可读程序指令可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络—包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。在一些实施例中,通过利用计算机可读程序指令的状态信息来个性化定制电子电路,例如可编程逻辑电路、现场可编程门阵列(FPGA)或可编程逻辑阵列(PLA),该电子电路可以执行计算机可读程序指令,从而实现本公开的各个方面。
这里参照根据本公开实施例的方法、装置(系统)和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其它可编程数据处理装置的处理器,从而生产出一种机器,使得这些指令在通过计算机或其它可编程数据处理装置的处理器执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
也可以把计算机可读程序指令加载到计算机、其它可编程数据处理装置、或其它设备上,使得在计算机、其它可编程数据处理装置或其它设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其它可编程数据处理装置、或其它设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,所述模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以 用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实施例,上述说明是示例性的,并非穷尽性的,并且也不限于所披露的各实施例。在不偏离所说明的各实施例的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实施例的原理、实际应用或对市场中技术的技术改进,或者使本技术领域的其它普通技术人员能理解本文披露的各实施例。

Claims (28)

  1. 一种基因变异识别方法,其特征在于,所述方法包括:
    获取基因变异候选位点对应的至少一个基因测序读段;
    根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;
    基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
  2. 根据权利要求1所述的方法,其特征在于,所述属性信息包括序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征,包括:
    根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
    获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
    根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
  3. 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
    确定所述至少一个基因测序读段在所述每个位点的基因类型;
    统计所述每个位点对应的每种基因类型的基因数量。
  4. 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
    根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;
    统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
  5. 根据权利要求2所述的方法,其特征在于,所述获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息,包括:
    根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;
    统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
  6. 根据权利要求1至5任意一项所述的方法,其特征在于,所述序列属性信息包括以下至少一种信息:
    参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
  7. 根据权利要求1至6任意一项所述的方法,其特征在于,所述属性信息包括非序列属性信息;根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的非序列特征,包括:
    获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
    根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
  8. 根据权利要求7所述的方法,其特征在于,所述非序列信息包括以下至少一种信息:
    对比质量;正负链偏好;基因测序读段长度;边缘偏好。
  9. 根据权利要求8所述的方法,其特征在于,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:
    根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;
    根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
  10. 根据权利要求8所述的方法,其特征在于,所述根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征,包括:
    根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;
    根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
  11. 根据权利要求1-10任意一项所述的方法,其特征在于,所述基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别,包括:
    将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
    基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
  12. 根据权利要求11所述的方法,其特征在于,所述基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别,包括:
    根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;
    在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
  13. 根据权利要求1至12任意一项所述的方法,其特征在于,所述获取基因变异候选位点对应的至少一个基因测序读段,包括:
    获取由体细胞基因进行基因测序得到的基因测序读段;
    将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;
    根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
    获取所述基因变异候选位点对应的至少一个基因测序读段。
  14. 一种基因变异识别装置,其特征在于,所述装置包括:
    获取模块,用于获取基因变异候选位点对应的至少一个基因测序读段;
    确定模块,用于根据所述至少一个基因测序读段的属性信息,确定所述基因变异候选位点的序列特征和非序列特征,其中,所述序列特征为与位点的位置相关的特征;
    识别模块,用于基于所述序列特征和所述非序列特征,对所述基因变异候选位点的基因变异进行识别。
  15. 根据权利要求14所述的装置,其特征在于,所述属性信息包括序列属性信息;所述确定模块,包括:
    第一确定子模块,用于根据所述基因变异候选位点的基因位置信息,确定所述基因变异候选位点所在的预设位点区间;
    第一获取子模块,用于获取所述至少一个基因测序读段在所述预设位点区间中每个位点的序列属性信息;其中,所述序列属性信息为与位点的位置相关的表征基因属性的信息;
    第一生成子模块,用于根据所述预设位点区间中每个位点的序列属性信息,生成所述基因变异候选位点的序列特征。
  16. 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于确定所述至少一个基因测序读段在所述每个位点的基因类型;统计所述每个位点对应的每种基因类型的基因数量。
  17. 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的缺失基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的缺失基因数量。
  18. 根据权利要求15所述的装置,其特征在于,所述第一获取子模块,具体用于根据每个基因测序读段的基因序列与参考基因组的基因序列进行比对的比对结果,确定每个基因测序读段在所述每个位点的插入基因的基因类型;统计所述至少一个基因测序读段在所述每个位点上每种基因类型的插入基因数量。
  19. 根据权利要求14至18任意一项所述的装置,其特征在于,所述序列属性信息包括以下至少一种信息:
    参考基因的基因类型;每种基因类型的基因数量;每种基因类型的缺失基因数量;每种基因类型的插入基因数量。
  20. 根据权利要求14至19任意一项所述的装置,其特征在于,所述属性信息包括非序列属性信息;所述确定模块,包括:
    第二获取子模块,用于获取所述至少一个基因测序读段的非序列属性信息;其中,所述非序列属性信息为与位点的位置不相关的表征基因属性的信息;
    第二确定子模块,用于根据所述至少一个基因测序读段的非序列属性信息,确定所述基因变异候选位点的非序列特征。
  21. 根据权利要求20任意一项所述的装置,其特征在于,所述非序列信息包括以下至少一种信息:
    对比质量;正负链偏好;基因测序读段长度;边缘偏好。
  22. 根据权利要求21所述的装置,其特征在于,所述第二确定子模块,具体用于根据每个基因测序读段中每个位点的对比质量,确定每个基因测序读段的对比质量;其中,所述对比质量用于表征基因测序读段中每个基因序列的基因测序的准确性;根据每个基因测序读段的对比质量,确定所述基因变异候选位点对应的非序列特征。
  23. 根据权利要求21所述的装置,其特征在于,所述第二确定子模块,具体用于根据每个基因测序读段所属基因链的正负链信息,确定所述至少一个基因测序读段所属基因链的正负链比例;根据所述正负链比例,确定所述基因变异候选位点对应的非序列特征。
  24. 根据权利要求14-23任意一项所述的装置,其特征在于,所述识别模块,包括:
    整合子模块,具体用于将所述序列特征和所述非序列特征进行特征整合,得到所述基因变异候选位点的整合特征;
    识别子模块,用于基于所述基因变异候选位点的整合特征,对所述基因变异候选位点的基因变异进行识别。
  25. 根据权利要求24所述的装置,其特征在于,所述识别子模块,具体用于根据所述基因变异候选位点的整合特征,得到所述基因变异候选位点的基因发生变异的变异值;在所述变异值大于或等于预设阈值的情况下,确定所述基因变异候选位点的基因存在变异。
  26. 根据权利要求14至25任意一项所述的装置,其特征在于,所述获取模块,具体用于,
    获取由体细胞基因进行基因测序得到的基因测序读段;
    将所述基因测序读段的基因序列与参考基因组的基因序列进行比对,得到比对结果;
    根据所述比对结果确定所述体细胞基因的基因存在异常的基因变异候选位点;
    获取所述基因变异候选位点对应的至少一个基因测序读段。
  27. 一种基因变异识别装置,其特征在于,包括:
    处理器;
    用于存储处理器可执行指令的存储器;
    其中,其中,所述处理器通过调用所述可执行指令实现如权利要求1至13中任意一项所述的方法。
  28. 一种非易失性计算机可读存储介质,其上存储有计算机程序指令,其特征在于,所述计算机程序指令被处理器执行时实现权利要求1至13中任意一项所述的方法。
PCT/CN2019/089499 2019-03-29 2019-05-31 一种基因变异识别方法、装置和存储介质 Ceased WO2020199336A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
JP2021514554A JP7064654B2 (ja) 2019-03-29 2019-05-31 遺伝子変異認識方法、装置および記憶媒体
SG11202011523VA SG11202011523VA (en) 2019-03-29 2019-05-31 Gene mutation identification method and apparatus, and storage medium
KR1020217020204A KR20210116454A (ko) 2019-03-29 2019-05-31 유전자 변이 인식 방법 및 장치 및 기억 매체
US17/102,136 US20210082539A1 (en) 2019-03-29 2020-11-23 Gene mutation identification method and apparatus, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910251891.0A CN109994155B (zh) 2019-03-29 2019-03-29 一种基因变异识别方法、装置和存储介质
CN201910251891.0 2019-03-29

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/102,136 Continuation US20210082539A1 (en) 2019-03-29 2020-11-23 Gene mutation identification method and apparatus, and storage medium

Publications (1)

Publication Number Publication Date
WO2020199336A1 true WO2020199336A1 (zh) 2020-10-08

Family

ID=67131990

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/089499 Ceased WO2020199336A1 (zh) 2019-03-29 2019-05-31 一种基因变异识别方法、装置和存储介质

Country Status (7)

Country Link
US (1) US20210082539A1 (zh)
JP (1) JP7064654B2 (zh)
KR (1) KR20210116454A (zh)
CN (1) CN109994155B (zh)
SG (1) SG11202011523VA (zh)
TW (1) TWI748263B (zh)
WO (1) WO2020199336A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113628683A (zh) * 2021-08-24 2021-11-09 慧算医疗科技(上海)有限公司 一种高通量测序突变检测方法、设备、装置及可读存储介质
US20240221954A1 (en) * 2021-10-28 2024-07-04 Chengdu Boe Optoelectronics Technology Co., Ltd. Disease prediction methods and devices, electronic devices, and computer readable storage media

Families Citing this family (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111081318B (zh) * 2019-12-06 2023-06-06 人和未来生物科技(长沙)有限公司 一种融合基因检测方法、系统和介质
CN111081314A (zh) * 2019-12-13 2020-04-28 北京市商汤科技开发有限公司 基因变异的识别方法及装置、电子设备和存储介质
CN111091873B (zh) * 2019-12-13 2023-07-18 北京市商汤科技开发有限公司 基因变异的识别方法及装置、电子设备和存储介质
CN111081313A (zh) * 2019-12-13 2020-04-28 北京市商汤科技开发有限公司 基因变异的识别方法及装置、电子设备和存储介质
CN111091867B (zh) * 2019-12-18 2021-11-09 中国科学院大学 基因变异位点筛选方法及系统
CN111304308B (zh) * 2020-03-02 2025-09-16 北京泛生子基因科技有限公司 一种审核高通量测序基因变异检测结果的方法
CN113517022B (zh) * 2021-06-10 2024-06-25 阿里巴巴达摩院(杭州)科技有限公司 基因检测方法、特征提取方法、装置、设备及系统
CN113539357B (zh) * 2021-06-10 2024-04-30 阿里巴巴达摩院(杭州)科技有限公司 基因检测方法、模型训练方法、装置、设备及系统
CN113299344A (zh) * 2021-06-23 2021-08-24 深圳华大医学检验实验室 基因测序分析方法、装置、存储介质和计算机设备
CN115458052B (zh) * 2022-08-16 2023-06-30 珠海横琴铂华医学检验有限公司 基于一代测序的基因突变分析方法、设备和存储介质
CN115620802B (zh) * 2022-09-02 2023-12-05 蔓之研(上海)生物科技有限公司 一种基因数据的处理方法及系统
CN120319304B (zh) * 2025-04-15 2026-02-03 中普康睿河北生物科技有限公司 用于基因多态性快速分型诊断的智能识别方法及系统

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014149134A2 (en) * 2013-03-15 2014-09-25 Guardant Health Inc. Systems and methods to detect rare mutations and copy number variation
CN105989246A (zh) * 2015-01-28 2016-10-05 深圳华大基因研究院 一种基于基因组组装的变异检测方法和装置
WO2016179049A1 (en) * 2015-05-01 2016-11-10 Guardant Health, Inc Diagnostic methods
CN107944228A (zh) * 2017-12-08 2018-04-20 广州漫瑞生物信息技术有限公司 一种基因测序变异位点的可视化方法

Family Cites Families (16)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
ES2704126T3 (es) * 2011-08-23 2019-03-14 Found Medicine Inc Moléculas de fusión KIF5B-RET y usos de las mismas
EP2959011A1 (en) * 2013-02-19 2015-12-30 Cergentis B.V. Sequencing strategies for genomic regions of interest
KR20160010277A (ko) * 2014-07-18 2016-01-27 에스케이텔레콤 주식회사 산모의 무세포 dna의 차세대 서열분석을 통한 태아의 단일유전자 유전변이의 예측방법
CN104293940B (zh) * 2014-09-30 2017-07-28 天津华大基因科技有限公司 构建测序文库的方法及其应用
CN104462869B (zh) * 2014-11-28 2017-12-26 天津诺禾致源生物信息科技有限公司 检测体细胞单核苷酸突变的方法和装置
JP6675164B2 (ja) 2015-07-28 2020-04-01 株式会社理研ジェネシス 変異判定方法、変異判定プログラムおよび記録媒体
JP6679065B2 (ja) 2015-10-07 2020-04-15 国立研究開発法人国立がん研究センター 稀少突然変異の検出方法、検出装置及びコンピュータプログラム
CN105574361B (zh) * 2015-11-05 2018-11-02 上海序康医疗科技有限公司 一种检测基因组拷贝数变异的方法
CN106529211A (zh) * 2016-11-04 2017-03-22 成都鑫云解码科技有限公司 变异位点的获取方法及装置
KR101936933B1 (ko) * 2016-11-29 2019-01-09 연세대학교 산학협력단 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스
CN106611106B (zh) 2016-12-06 2019-05-03 北京荣之联科技股份有限公司 基因变异检测方法及装置
CN106683081B (zh) * 2016-12-17 2020-10-30 复旦大学 基于影像组学的脑胶质瘤分子标记物无损预测方法和预测系统
KR102035615B1 (ko) 2017-08-07 2019-10-23 연세대학교 산학협력단 유전자 패널에 기초한 염기서열의 변이 검출방법 및 이를 이용한 염기서열의 변이 검출 디바이스
CN108021788B (zh) * 2017-12-06 2022-08-05 北京新合睿恩生物医疗科技有限公司 基于细胞游离dna的深度测序数据提取生物标记物的方法和装置
EP3587586A1 (en) * 2018-06-22 2020-01-01 Julius-Maximilians-Universität Würzburg Method for statistically determining a quantification of old and new rna
CN109326316B (zh) * 2018-09-18 2020-10-09 哈尔滨工业大学(深圳) 一种癌症相关SNP、基因、miRNA和蛋白质相互作用的多层网络模型构建方法和应用

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2014149134A2 (en) * 2013-03-15 2014-09-25 Guardant Health Inc. Systems and methods to detect rare mutations and copy number variation
CN105989246A (zh) * 2015-01-28 2016-10-05 深圳华大基因研究院 一种基于基因组组装的变异检测方法和装置
WO2016179049A1 (en) * 2015-05-01 2016-11-10 Guardant Health, Inc Diagnostic methods
CN107944228A (zh) * 2017-12-08 2018-04-20 广州漫瑞生物信息技术有限公司 一种基因测序变异位点的可视化方法

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113628683A (zh) * 2021-08-24 2021-11-09 慧算医疗科技(上海)有限公司 一种高通量测序突变检测方法、设备、装置及可读存储介质
CN113628683B (zh) * 2021-08-24 2024-04-09 慧算医疗科技(上海)有限公司 一种高通量测序突变检测方法、设备、装置及可读存储介质
US20240221954A1 (en) * 2021-10-28 2024-07-04 Chengdu Boe Optoelectronics Technology Co., Ltd. Disease prediction methods and devices, electronic devices, and computer readable storage media

Also Published As

Publication number Publication date
CN109994155B (zh) 2021-08-20
US20210082539A1 (en) 2021-03-18
CN109994155A (zh) 2019-07-09
KR20210116454A (ko) 2021-09-27
TW202036582A (zh) 2020-10-01
SG11202011523VA (en) 2020-12-30
JP7064654B2 (ja) 2022-05-10
JP2022500773A (ja) 2022-01-04
TWI748263B (zh) 2021-12-01

Similar Documents

Publication Publication Date Title
TWI748263B (zh) 一種基因變異辨識方法、裝置和儲存介質
Ono et al. PBSIM2: a simulator for long-read sequencers with a novel generative model of quality scores
Sessegolo et al. Transcriptome profiling of mouse samples using nanopore sequencing of cDNA and RNA molecules
Hu et al. Bayesian detection of convergent rate changes of conserved noncoding elements on phylogenetic trees
Ayling et al. New approaches for metagenome assembly with short reads
Rakocevic et al. Fast and accurate genomic analyses using genome graphs
Girgis Red: an intelligent, rapid, accurate tool for detecting repeats de-novo on the genomic scale
Modolo et al. UrQt: an efficient software for the Unsupervised Quality trimming of NGS data
TWI740262B (zh) 一種基因變異識別方法、裝置和儲存介質
CN111402951B (zh) 拷贝数变异预测方法、装置、计算机设备和存储介质
CN111292802A (zh) 用于检测突变的方法、电子设备和计算机存储介质
Sarmashghi et al. Estimating repeat spectra and genome length from low-coverage genome skims with RESPECT
Wang et al. Tool evaluation for the detection of variably sized indels from next generation whole genome and targeted sequencing data
CN109979530B (zh) 一种基因变异识别方法、装置和存储介质
Ferrario et al. Transferring entropy to the realm of GxG interactions
CN111933214A (zh) 用于检测rna水平体细胞基因变异的方法、计算设备
Huang et al. Reveel: large-scale population genotyping using low-coverage sequencing data
Bonham-Carter et al. Cellular proliferation biases clonal lineage tracing and trajectory inference
Fenton et al. Distinguishing multiple-merger from Kingman coalescence using two-site frequency spectra
US20160026756A1 (en) Method and apparatus for separating quality levels in sequence data and sequencing longer reads
CN110570908B (zh) 测序序列多态识别方法及装置、存储介质、电子设备
HK40007439B (zh) 一种基因变异识别方法、装置和存储介质
HK40007439A (zh) 一种基因变异识别方法、装置和存储介质
Zheng et al. CIGenotyper: A machine learning approach for genotyping complex indel calls
HK40006878B (zh) 一种基因变异识别方法、装置和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19923296

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021514554

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 04.02.2022)

122 Ep: pct application non-entry in european phase

Ref document number: 19923296

Country of ref document: EP

Kind code of ref document: A1