WO2024254824A1 - 群体变异检测方法、装置、电子设备和存储介质 - Google Patents
群体变异检测方法、装置、电子设备和存储介质 Download PDFInfo
- Publication number
- WO2024254824A1 WO2024254824A1 PCT/CN2023/100428 CN2023100428W WO2024254824A1 WO 2024254824 A1 WO2024254824 A1 WO 2024254824A1 CN 2023100428 W CN2023100428 W CN 2023100428W WO 2024254824 A1 WO2024254824 A1 WO 2024254824A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- variation data
- genome
- population
- variation
- data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
Definitions
- the present disclosure relates to the field of biological information processing technology, and in particular to a population variation detection method, device, electronic device and storage medium.
- Population genomics By analyzing DNA sequence polymorphism data, various forces that have acted on the genome are inferred, and the process of biological evolution is explored.
- the research content of population genomics mainly includes changes in population structure and quantity, population mutation, natural selection, genetic drift, etc., as well as the relationship between these and genetic variation in the genome.
- Medical research on population genetics is to explore the frequency of genetic diseases, inheritance mode, gene frequency and change patterns, so as to understand the occurrence and spread of genetic diseases in the population, and provide important information and measures for the prevention, monitoring and treatment of genetic diseases.
- population joint variation detection is a key link in population genomics research.
- how to efficiently perform population joint variation detection on large-scale samples to obtain population variation data corresponding to the samples is an urgent technical issue.
- the present disclosure provides a population variation detection method, device, electronic device and storage medium.
- an embodiment of the present disclosure proposes a population variation detection method, the method comprising: obtaining genome variation data corresponding to each of a plurality of samples; determining a division method of the genome variation data according to population information of the plurality of samples; dividing each of the genome variation data according to the division method to obtain a plurality of variation data segments corresponding to each of the genome variation data, wherein the order of the variation data segments is determined based on the position of the variation data segments in the genome variation data; merging variation data segments with the same order in each of the genome variation data to obtain merged variation data; performing population variation detection on each of the merged variation data to obtain population genome variation data of each of the merged variation data; merging each of the population genome variation data to obtain target population genome variation data of the plurality of samples.
- a population variation detection device comprising: an acquisition module, used to acquire genome variation data corresponding to each of a plurality of samples; a first determination module, used to determine a division method of the genome variation data according to the population information of the plurality of samples; a first division module, used to divide each of the genome variation data according to the division method, so as to obtain a plurality of variation data segments corresponding to each of the genome variation data, wherein the sorting of the variation data segments is based on the order of the variation data segments in the genome variation data.
- the position of each of the genome variation data is determined by the position in the variant data; a first merging processing module is used to merge the variation data fragments with the same order in each of the genome variation data to obtain merged variation data; a first population variation detection module is used to perform population variation detection on each of the merged variation data respectively to obtain the population genome variation data of each of the merged variation data; a second merging processing module is used to merge each of the population genome variation data to obtain the target population genome variation data of the multiple samples.
- Another aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the population variation detection method disclosed in the embodiment of the present disclosure when executing the program.
- Another aspect of the present disclosure is directed to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the population variation detection method disclosed in the embodiment of the present disclosure.
- the division method of the genome variation data is determined according to the population information of the multiple samples, and each genome variation data is divided according to the division method to obtain multiple variation data fragments corresponding to each genome variation data and the variation data fragments with the same order in each genome variation data are merged to obtain merged variation data; population variation detection is performed on each merged variation data to obtain the population genome variation data of each merged variation data, and each population genome variation data is merged to obtain the target population genome variation data of multiple samples. Therefore, in the process of population variation detection, by performing population anomaly detection on multiple local merged data respectively and merging the detection results, the efficiency of obtaining complete population genome variation data of multiple samples can be improved.
- FIG1 is a schematic diagram of a flow chart of a population variation detection method according to an embodiment of the present disclosure
- FIG2 is a schematic diagram of a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- FIG3 is a schematic diagram of a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- FIG4 is a schematic diagram of a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- FIG5 is a schematic diagram of a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- FIG6 is a schematic diagram of a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- FIG7 is a schematic diagram of the structure of a population variation detection device according to an embodiment of the present disclosure.
- FIG8 is a schematic diagram of the structure of a population variation detection device according to another embodiment of the present disclosure.
- FIG. 9 is a block diagram of an electronic device according to an embodiment of the present disclosure.
- sample refers to any sample of a subject, which can reflect a biological state associated with the subject, and it includes cell-free DNA.
- samples include, but are not limited to, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, saliva, sweat, and tears of the subject.
- a biological sample may include any tissue or material derived from a living or dead subject. The sample may be a cell-free sample.
- biological tissue or fluid may be or include amniotic fluid, aqueous humor, ascites, bile, bone marrow, blood, breast milk, cerebrospinal fluid, cerumen, chyle, chyme, ejaculate, endolymph, exudate, feces, gastric acid, gastric juice, lymph, mucus, pericardial fluid, perilymph, peritoneal fluid, pleural fluid, pus, inflammatory secretions, saliva, sebum, semen, serum, smegma, sputum, synovial fluid, sweat, tears, urine, vaginal secretions, vitreous fluid, vomitus, and/or combinations or one or more components thereof.
- biological fluid can be or include intracellular fluid, extracellular fluid, intravascular fluid (plasma), interstitial fluid, lymph and/or transcellular fluid.
- biological fluid can be or include plant exudate.
- biological tissue or sample can be obtained, for example, by suction, biopsy (for example, fine needle or tissue biopsy), swab (for example, oral, nasal, skin or vaginal swab), scraping, surgery, washing or lavage (for example, bronchoalveolar, duct, nose, eye, oral cavity, uterus, vagina or other washing or lavage).
- biological sample is or includes cells obtained from individual.
- sample is "primary sample” obtained directly from a source of interest by any suitable means.
- sample refers to the product obtained by processing the primary sample (for example, by removing one or more components thereof and/or by adding one or more agents thereto). For example, using semipermeable membrane filtration.
- processed samples may include, for example, nucleic acids (including DNA, RNA), proteins or chromatin extracted from the sample or obtained by subjecting the primary sample to one or more techniques (such as amplification or reverse transcription of nucleic acids, separation and/or purification of certain components, etc.).
- Fig. 1 is a flow chart of a population variation detection method according to an embodiment of the present disclosure. It should be noted that the population variation detection method provided in this embodiment is performed by a population variation detection device, wherein the population variation detection device in this embodiment can be implemented by software and/or hardware, and the population variation detection device can be an electronic device, or can be configured in an electronic device.
- the electronic device may include but is not limited to a terminal device, a server, a computer cluster, etc., which is not specifically limited in this embodiment.
- the terminal device may include but is not limited to a personal computer (PC), a mobile device, a tablet computer, etc., and this embodiment does not make any specific limitations on this.
- the population variation detection method may include:
- Step 101 obtaining genome variation data corresponding to each of a plurality of samples.
- variation detection may be performed on multiple samples separately to obtain genome variation detection format (Genome Variant Call Format) GVCF files corresponding to each of the multiple samples, wherein the GVCF files may include genome variation data.
- genome variation detection format Gene Variant Call Format
- Step 102 Determine a division method for genome variation data based on population information of multiple samples.
- the group information is used to indicate the information of the group to which the sample belongs.
- the group information may include but is not limited to group identification information, group name information, etc.
- a possible implementation method for determining the division method of genome variation data according to the population information of multiple samples may be: comparing whether the population information of multiple samples is the same to obtain a comparison result; and determining the division method of genome variation data according to the comparison result. In other words, determining whether the multiple samples are from the same population according to the population information corresponding to each of the multiple samples, and determining the division method of genome variation data according to whether the multiple samples are from the same sample.
- the division method of the genome variation data is determined to be the first division method. That is, after determining that the multiple samples are from the same population according to the population information corresponding to each of the multiple samples, that is, the populations to which the multiple samples belong are the same, and since the multiple samples from the same population have a common origin, the first division method can be used to divide the genome variation data, wherein the first division method is used to indicate a division method based on the density of variation sites.
- the density of variant sites refers to the number of variant sites in a sample of unit length.
- the division method of the genomic variation data is determined to be the second division method. That is, when it is determined based on the population information of the multiple samples that the multiple samples come from different populations, because the multiple samples come from different populations, the population sources are wide, the population heterozygosity is high, and the number of alleles at the same site is large, therefore, it can be determined that the division method of the genomic variation data is the second division method, wherein the second division method is used to indicate the division based on the allele density.
- heterozygosity is also called the average heterozygosity or heterozygosity of a population. It is another indicator of population genetic variation.
- the metric parameter refers to the frequency with which the alleles at a certain locus are heterozygous.
- allele density refers to the number of alleles per unit length of sample.
- alleles refer to genes that control different forms of the same trait at the same position on a pair of homologous chromosomes. Different alleles produce changes in genetic characteristics such as hair color or blood type.
- a multiple allele For example, the human ABO blood type gene locus is at the end of the long arm of chromosome 9. The alleles at this locus are A, B, and O. Therefore, the human ABO blood type is determined by three multiple alleles. When Joint-Calling is performed on the population genome, the number of alleles at the same locus can reach dozens.
- Step 103 dividing each genome variation data according to the division method to obtain a plurality of variation data segments corresponding to each genome variation data, wherein the order of the variation data segments is determined based on the positions of the variation data segments in the genome variation data.
- each genome variation data is divided according to the division method to obtain multiple variation data fragments corresponding to each genome variation data.
- One possible implementation method is: for each genome variation data, the genome variation data is divided according to the variation site density to obtain multiple variation data fragments corresponding to the genome variation data.
- the mutation sites in the genome mutation data are not evenly distributed.
- the distribution of the mutation sites in the genome mutation data is statistically analyzed to obtain the mutation site statistics of the genome mutation data, and according to the mutation site statistics, the genome mutation data is divided based on the mutation site density to obtain multiple mutation data segments corresponding to the genome mutation data, wherein the lengths corresponding to the multiple mutation data segments are different, but the difference in the number of mutation sites on the multiple mutation data segments is within a first preset number, for example, the number of mutation sites on the multiple mutation data segments is basically the same.
- the first preset number may be pre-set.
- the average value of the variant sites assigned to each variant data segment can be determined based on a preset number of segments and the total number of variant sites in the genomic variation data, and the starting and ending positions of each variant data segment in the genomic variation data can be determined based on the average value, and the genomic variation data can be divided based on the starting and ending positions of each variant data segment to obtain a preset number of variant data segments.
- each genome variation data is divided according to the division method to obtain multiple variation data segments corresponding to each genome variation data.
- One possible implementation method is: for each genome variation data, the genome variation data is divided according to the allele density to obtain multiple variation data segments corresponding to the genome variation data.
- the allele information corresponding to each variation site in the genome variation data can be determined, and the allele information in the genome variation data can be counted to obtain the statistical results of the allele information of the genome variation data, and according to the statistical results, the genome variation data is divided based on the allele density to obtain multiple variation data segments corresponding to the genome variation data, wherein the lengths corresponding to the variation data segments are different, and the difference between the numbers of alleles in the multiple variation data segments is within a second preset number.
- the numbers of alleles in the multiple variation data segments can be basically the same.
- Step 104 merging the variation data segments with the same order in each genome variation data to obtain merged variation data.
- Step 105 performing population variation detection on each merged variation data respectively to obtain population genome variation data of each merged variation data.
- multiple idle computing nodes may be called, and each computing node performs population variation detection on a merged variation data, wherein the merged variation data processed by different nodes are different. That is to say, in some examples, population variation detection may be performed on multiple merged variation data respectively through multiple computing nodes, wherein the merged variation data processed by different computing nodes are different.
- population variation detection may be performed on multiple merged variation data respectively through multiple computing nodes, wherein the merged variation data processed by different computing nodes are different.
- a possible implementation method of performing population variation detection on each merged variation data separately to obtain population genome variation data of each merged variation data is as follows: for each computing node: obtaining an unprocessed target merged variation data from multiple merged variation data; calling the computing node to perform population variation detection on the target merged variation data to obtain population genome variation data corresponding to the target merged variation data; obtaining the next unprocessed target merged variation data from the multiple merged variation data until population variation detection is completed for all the multiple merged variation data.
- Step 106 merging the genome variation data of each population to obtain the genome variation data of the target population of multiple samples.
- Step 202 Determine a division method for genome variation data based on population information of multiple samples.
- Step 206 determining the chromosome region corresponding to the genome variation data of each population.
- the population information of multiple samples can be combined to determine the way to divide each genome variation data, and each genome variation data can be divided based on the determined division method, and subsequent data can be obtained based on the divided data.
- the process is exemplarily described below in conjunction with Figure 3.
- Step 301 obtaining genome variation data corresponding to each of a plurality of samples.
- step 301 please refer to the relevant description of the embodiment of the present disclosure, and will not be repeated here.
- Step 302 compare whether the group information of multiple samples is the same to obtain a comparison result.
- step 304 please refer to the relevant description in the embodiment of the present disclosure, and will not be repeated here.
- the second division method is used to indicate a division method based on allele density.
- Step 306 when the division method is the second division method, for each genome variation data, the genome variation data is divided according to the allele density to obtain a plurality of variation data segments corresponding to the genome variation data.
- step 307 please refer to the relevant description of the embodiment of the present disclosure, which will not be repeated here.
- Step 406 performing population variation detection on each merged variation data respectively to obtain the population genome variation data of each merged variation data.
- Step 408 merging the population genome variation data in the same chromosome region to obtain the target population genome variation data of multiple samples.
- FIG5 is a flow chart of a population variation detection method according to another embodiment of the present disclosure.
- the method may further include:
- Step 501 for each genome variation data, the multiple variation data segments of the genome variation data are divided again according to preset lengths to obtain multiple variation data sub-segments corresponding to the genome variation data.
- the order of the variant data sub-segments is determined based on the positions of the variant data sub-segments in the variant data segment to which they belong.
- Step 502 merge the variation data sub-segments with the same order in each genome variation data to obtain a merged variation data segment.
- Step 504 merging the population genome variation data of each merged variation data segment to obtain merged genome variation data corresponding to multiple samples.
- the genomic data is divided based on the division method determined by the population information of multiple samples.
- the multiple variant data segments of the genomic variant data are further divided based on the preset length to obtain variant data sub-segments, and the variant data sub-segments with the same order in each genomic variant data are merged to obtain merged variant data segments, and population variation detection is performed on each merged variant data segment to obtain the population genomic variation data of each merged variant data segment, and the population genomic variation data of each merged variant data segment are merged to obtain the merged genomic variation number corresponding to multiple samples. Therefore, by dividing the genomic data in two division methods and performing subsequent population variation detection based on the divided data, the accuracy of the subsequent population genomic variation data can be improved.
- population anomaly detection is performed on multiple local merged data separately to obtain multiple local population genome anomaly data, and each local population genome anomaly data is merged to obtain complete population genome data, thereby improving the efficiency of obtaining complete population genome data.
- FIG. 7 is a schematic diagram of the structure of a population variation detection device according to an embodiment of the present disclosure.
- the population variation detection device 700 includes: an acquisition module 701, a first determination module 702, a first division module 703, a first merging processing module 704, a first population variation detection module 705, and a second merging processing module 706, wherein:
- a first determination module 702 is used to determine a division method of genome variation data according to population information of multiple samples
- a first partitioning module 703 is used to partition each genome variation data according to a partitioning method to obtain a plurality of variation data segments corresponding to each genome variation data, wherein the order of the variation data segments is determined based on the positions of the variation data segments in the genome variation data;
- a first merging processing module 704 is used to merge the variation data segments with the same order in each genome variation data to obtain merged variation data;
- a first population variation detection module 705 is used to perform population variation detection on each merged variation data to obtain population genome variation data of each merged variation data;
- the second merging processing module 706 is used to merge the genome variation data of each population to obtain the genome variation data of the target population of multiple samples.
- a comparison unit 7021 is used to compare whether the group information of multiple samples is the same to obtain a comparison result
- the determination unit 7022 is specifically used to: when the comparison result is that the population information of multiple samples is the same, determine that the division method of the genome variation data is a first division method; when the comparison result is that the population information of multiple samples is different, determine that the division method of the genome variation data is a second division method, wherein the first division method is used to indicate a division method based on the density of variation sites, and the second division method is used to indicate a division method based on the density of alleles.
- the first division module 703 is specifically used to: for each genome variation data, divide the genome variation data according to the variation site density to obtain multiple variation data segments corresponding to the genome variation data.
- the first division module 703 is specifically used to: for each genome variation data, divide the genome variation data according to the allele density to obtain multiple variation data segments corresponding to the genome variation data.
- the device may further include:
- the second division module 707 is used to divide each genome variation data according to a preset length to obtain a plurality of variant sub-data corresponding to the genome variation data;
- the first partitioning module 703 is specifically used to: partition the multiple variant sub-data according to the partitioning method to obtain multiple variant data segments corresponding to the genome variation.
- the device may further include:
- a third division module 708 is used to divide the multiple variant data segments of the genome variant data again according to preset lengths for each genome variant data, so as to obtain multiple variant data sub-segments corresponding to the genome variant data, wherein the order of the variant data sub-segments is determined based on the position of the variant data sub-segments in the variant data segment to which they belong;
- the population variation detection device of the disclosed embodiment in the process of processing the genome variation data corresponding to each of the multiple samples, determines the division method of the genome variation data according to the population information of the multiple samples, and divides each genome variation data according to the division method to obtain multiple variation data fragments corresponding to each genome variation data, and merges the variation data fragments with the same order in each genome variation data to obtain merged variation data; performs population variation detection on each merged variation data to obtain the population genome variation data of each merged variation data, and merges each population genome variation data to obtain the target population genome variation data of multiple samples. Therefore, in the process of population variation detection, by performing population anomaly detection on multiple local merged data respectively and merging the detection results, the efficiency of obtaining complete population genome variation data of multiple samples can be improved.
- the present disclosure also provides an electronic device and a readable storage medium.
- the processor 920 is used to implement the population variation detection method of the above embodiment when executing the program.
Landscapes
- Bioinformatics & Cheminformatics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Chemical & Material Sciences (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Bioinformatics & Computational Biology (AREA)
- Analytical Chemistry (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
本公开提出了一种群体变异检测方法、装置、电子设备和存储介质,其中,该方法包括:根据多个样本的群体信息,确定基因组变异数据的划分方式,并根据划分方式,划分各个基因组变异数据,以得到各个基因组变异数据各自对应的多个变异数据片段;将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个合并变异数据进行群体变异检测以得到各个合并变异数据的群体基因组变异数据,并对各个群体基因组变异数据进行合并以得到多个样本的目标群体基因组变异数据。由此,在群体变异检测过程中,通过分别对多个局部合并数据进行群体异常检测,并对检测结果进行合并,可提高获取完整的群体基因组变异数据的效率。
Description
本公开涉及生物信息处理技术领域,尤其涉及一种群体变异检测方法、装置、电子设备和存储介质。
群体基因组学:通过分析DNA序列多态数据,推测曾经作用于基因组的各种力量,进而探讨生物演化的过程。群体基因组学的研究内容主要包括群体结构和数量变化、群体突变、自然选择、遗传漂变等,以及这些与基因组中的遗传变异关系。医学研究群体遗传是要探讨遗传病的发病频率,遗传方式及其基因频率和变化的规律,从而了解遗传病在群体中的发生和散布的规律,为预防、监测和治疗遗传病提供重要的信息和措施。
在群体基因组学研究中,群体联合变异检测是群体基因组学研究的一个关键环节。在群体联合变异检测的过程中,如何高效对大规模样本进行群体联合变异检测,以得到样本对应的群体变异数据是目前亟需的技术问题。
发明内容
本公开提出一种群体变异检测方法、装置、电子设备和存储介质。
本公开一方面实施例提出一种群体变异检测方法,所述方法包括:获取多个样本各自对应的基因组变异数据;根据所述多个样本的群体信息,确定所述基因组变异数据的划分方式;根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,其中,所述变异数据片段的排序是基于所述变异数据片段在所述基因组变异数据中的位置确定的;将各个所述基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个所述合并变异数据进行群体变异检测,以得到各个所述合并变异数据各自的群体基因组变异数据;对各个所述群体基因组变异数据进行合并,以得到所述多个样本的目标群体基因组变异数据。
本公开另一方面实施例提出一种群体变异检测装置,所述装置包括:获取模块,用于获取多个样本各自对应的基因组变异数据;第一确定模块,用于根据所述多个样本的群体信息,确定所述基因组变异数据的划分方式;第一划分模块,用于根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,其中,所述变异数据片段的排序是基于所述变异数据片段在所述基因组变
异数据中的位置确定的;第一合并处理模块,用于将各个所述基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;第一群体变异检测模块,用于分别对各个所述合并变异数据进行群体变异检测,以得到各个所述合并变异数据各自的群体基因组变异数据;第二合并处理模块,用于对各个所述群体基因组变异数据进行合并,以得到所述多个样本的目标群体基因组变异数据。
本公开另一方面实施例提出了一种电子设备,包括:存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述程序时实现本公开实施例所公开的群体变异检测方法。
本公开另一方面实施例提出了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现本公开实施例所公开的群体变异检测方法。
根据本公开实施例提供的技术方案,在多个样本各自对应的基因组变异数据进行处理的过程中,根据多个样本的群体信息,确定基因组变异数据的划分方式,并根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段以及将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,以及对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。由此,在群体变异检测的过程中,通过分别对多个局部合并数据进行群体异常检测,并对检测结果进行合并,可提高获取多个样本的完整的群体基因组变异数据的效率。
应当理解,本部分所描述的内容并非旨在标识本公开的实施例的关键或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的说明书而变得容易理解。
附图用于更好地理解本方案,不构成对本公开的限定。其中:
图1是根据本公开一个实施例的群体变异检测方法的流程示意图;
图2是根据本公开另一个实施例的群体变异检测方法的流程示意图;
图3是根据本公开另一个实施例的群体变异检测方法的流程示意图;
图4是根据本公开另一个实施例的群体变异检测方法的流程示意图;
图5是根据本公开另一个实施例的群体变异检测方法的流程示意图;
图6是根据本公开另一个实施例的群体变异检测方法的流程示意图;
图7是根据本公开一个实施例的群体变异检测装置的结构示意图;
图8是根据本公开另一个实施例的群体变异检测装置的结构示意图;
图9是根据本公开一个实施例的电子设备的框图。
下面详细描述本发明的实施例,实施例的示例在附图中示出,其中自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面通过参考附图描述的实施例是示例性的,旨在用于解释本发明,而不能理解为对本发明的限制。
如本文公开的,术语“样本”是指受试者的任何样本,其可反应与受试者相关联的生物状态,并且其包括无细胞DNA。样本的实例包括但不限于受试者的血液、全血、血浆、血清、尿液、脑脊液、唾液、汗液、泪液。生物样本可包括来源于活的或死的受试者的任何组织或材料。样本可为无细胞样本。在一些实施方案中,生物组织或流体可以是或包括羊水、房水、腹水、胆汁、骨髓、血液、母乳、脑脊液、耳垢、乳糜、食糜(chime)、一次射出的精液(ejaculate)、内淋巴、渗出物、粪便、胃酸、胃液、淋巴液、粘液、心包液、外淋巴、腹膜液、胸膜液、脓液、发炎性分泌物、唾液、皮脂、精液、血清、阴垢、痰、滑液、汗液、泪液、尿液、阴道分泌物、玻璃体液、呕吐物,以及/或者它们的组合或一种或多种组分。在一些实施方案中,生物流体可以是或包括细胞内液、细胞外液、血管内液(血浆)、间质液、淋巴液和/或跨细胞液。在一些实施方案中,生物流体可以是或包括植物渗出物。在一些实施方案中,生物组织或样本可以例如通过抽吸、活检(例如,细针或组织活检)、拭子(例如,口腔、鼻、皮肤或阴道拭子)、刮擦、手术、洗涤或灌洗(例如,支气管肺泡、导管、鼻、眼、口腔、子宫、阴道或其他洗涤或灌洗)来获取。在一些实施方案中,生物样本是或包括从个体获得的细胞。在一些实施方案中,样本是通过任何合适的手段直接从感兴趣的来源获得的“初次样本”。在一些实施方案中,如从上下文来看将显而易见的,术语“样本”是指通过处理初次样本(例如,通过去除其一种或多种组分和/或通过向其中添加一种或多种剂)获得的制品。例如,使用半透膜过滤。此类“经处理的样本”可包括例如从样本中提取或通过使初次样本经受一种或多种技术(诸如核酸的扩增或逆转录,某些组分的分离和/或纯化等)而获得的核酸(包括DNA、RNA)、蛋白质或染色质等。
下面参考附图描述本公开实施例的群体变异检测方法、装置、电子设备和存储介质。
图1是根据本公开一个实施例的群体变异检测方法的流程示意图。其中,需要说明的是,本实施例提供的群体变异检测方法由群体变异检测装置执行,其中,本实施例中的群体变异检测装置可以由软件和/或硬件的方式实现,该群体变异检测装置可以为电子设备,或者可以配置在电子设备中。
其中,该电子设备可以包括但不限于终端设备、服务器、计算机集群等,该实施例对此不作具体限定。
其中,终端设备可以包括但不限个人计算机(Personal Computer,PC)、移动设备、平板电脑等,该实施例对此不作具体限定。
如图1所示,该群体变异检测方法可以包括:
步骤101,获取多个样本各自对应的基因组变异数据。
在一些示例性的实施方式中,可通过分别对多个样本进行变异检测,以得到多个样本各自对应的基因组变异检测格式(Genome Variant Call Format)GVCF文件,其中,GVCF文件可以包括基因组变异数据。
步骤102,根据多个样本的群体信息,确定基因组变异数据的划分方式。
其中,群体信息用于表示样本所属于群体的信息,例如,群体信息可以包括但不限于群体标识信息、群体名称信息等。
在一些示例性的实施方式中,根据多个样本的群体信息,确定基因组变异数据的划分方式的一种可能实现方式可以为:比较多个样本的群体信息是否相同,以得到比较结果;根据比较结果,确定基因组变异数据的划分方式。也就是说,根据多个样本各自对应的群体信息来确定多个样本是否来自于同一个群体,并根据多个样本是否来自于同一个样本,确定对基因组变异数据进行划分的划分方式。
作为一种示例中,在比较结果为多个样本的群体信息相同的情况下,确定基因组变异数据的划分方式为第一划分方式。也就是说,在根据多个样本各自对应的群体信息确定多个样本来自同一个群体,即,多个样本所属于的群体是相同的,由于来自同一个群体上的多个样本有着共同的起源,因此,可采用第一划分方式对基因组变异数据进行划分,其中,第一划分方式用于指示基于变异位点密度进行划分的方式。
其中,变异位点密度是指变异位点在单位长度样本上变异位点的数量。
其中,单位长度是根据需求设定的一个参考标准,单位长度就是可供参考的标准,它没有固定值,依设定而变动的。
在另一些示例中,在比较结果为多个样本的群体信息不相同的情况下,确定基因组变异数据的划分方式为第二划分方式。也就是说,在基于多个样本的群体信息确定多个样本来自于不同于群体的情况下,由于多个样本来自于不同群体,群体来源广泛,群体杂合性高,同一个位点等位基因数量较多,因此,可确定对基因组变异数据进行划分的方式为第二划分方式,其中,第二划分方式为用于指示基于等位基因密度进行划分的方式。
其中,杂合性:杂合性又称群体的平均杂合性或杂合度,它是群体遗传变异的另一个
度量参数,是指某一基因座上的等位基因是杂合体的频率。
其中,等位基因密度是指等位基因在单位长度样本上等位基因的数量。
其中,等位基因是指位于一对同源染色体相同位置上控制同一性状不同形态的基因,不同的等位基因产生例如发色或血型等遗传特征的变化。在一个群体内,同源染色体的某个相同座位上的等位基因超过2个以上时,就称作复等位基因。例如,人类ABO血型基因座位是在9号染色体长臂的末端,在这个座位上的等位基因,有A、B、O三个基因,因此人类的ABO血型是由3个复等位基因决定的。群体基因组联合变异检测Joint-Calling时,同一个位点的等位基因数量可达几十个。
步骤103,根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段,其中,变异数据片段的排序是基于变异数据片段在基因组变异数据中的位置确定的。
在一些示例性的实施方式中,在划分方式为第一划分方式的情况下,根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段的一种可能实现方式为:针对每个基因组变异数据,将基因组变异数据按照变异位点密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
在本示例中,基因组变异数据中的变异位点并不是均匀分布的,为了充分利用计算资源和降低计算时间,因此,针对每个基因组变异数据,本示例中对该基因组变异数据中的变异位点的分布情况进行统计,以得到该基因组变异数据的变异位点统计结果,并根据变异位点统计结果,基于变异位点密度对该基因组变异数据进行划分,以得到该基因组变异数据对应的多个变异数据片段,其中,多个变异数据片段所对应的长度是不同的,但是多个变异数据片段上的变异位点数量之间的数量差在第一预设数量之内,例如,多个变异数据片段上的变异位点数量基本一致。
其中,第一预设数量可以是预先设置的。
在一些示例中,可根据预设片段数量和该基因组变异数据中的变异位点的总数,确定出分配给每个变异数据片段的变异位点的平均值,并根据平均值,确定出基因组变异数据中每个变异数据片段的起止位置,并根据每个变异数据片段的起止位置,对该基因组变异数据进行划分,以得到预设片段数量个变异数据片段。
在另一些示例性的实施方式中,在划分方式为第二划分方式的情况下,根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段的一种可能实现方式为:针对每个基因组变异数据,将基因组变异数据按照等位基因密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
具体而言,针对每个基因组变异数据,可确定出该基因组变异数据中每个变异位点上所对应的等位基因信息,并对该基因组变异数据中的等位基因信息进行统计,以得到该基因组变异数据的等位基因信息的统计结果,并根据统计结果,基于等位基因密度对该基因组变异数据进行划分,以得到该基因组变异数据对应的多个变异数据片段,其中,变异数据片段所对应的长度是不同的,并且多个变异数据片段中等位基因的数量之间的差值在第二预设数量内例如,多个变异数据片段中等位基因的数量可以是基本相同的。
在一些示例中,可根据预设片段数量和该基因组变异数据中的等位基因的总数,确定出分配给每个变异数据片段的等位基因的平均值,并根据平均值,确定出基因组变异数据中每个变异数据片段的起止位置,并根据每个变异数据片段的起止位置,对该基因组变异数据进行划分,以得到预设片段数量个变异数据片段。
步骤104,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据。
其中,可以理解的是,本示例中的合并变异数据有多个。其中,合并变异数据的数量与对一个基因组变异数据进行划分所得到的基因变异数据片段的数量是相同的。
步骤105,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据。
在一些示例性的实施方式中,可调用多个处于空闲状态的计算节点,并且每一个计算节点对一个合并变异数据进行群体变异检测,其中,不同节点所处理的合并变异数据是不同的。也就是说,在一些示例中,可通过多个计算节点,分别对多个合并变异数据进行群体变异检测,其中,不同计算节点所处理的合并变异数据是不同的。由此,通过多个计算节点,分别对多个合并变异数据进行群体变异检测,可加速联合变异检测过程,提高获取多个合并变异数据各自对应的群体基因组变异数据的效率。
在另一些示例性的实施方式中,在通过多个计算节点对多个合并变异数据进行群体变异检测的情况下,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据的一种可能实现方式为:针对每个计算节点:从多个合并变异数据中获取一个未处理的目标合并变异数据;调用计算节点对目标合并变异数据进行群体变异检测,以得到目标合并变异数据所对应的群体基因组变异数据;从多个合并变异数据获取下一个未处理的目标合并变异数据,直至多个合并变异数据均完成群体变异检测。
步骤106,对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。
本公开实施例的群体变异检测方法,在多个样本各自对应的基因组变异数据进行处理
的过程中,根据多个样本的群体信息,确定基因组变异数据的划分方式,并根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段以及将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,以及对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。由此,在群体变异检测的过程中,通过分别对多个局部合并数据进行群体异常检测,并对检测结果进行合并,可提高获取多个样本的完整的群体基因组变异数据的效率。
图2是根据本公开另一个实施例的群体变异检测方法的流程示意图。其中,需要说明的是,本实施例是对上述实施例的进一步细化。
如图2所示,该群体变异检测方法,可以包括:
步骤201,获取多个样本各自对应的基因组变异数据。
步骤202,根据多个样本的群体信息,确定基因组变异数据的划分方式。
步骤203,根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段,其中,变异数据片段的排序是基于变异数据片段在基因组变异数据中的位置确定的。
步骤204,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据。
步骤205,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据。
其中,需要说明的是,关于步骤201至步骤205的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
步骤206,确定各个群体基因组变异数据所对应的染色体区域。
步骤207,将处于同一个染色体区域的群体基因组变异数据进行合并处理,以得到多个样本的目标群体基因组变异数据。
其中,目标群体基因组变异数据可以但不限于25个染色区域各自对应的群体基因合并结果。其中,每个染色区域各自对应的群体基因合并结果是通过对对应染色区域所对应的群体基因组变异数据进行合并处理而得到的。
在本示例中,在多个样本各自对应的基因组变异数据进行处理的过程中,根据多个样本的群体信息,确定基因组变异数据的划分方式,并根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段以及将各个基
因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据,并分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,按照染色体区域对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。由此,通过按照染色体区域对局部群体基因组变异数据进行合并,可方便简单地得到多个样本的每个染色区域所对应的全部群体基因组变异数据。
基于上述任意一个实施例的基础上,为了可以提高所得到的群体基因组差异数据的准确性,在一些示例中,可结合多个样本的群体信息来确定对各个基因组变异数据进行划分的方式,并基于所确定出的划分方式分别对各个基因组变异数据进行划分,并基于划分后的数据进行后续出,为了可以清楚理解该过程,下面结合图3对该过程进行示例性描述。
图3是根据本公开另一个实施例的群体变异检测方法的流程示意图。
如图3所示,该群体变异检测方法可以包括:
步骤301,获取多个样本各自对应的基因组变异数据。
其中,关于步骤301的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
步骤302,比较多个样本的群体信息是否相同,以得到比较结果。
步骤303,在比较结果为多个样本的群体信息相同的情况下,确定基因组变异数据的划分方式为第一划分方式。
其中,第一划分方式用于指示基于变异位点密度进行划分的方式。
步骤304,在划分方式为第一划分方式的情况下,针对每个基因组变异数据,将基因组变异数据按照变异位点密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
其中,关于步骤304的具体实现方式,可参见本公开实施例中的相关描述,此处不再赘述。
步骤305,在比较结果为多个样本的群体信息不相同的情况下,确定基因组变异数据的划分方式为第二划分方式。
其中,第二划分方式用于指示基于等位基因密度进行划分的方式。
步骤306,在划分方式为第二划分方式的情况下,针对每个基因组变异数据,将基因组变异数据按照等位基因密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
其中,变异数据片段的排序是基于变异数据片段在基因组变异数据中的位置确定的。
其中,需要说明的是,关于步骤306的具体实现方式,可参考本公开实施例的相关描述,此处不再赘述。
步骤307,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合
并变异数据。
其中,关于步骤307的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
步骤308,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据。
步骤309,对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。
其中,需要说明的是,关于步骤307至步骤309的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
在本示例中,基于多个样本的群体信息来判断多个样本是否来自于同一个群体,并基于判断结果来确定对基因组变异数据的划分方式,并基于所确定出划分方式对各个基因组变异数据进行划分,得到各个基因组变异数据各自对应的多个变异数据片段以及将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据,并分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,按照染色体区域对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。由此,可以提高所得到的群体基因组变异数据的准确性。
在本公开的一个实施例中,在一些场景中,基因组上可能有一段很长距离上没有变异位点的区域,为了方便后续基于染色体区域对各个局部群体基因组变异数据进行合并处理,在采用基于多个样本的群体信息所确定出的划分方式对基因组数据进行划分之前,还可以基于预设划分长度对基因组数据进行划分,为了可以清楚理解本公开,下面结合图4对该实施例的方法进行示例性描述。
图4是根据本公开另一个实施例的群体变异检测方法的流程示意图。
如图4所示,该群体变异检测方法可以包括:
步骤401,获取多个样本各自对应的基因组变异数据。
步骤402,根据多个样本的群体信息,确定基因组变异数据的划分方式。
其中,需要说明的是,关于步骤401和步骤402的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
步骤403,针对每个基因组变异数据,按照预设长度对基因组变异数据进行划分,以得到基因组变异数据对应的多个变异子数据。
其中,预设长度是在群体变异检测装置中预先设置的长度,在实际应用中,可根据实际应用需求来确定该预设长度的取值,该实施例对此不作具体限定。
其中,需要说明的是,步骤402和步骤403的执行顺序不分先后。
步骤404,根据划分方式,对多个变异子数据进行划分,以得到基因组变异所对应的多个变异数据片段。
其中,变异数据片段的排序是基于变异数据片段在基因组变异数据中的位置确定的。
步骤405,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据。
步骤406,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据。
步骤407,确定各个群体基因组变异数据所对应的染色体区域。
步骤408,将处于同一个染色体区域的群体基因组变异数据进行合并处理,以得到多个样本的目标群体基因组变异数据。
其中,需要说明的是,关于步骤405至步骤408的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
在本示例中,基于预设长度对基因组变异数据进行划分,以得到基因组变异数据的多个变异子数据,并基于多个样本的群体信息所确定出的划分方式继续对各个变异子数据进行划分,以得到变异数据片段,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,确定各个群体基因组变异数据所对应的染色体区域,以及将处于同一个染色体区域的群体基因组变异数据进行合并处理,以得到多个样本的目标群体基因组变异数据。由此,通过两种划分方式对基因组数据进行划分,并基于划分后的数据进行后续群体变异检测,可提高后续所得到的目标群体基因组变异数据的准确性。
在本公开的一个实施例中,图5是根据本公开另一个实施例的群体变异检测方法的流程示意图。
如图5所示,该方法还可以包括:
步骤501,针对每个基因组变异数据,按照预设长度,分别对基因组变异数据的多个变异数据片段进行再次划分,以得到基因组变异数据对应的多个变异数据子片段。
其中,变异数据子片段的排序是基于变异数据子片段在其所属于的变异数据片段中的位置确定的。
其中,预设长度是在群体变异检测装置中预先设置的长度,在实际应用中,可根据实际应用需求来确定该预设长度的取值,该实施例对此不作具体限定。
步骤502,将各个基因组变异数据中具有相同排序的变异数据子片段进行合并,以得到合并变异数据片段。
步骤503,分别对各个合并变异数据片段进行群体变异检测,以得到各个合并变异数据片段各自的群体基因组变异数据。
步骤504,对各个合并变异数据片段各自的群体基因组变异数据进行合并处理,以得到多个样本对应的合并基因组变异数据。
在本示例中,基于多个样本的群体信息所确定出的划分方式对基因组数据进行划分,得到基因组变异数据的多个变异数据片段后,继续基于预设长度分别对基因组变异数据的多个变异数据片段进行划分,以得到变异数据子片段,将各个基因组变异数据中具有相同排序的变异数据子片段进行合并,以得到合并变异数据片段,分别对各个合并变异数据片段进行群体变异检测,以得到各个合并变异数据片段各自的群体基因组变异数据,对各个合并变异数据片段各自的群体基因组变异数据进行合并处理,以得到多个样本对应的合并基因组变异数。由此,通过两种划分方式对基因组数据进行划分,并基于划分后的数据进行后续群体变异检测,可提高后续所得到的群体基因组变异数据的准确性。
为了可以清楚了解本公开,下面结合图6,对该实施例的群体变异检测方法进行示例性描述。
图6是根据本公开另一个实施例的群体变异检测方法的流程示意图。
如图6所示,该群体变异检测方法可以包括:
步骤601,分别多个样本各自对应的基因组变异数据进行划分,以得到每个基因组变异数据各自对应的多个变异数据片段。
其中,需要说明的是,关于对基因组变异数据进行划分的划分方式,可基于多个样本的群体信息来确定。
其中,关于基于多个样本的群体信息确定对基因组变异数据进行划分的划分方式的具体描述,可参见本公开实施例的相关描述,此处不再赘述。
步骤602,将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据。
步骤603,分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据。
步骤604,根据各个群体基因组变异数据所对应的染色体区域,将处于同一个染色体区域的群体基因组变异数据进行合并处理,以得到多个样本的对应染色体区域的目标群体基因组变异数据。
其中,需要说明的是,关于步骤602至步骤604的具体实现方式,可参见本公开实施例的相关描述,此处不再赘述。
在本示例中,在群体变异检测的过程中,通过分别对多个局部合并数据进行群体异常检测,以得到多个局部群体基因组异常数据,并对各个局部群体基因组异常数据进行合并处理,以得到完整群体基因组数据,提高了获取完整的群体基因组数据的效率。
在一些示例性的实施方式中,在得到多个样本对应的完整的群体基因组变异数据后,如果又收集到新样本,为了节省大量计算资源和时间消耗,在本示例中,还可以采用本公开的方式对新样本进行处理,以得到新样本对应的群体基因组变异数据,然后,将原先的多个样本对应的群体基因组变异数据和新样本对应的群体基因变异数据进行合并处理,以得到最终的群体的群体基因组变异数据。
与上述几种实施例提供的群体变异检测方法相对应,本公开的一种实施例还提供一种群体变异检测装置,由于本公开实施例提供的群体变异检测装置与上述几种实施例提供的群体变异检测方法相对应,因此,在群体变异检测方法的实施方式也适用于本实施例的群体变异检测装置,在本实施例中不再详细描述。
图7是根据本公开一个实施例的群体变异检测装置的结构示意图。
如图7所示,该群体变异检测装置700包括:获取模块701、第一确定模块702、第一划分模块703、第一合并处理模块704、第一群体变异检测模块705和第二合并处理模块706,其中:
获取模块701,用于获取多个样本各自对应的基因组变异数据;
第一确定模块702,用于根据多个样本的群体信息,确定基因组变异数据的划分方式;
第一划分模块703,用于根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段,其中,变异数据片段的排序是基于变异数据片段在基因组变异数据中的位置确定的;
第一合并处理模块704,用于将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;
第一群体变异检测模块705,用于分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据;
第二合并处理模块706,用于对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。
在本公开的一个实施例中,第二合并处理模块706,具体用于:确定各个群体基因组变异数据所对应的染色体区域;将处于同一个染色体区域的群体基因组变异数据进行合并处
理,以得到多个样本的目标群体基因组变异数据。
在本公开的一个实施例中,在图7所示的装置实施例的基础上,如图8所示,第一确定模块702可以包括:
比较单元7021,用于比较多个样本的群体信息是否相同,以得到比较结果;
确定单元7022,用于根据比较结果,确定基因组变异数据的划分方式。
在本公开的一个实施例中,确定单元7022,具体用于:在比较结果为多个样本的群体信息相同的情况下,确定基因组变异数据的划分方式为第一划分方式;在比较结果为多个样本的群体信息不相同的情况下,确定基因组变异数据的划分方式为第二划分方式,其中,第一划分方式用于指示基于变异位点密度进行划分的方式,第二划分方式用于指示基于等位基因密度进行划分的方式。
在本公开的一个实施例中,在划分方式为第一划分方式的情况下,第一划分模块703,具体用于:针对每个基因组变异数据,将基因组变异数据按照变异位点密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
在本公开的一个实施例中,在划分方式为第二划分方式的情况下,第一划分模块703,具体用于:针对每个基因组变异数据,将基因组变异数据按照等位基因密度进行划分,以得到基因组变异数据对应的多个变异数据片段。
在本公开的一个实施例中,如图8所示,该装置还可以包括:
第二划分模块707,用于针对每个基因组变异数据,按照预设长度对基因组变异数据进行划分,以得到基因组变异数据对应的多个变异子数据;
第一划分模块703,具体用于:根据划分方式,对多个变异子数据进行划分,以得到基因组变异所对应的多个变异数据片段。
在本公开的一个实施例中,如图8所示,该装置还可以包括:
第三划分模块708,用于针对每个基因组变异数据,按照预设长度,分别对基因组变异数据的多个变异数据片段进行再次划分,以得到基因组变异数据对应的多个变异数据子片段,其中,变异数据子片段的排序是基于变异数据子片段在其所属于的变异数据片段中的位置确定的;
第三合并处理模块709,用于将各个基因组变异数据中具有相同排序的变异数据子片段进行合并,以得到合并变异数据片段;
第二群体变异检测模块710,用于分别对各个合并变异数据片段进行群体变异检测,以得到各个合并变异数据片段各自的群体基因组变异数据。
在本公开的一个实施例中,在通过多个计算节点对多个合并变异数据进行群体变异检
测的情况下,第一群体变异检测模块705,具体用于:针对每个计算节点:从多个合并变异数据中获取一个未处理的目标合并变异数据;调用计算节点对目标合并变异数据进行群体变异检测,以得到目标合并变异数据所对应的群体基因组变异数据;从多个合并变异数据获取下一个未处理的目标合并变异数据,直至多个合并变异数据均完成群体变异检测。
本公开实施例的群体变异检测装置,在多个样本各自对应的基因组变异数据进行处理的过程中,根据多个样本的群体信息,确定基因组变异数据的划分方式,并根据划分方式,分别对各个基因组变异数据进行划分,以得到各个基因组变异数据各自对应的多个变异数据片段以及将各个基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个合并变异数据进行群体变异检测,以得到各个合并变异数据各自的群体基因组变异数据,以及对各个群体基因组变异数据进行合并,以得到多个样本的目标群体基因组变异数据。由此,在群体变异检测的过程中,通过分别对多个局部合并数据进行群体异常检测,并对检测结果进行合并,可提高获取多个样本的完整的群体基因组变异数据的效率。
根据本公开的实施例,本公开还提供了一种电子设备和一种可读存储介质。
图9是根据本公开一个实施例的电子设备的结构框图。
如图9所示,该电子设备900包括:存储器910、处理器920及存储在存储器910上并可在处理器920上运行的计算机指令。
处理器920执行指令时实现上述实施例中提供的群体变异检测方法。
进一步地,电子设备900还包括:
通信接口930,用于存储器910和处理器920之间的通信。
存储器910,用于存放可在处理器920上运行的计算机指令。
存储器910可能包含高速RAM存储器,也可能还包括非易失性存储器(non-volatile memory),例如至少一个磁盘存储器。
处理器920,用于执行程序时实现上述实施例的群体变异检测方法。
如果存储器910、处理器920和通信接口930独立实现,则通信接口930、存储器910和处理器920可以通过总线相互连接并完成相互间的通信。总线可以是工业标准体系结构(Industry Standard Architecture,简称为ISA)总线、外部设备互连(Peripheral Component,简称为PCI)总线或扩展工业标准体系结构(Extended Industry Standard Architecture,简称为EISA)总线等。总线可以分为地址总线、数据总线、控制总线等。为便于表示,图9中仅用一条粗线表示,但并不表示仅有一根总线或一种类型的总线。
可选的,在具体实现上,如果存储器910、处理器920及通信接口930,集成在一块芯
片上实现,则存储器910、处理器920及通信接口930可以通过内部接口完成相互间的通信。
处理器920可能是一个中央处理器(Central Processing Unit,简称为CPU),或者是特定集成电路(Application Specific Integrated Circuit,简称为ASIC),或者是被配置成实施本公开实施例的一个或多个集成电路。
本公开另一方面实施例提出了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现本公开实施例任一的群体变异检测方法。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本发明的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不必须针对的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任一个或多个实施例或示例中以合适的方式结合。此外,在不相互矛盾的情况下,本领域的技术人员可以将本说明书中描述的不同实施例或示例以及不同实施例或示例的特征进行结合和组合。
此外,术语“第一”、“第二”仅用于描述目的,而不能理解为指示或暗示相对重要性或者隐含指明所指示的技术特征的数量。由此,限定有“第一”、“第二”的特征可以明示或者隐含地包括至少一个该特征。在本发明的描述中,“多个”的含义是至少两个,例如两个,三个等,除非另有明确具体的限定。
尽管上面已经示出和描述了本发明的实施例,可以理解的是,上述实施例是示例性的,不能理解为对本发明的限制,本领域的普通技术人员在本发明的范围内可以对上述实施例进行变化、修改、替换和变型。
Claims (20)
- 一种群体变异检测方法,其特征在于,所述方法包括:获取多个样本各自对应的基因组变异数据;根据所述多个样本的群体信息,确定所述基因组变异数据的划分方式;根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,其中,所述变异数据片段的排序是基于所述变异数据片段在所述基因组变异数据中的位置确定的;将各个所述基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;分别对各个所述合并变异数据进行群体变异检测,以得到各个所述合并变异数据各自的群体基因组变异数据;对各个所述群体基因组变异数据进行合并,以得到所述多个样本的目标群体基因组变异数据。
- 如权利要求1所述的方法,其特征在于,所述对各个所述群体基因组变异数据进行合并,以得到所述多个样本的目标群体基因组变异数据,包括:确定各个所述群体基因组变异数据所对应的染色体区域;将处于同一个染色体区域的所述群体基因组变异数据进行合并处理,以得到所述多个样本的目标群体基因组变异数据。
- 如权利要求1所述的方法,其特征在于,所述根据所述多个样本的群体信息,确定所述基因组变异数据的划分方式,包括:比较所述多个样本的群体信息是否相同,以得到比较结果;根据所述比较结果,确定所述基因组变异数据的划分方式。
- 如权利要求3所述的方法,其特征在于,所述根据所述比较结果,确定所述基因组变异数据的划分方式,包括:在所述比较结果为所述多个样本的群体信息相同的情况下,确定所述基因组变异数据的划分方式为第一划分方式;在所述比较结果为所述多个样本的群体信息不相同的情况下,确定所述基因组变异数据的划分方式为第二划分方式,其中,所述第一划分方式用于指示基于变异位点密度进行划分的方式,所述第二划分方式用于指示基于等位基因密度进行划分的方式。
- 如权利要求4所述的方法,其特征在于,在所述划分方式为所述第一划分方式的情况 下,所述根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,包括:针对每个所述基因组变异数据,将所述基因组变异数据按照变异位点密度进行划分,以得到所述基因组变异数据对应的多个变异数据片段。
- 如权利要求4所述的方法,其特征在于,在所述划分方式为第二划分方式的情况下,所述根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,包括:针对每个所述基因组变异数据,将所述基因组变异数据按照等位基因密度进行划分,以得到所述基因组变异数据对应的多个变异数据片段。
- 如权利要求2所述的方法,其特征在于,根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段之前,所述方法还包括:针对每个所述基因组变异数据,按照预设长度对所述基因组变异数据进行划分,以得到所述基因组变异数据对应的多个变异子数据;所述根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,包括:根据所述划分方式,对所述多个变异子数据进行划分,以得到所述基因组变异所对应的多个变异数据片段。
- 如权利要求1所述的方法,其特征在于,所述方法还包括:针对每个基因组变异数据,按照预设长度,分别对所述基因组变异数据的多个变异数据片段进行再次划分,以得到所述基因组变异数据对应的多个变异数据子片段,其中,所述变异数据子片段的排序是基于所述变异数据子片段在其所属于的变异数据片段中的位置确定的;将各个所述基因组变异数据中具有相同排序的变异数据子片段进行合并,以得到合并变异数据片段;分别对各个所述合并变异数据片段进行群体变异检测,以得到各个所述合并变异数据片段各自的群体基因组变异数据。
- 如权利要求1-8中任一项所述的方法,其特征在于,在通过多个计算节点对多个所述合并变异数据进行群体变异检测的情况下,所述分别对各个所述合并变异数据进行群体变异检测,以得到各个所述合并变异数据各自的群体基因组变异数据,包括:针对每个计算节点:从多个所述合并变异数据中获取一个未处理的目标合并变异数据;调用所述计算节点对所述目标合并变异数据进行群体变异检测,以得到所述目标合并变异数据所对应的群体基因组变异数据;从多个所述合并变异数据获取下一个未处理的目标合并变异数据,直至多个所述合并变异数据均完成群体变异检测。
- 一种群体变异检测装置,其特征在于,所述装置包括:获取模块,用于获取多个样本各自对应的基因组变异数据;第一确定模块,用于根据所述多个样本的群体信息,确定所述基因组变异数据的划分方式;第一划分模块,用于根据所述划分方式,分别对各个所述基因组变异数据进行划分,以得到各个所述基因组变异数据各自对应的多个变异数据片段,其中,所述变异数据片段的排序是基于所述变异数据片段在所述基因组变异数据中的位置确定的;第一合并处理模块,用于将各个所述基因组变异数据中具有相同排序的变异数据片段进行合并,以得到合并变异数据;第一群体变异检测模块,用于分别对各个所述合并变异数据进行群体变异检测,以得到各个所述合并变异数据各自的群体基因组变异数据;第二合并处理模块,用于对各个所述群体基因组变异数据进行合并,以得到所述多个样本的目标群体基因组变异数据。
- 如权利要求10所述的装置,其特征在于,所述第二合并处理模块,具体用于:确定各个所述群体基因组变异数据所对应的染色体区域;将处于同一个染色体区域的所述群体基因组变异数据进行合并处理,以得到所述多个样本的目标群体基因组变异数据。
- 如权利要求10所述的装置,其特征在于,所述第一确定模块,包括:比较单元,用于比较所述多个样本的群体信息是否相同,以得到比较结果;确定单元,用于根据所述比较结果,确定所述基因组变异数据的划分方式。
- 如权利要求12所述的装置,其特征在于,所述确定单元,具体用于:在所述比较结果为所述多个样本的群体信息相同的情况下,确定所述基因组变异数据的划分方式为第一划分方式;在所述比较结果为所述多个样本的群体信息不相同的情况下,确定所述基因组变异数据的划分方式为第二划分方式,其中,所述第一划分方式用于指示基于变异位点密度进行划分的方式,所述第二划分方式用于指示基于等位基因密度进行划分的方式。
- 如权利要求13所述的装置,其特征在于,在所述划分方式为所述第一划分方式的情况下,所述第一划分模块,具体用于:针对每个所述基因组变异数据,将所述基因组变异数据按照变异位点密度进行划分,以得到所述基因组变异数据对应的多个变异数据片段。
- 如权利要求13所述的装置,其特征在于,在所述划分方式为第二划分方式的情况下,所述第一划分模块,具体用于:针对每个所述基因组变异数据,将所述基因组变异数据按照等位基因密度进行划分,以得到所述基因组变异数据对应的多个变异数据片段。
- 如权利要求10所述的装置,其特征在于,所述装置还包括:第二划分模块,用于针对每个所述基因组变异数据,按照预设长度对所述基因组变异数据进行划分,以得到所述基因组变异数据对应的多个变异子数据;所述第一划分模块,具体用于:根据所述划分方式,对所述多个变异子数据进行划分,以得到所述基因组变异所对应的多个变异数据片段。
- 如权利要求11所述的装置,其特征在于,所述装置还包括:第三划分模块,用于针对每个基因组变异数据,按照预设长度,分别对所述基因组变异数据的多个变异数据片段进行再次划分,以得到所述基因组变异数据对应的多个变异数据子片段,其中,所述变异数据子片段的排序是基于所述变异数据子片段在其所属于的变异数据片段中的位置确定的;第三合并处理模块,用于将各个所述基因组变异数据中具有相同排序的变异数据子片段进行合并,以得到合并变异数据片段;第二群体变异检测模块,用于分别对各个所述合并变异数据片段进行群体变异检测,以得到各个所述合并变异数据片段各自的群体基因组变异数据。
- 如权利要求10-17中任一项所述的装置,其特征在于,在通过多个计算节点对多个所述合并变异数据进行群体变异检测的情况下,所述第一群体变异检测模块,具体用于:针对每个计算节点:从多个所述合并变异数据中获取一个未处理的目标合并变异数据;调用所述计算节点对所述目标合并变异数据进行群体变异检测,以得到所述目标合并变异数据所对应的群体基因组变异数据;从多个所述合并变异数据获取下一个未处理的目标合并变异数据,直至多个所述合并变异数据均完成群体变异检测。
- 一种电子设备,其特征在于,包括:存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,其特征在于,所述处理器执行所述程序时实现如权利要求1-9中任一所述的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,该程序被处理器执行时实现如权利要求1-9中任一所述的方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/100428 WO2024254824A1 (zh) | 2023-06-15 | 2023-06-15 | 群体变异检测方法、装置、电子设备和存储介质 |
| PCT/CN2023/105780 WO2024254925A1 (zh) | 2023-06-15 | 2023-07-04 | 一种对高性能测序数据Joint-Calling加速的方法、装置、电子设备和存储介质 |
| CN202380099345.4A CN121359205A (zh) | 2023-06-15 | 2023-07-04 | 一种对高性能测序数据Joint-Calling加速的方法、装置、电子设备和存储介质 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/100428 WO2024254824A1 (zh) | 2023-06-15 | 2023-06-15 | 群体变异检测方法、装置、电子设备和存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024254824A1 true WO2024254824A1 (zh) | 2024-12-19 |
Family
ID=93851093
Family Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/100428 Ceased WO2024254824A1 (zh) | 2023-06-15 | 2023-06-15 | 群体变异检测方法、装置、电子设备和存储介质 |
| PCT/CN2023/105780 Ceased WO2024254925A1 (zh) | 2023-06-15 | 2023-07-04 | 一种对高性能测序数据Joint-Calling加速的方法、装置、电子设备和存储介质 |
Family Applications After (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/105780 Ceased WO2024254925A1 (zh) | 2023-06-15 | 2023-07-04 | 一种对高性能测序数据Joint-Calling加速的方法、装置、电子设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121359205A (zh) |
| WO (2) | WO2024254824A1 (zh) |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106446254A (zh) * | 2016-10-14 | 2017-02-22 | 北京百度网讯科技有限公司 | 文件检测方法和装置 |
| CN107403076A (zh) * | 2016-05-18 | 2017-11-28 | 华为技术有限公司 | Dna序列的处理方法及设备 |
| CN109698010A (zh) * | 2017-10-23 | 2019-04-30 | 北京哲源科技有限责任公司 | 一种针对基因数据的处理方法 |
| CN110767264A (zh) * | 2019-10-15 | 2020-02-07 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、装置和计算机可读存储介质 |
| WO2022165430A1 (en) * | 2021-02-01 | 2022-08-04 | Google Llc | Structural variant evaluation through iterative genome construction |
| CN115641910A (zh) * | 2022-10-20 | 2023-01-24 | 哈尔滨工业大学 | 一种三代群体基因组结构变异联合检测方法 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DK3511422T3 (da) * | 2013-11-12 | 2023-02-06 | Population Bio Inc | Fremgangsmåder og sammensætninger til diagnosticering, prognose og behandling af endometriose |
| CN112735516A (zh) * | 2020-12-29 | 2021-04-30 | 上海派森诺生物科技股份有限公司 | 一种无参考基因组的群体变异检测分析方法 |
| CN114913918A (zh) * | 2021-02-10 | 2022-08-16 | 中国科学院脑科学与智能技术卓越创新中心 | 一种针对孤独症的高通量测序数据分析方法及装置 |
| CN115331730B (zh) * | 2022-07-20 | 2025-10-17 | 郑州金域临床检验中心有限公司 | 核基因组拷贝数变异检测方法及装置、设备、存储介质 |
| CN115910199B (zh) * | 2022-11-01 | 2023-07-14 | 哈尔滨工业大学 | 一种基于比对框架的三代测序数据结构变异检测方法 |
-
2023
- 2023-06-15 WO PCT/CN2023/100428 patent/WO2024254824A1/zh not_active Ceased
- 2023-07-04 CN CN202380099345.4A patent/CN121359205A/zh active Pending
- 2023-07-04 WO PCT/CN2023/105780 patent/WO2024254925A1/zh not_active Ceased
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107403076A (zh) * | 2016-05-18 | 2017-11-28 | 华为技术有限公司 | Dna序列的处理方法及设备 |
| CN106446254A (zh) * | 2016-10-14 | 2017-02-22 | 北京百度网讯科技有限公司 | 文件检测方法和装置 |
| CN109698010A (zh) * | 2017-10-23 | 2019-04-30 | 北京哲源科技有限责任公司 | 一种针对基因数据的处理方法 |
| CN110767264A (zh) * | 2019-10-15 | 2020-02-07 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、装置和计算机可读存储介质 |
| WO2022165430A1 (en) * | 2021-02-01 | 2022-08-04 | Google Llc | Structural variant evaluation through iterative genome construction |
| CN115641910A (zh) * | 2022-10-20 | 2023-01-24 | 哈尔滨工业大学 | 一种三代群体基因组结构变异联合检测方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024254925A1 (zh) | 2024-12-19 |
| CN121359205A (zh) | 2026-01-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12365933B2 (en) | Systems and methods for epigenetic analysis | |
| US20250215518A1 (en) | Systems and methods for analyzing viral nucleic acids | |
| CN104302781B (zh) | 一种检测染色体结构异常的方法及装置 | |
| CN118824377A (zh) | 基因与疾病关联分析方法、装置、计算机设备及存储介质 | |
| CN107273204B (zh) | 用于基因分析的资源分配方法和装置 | |
| CN108830045A (zh) | 一种基于多组学的生物标记物系统筛选方法 | |
| CN112823391B (zh) | 基于检测限的质量控制度量 | |
| CN115273982A (zh) | 基于转录组测序数据的非编码circRNA生物信息分析方法、装置、终端及介质 | |
| JP7735457B2 (ja) | コピー数バリアントコーラ | |
| CN112687341A (zh) | 一种以断点为中心的染色体结构变异鉴定方法 | |
| WO2024254824A1 (zh) | 群体变异检测方法、装置、电子设备和存储介质 | |
| WO2019242187A1 (zh) | 检测染色体拷贝数异常的方法、装置和存储介质 | |
| CN115148288A (zh) | 一种微生物识别的方法、识别装置及相关设备 | |
| CN116312794B (zh) | 一种融合单细胞分析方法的甲基化样本聚类方法 | |
| CN118471329A (zh) | 基因变异筛查报告方法、系统、电子设备和存储介质 | |
| CN117238368B (zh) | 分子遗传标记分型方法和装置、生物个体识别方法和装置 | |
| CN115206424B (zh) | 一种全同胞关系鉴定方法、系统、设备及存储介质 | |
| CN104182654B (zh) | 基于蛋白‑蛋白相互作用网络的基因集鉴定方法 | |
| CN115273976B (zh) | 一种半同胞关系鉴定方法、系统、设备及存储介质 | |
| CN118098345B (zh) | 一种染色体非整倍体的检测方法、装置、设备及存储介质 | |
| Tenger-Trolander et al. | Genomic resources for the scuttle fly Megaselia abdita: A Model Organism for Comparative Developmental Studies in Flies | |
| CN118588171A (zh) | 基于序列相似度对靶向序列进行准种分析的方法、装置及应用 | |
| WO2024254825A1 (zh) | 一种用于基因检测的加速方法、装置及电子设备 | |
| CN120279993A (zh) | 一种提高预测增强子数量和准确性的生物信息学方法 | |
| Nguyen et al. | Comparison of Galaxy and Unix tools for analyzing the exome sequencing data from syndactyly abnormalities |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23941059 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |