WO2016141516A1 - 获取子代特异性序列、检测子代新突变的方法和装置 - Google Patents
获取子代特异性序列、检测子代新突变的方法和装置 Download PDFInfo
- Publication number
- WO2016141516A1 WO2016141516A1 PCT/CN2015/073816 CN2015073816W WO2016141516A1 WO 2016141516 A1 WO2016141516 A1 WO 2016141516A1 CN 2015073816 W CN2015073816 W CN 2015073816W WO 2016141516 A1 WO2016141516 A1 WO 2016141516A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sub
- progeny
- sequencing result
- reading
- read
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
Definitions
- the present invention relates to the field of biological information. Specifically, the present invention relates to a method for obtaining a progeny-specific sequence, a method for detecting a new mutation in a progeny, and a device for detecting a new mutation in a progeny.
- the present invention provides a method for obtaining a progeny-specific sequence, the method comprising: obtaining a first sequencing result and a second sequencing result, wherein the first sequencing result is a genome sequencing result of the target progeny, The second sequencing result is the genome sequencing result of the father and mother of the offspring, the first sequencing result includes a plurality of reads of length K 1 , and the second sequencing result includes a plurality of reads of length K 2 ; a first sequencing result and a read in the second sequencing result, obtaining a first sub-reading set and a second sub-reading set, the first sub-reading set being composed of a plurality of first sub-reading subsets, a first sub- The subset of reads consists of all sub-reads of length K appearing in one of the first sequencing results, and the second set of sub-reads is composed of a plurality of subsets of second sub-reads, a subset of second sub-reads Compensating for all sub-reads
- Figure 1 illustrates the method concept.
- Obtaining the results of genome sequencing includes sequencing library construction of at least a part of the genome sequence, and sequencing the constructed sequencing library.
- the type and preparation method of the sequencing library can be performed according to the selected sequencing method, and the optional sequencing method is based on Sequencing platforms include, but are not limited to, CG (Complete Genomics) CGA, Illumina/Solexa, Life Technologies/Ion Torrent, and Roche 454, for the preparation of single-ended or multi-end, stranded or circular sequencing libraries based on the selected sequencing platform.
- the father and mother of the so-called offspring refer to the biological parents of the offspring and are the source individuals of the genetic information of the offspring.
- the genome sequencing results of the target progeny and their parents can be obtained simultaneously by co-sequencing, or can be separately sequenced.
- a tag is introduced at the time of library construction to distinguish nucleic acids from progeny and from parents, and the constructed library is mixed on the machine. Sequencing, and then sequencing the sub-generation sequencing results and their parents' sequencing results.
- the daughter sequencing library and its parent sequencing library are separately subjected to double-end sequencing, and the corresponding sequencing results are obtained.
- the two sequencing results include multiple pairs of read pairs, and two pairs of read pairs.
- the reads are from the two ends of an insert, which is usually obtained by interrupting the genome of the target individual or is a free DNA fragment, and the sequencing library is a band suitable for sequencing on a certain sequencing platform. Insert the insert of the linker.
- the read segment corresponding to the sub-read segment refers to the read segment on which the sub-read segment appears.
- the sub-read segment and the read segment usually do not uniquely correspond, that is, generally one sub-read segment corresponds to multiple read segments.
- neither the first sequencing result nor the second sequencing result includes a read that meets any of the following descriptions, including N exceeding 10%, mass value below 10, being contaminated by the joint,
- N refers to the indeterminate base in the read
- the quality value is the value assigned to the read by the sequencing platform
- the quality value is -10*lg(p), where p is the probability of error detection, when a certain position is read
- the probability of error is 0.1
- the mass value is 10
- the junction-contaminated finger refers to the sequencing link included in the reads. Removing these reads will improve the overall quality of the reads contained in the sequencing results, and will also reduce the need for subsequent fast detection analysis of the machine memory.
- Each sub-segment of length K appearing in a read segment indicates that each subsequence in a sequence is present, and the number of subsequences appearing in one sequence is (the length of the sequence - the length of the subsequence +1)
- the 3 bp sub-readings appearing in the 7 bp segment AGAGAGT are AGA, GAG, AGA, GAG, and AGT, with a total of 5 (ie, 7-3+1) 3 bp sub-reads, these 5 sub-reads.
- the segments form a subset of said sub-readers.
- a sub-reading of length K is referred to as a Kmer.
- the first sub-segment set and the second sub-segment set are obtained by using the read sequence in the first sequencing result and the second sequencing result, respectively, including using the first sequencing result All reads, obtaining a subset of all first sub-reading segments, all first subsets of sub-reading segments constitute the first set of sub-reading segments, and all subsets in the second sequencing result are used to obtain all subsets of second sub-reading segments, All of the second subset of sub-reading segments constitute the second set of sub-reading segments.
- the first sub-segment set and the second sub-segment set are obtained by using the read sequence in the first sequencing result and the second sequencing result, respectively, including, from the first sequencing result, the read i Start with the jth nucleotide at one end, copy the K contiguous nucleotides of the read in the other end direction into a Kmer, and take the values in ⁇ 1, 2, ..., m-1, m ⁇ , corresponding to Each i,j takes a value in ⁇ 1,2,...,(K 1 -K), (K 1 -K+1) ⁇ to obtain the first sub-read set, i is the first sequencing
- the number of reads in the result, m is the total number of reads included in the first sequencing result, and the acquisition of the second sub-read set is similar, from the z-th core at the end of the read w in the second sequencing result Start with glycosylation, copy the K contiguous nucleotides of the read in the other end direction to a Kmer,
- the above process of acquiring the sub-reading segment on the machine includes inputting a read segment, and transforming each input into a fixed-length output using a hash algorithm (hashing algorithm), the output being a hash value, that is, a sub-reading segment.
- This conversion is a compression map, where the space of the hash value is usually much smaller than the input (pre-mapped) space.
- Converting the reads to Kmer facilitates rapid comparison and calculation of specific Kmers.
- the memory demand is related to the number and type of kmer. If the kmer is too long, there are many types of kmer. Using machine storage analysis will take up more machine memory, and if Kmer is too short, the number of Kmer will be many.
- the use of machine storage analysis also occupies more machine memory, the number and type of kmer is related to the length of Kmer, that is, the size of K.
- the first sequencing result and the second sequencing result include a read length of not less than 50 bp, that is, K 1 ⁇ 50 bp, K 2 ⁇ 50 bp, and preferably, the sub-reading segment may be set.
- the length is 19 to 31 bp, which makes the requirements for machine memory low, and enables subsequent comparative analysis to be performed quickly.
- the progeny-specific Kmer is filtered prior to extracting the corresponding reads, including filtering out the progeny-specific Kmers having a frequency less than two.
- the frequency of a sub-read in a set refers to the number of sub-reads in the set.
- the following embodiments can also be used to achieve the same effect as the embodiment. After obtaining the first sub-reader set and the second sub-reader set, the Kmer whose frequency is less than 2 in each Kmer set is filtered out, so that the same can be reduced. The resulting false positives for specific reads can also reduce the time required for the next comparative run.
- the method further comprises: after obtaining the progeny-specific sequence, filtering the progeny-specific sequence, including, based on the progeny-specific Kmer frequency, filtering out in the first sub-
- the read segment has a read set corresponding to the progeny-specific sub-reading segment having a frequency less than 2, and/or, filtering out the number of sub-specific sub-reading segments occurring thereon that have fewer than two reads.
- the frequency of a sub-read in a set refers to the number of sub-reads in the set.
- Figure 2 illustrates the flow of obtaining and filtering progeny-specific reads.
- the Kmer count can be directly used by the hash table.
- the hash table is a data structure that exchanges space for time.
- the first sequencing result is built into a hash table.
- the key in the hash table is Kmers, and the stored value is stored.
- the read-out refers to the reading of the number of progeny-specific sub-readings that appear on the sub-readings that are less than 2, and if the number of progeny-specific Kmers present in a reading is less than 2, it is filtered out, for example If a 2 seed-specific Kmer appears on a progeny-specific sequence, regardless of the frequency of each of the two specific Kmers, the progeny-specific sequence is retained; if only one species is present on a progeny-specific sequence Generation-specific Kmer, the progeny-specific Kmer appears no less than 2 times on the progeny-specific sequence, and the progeny-specific sequence is also retained. Filtration of at least one of the above-described progeny-specific sequences facilitates the reduction of obtaining false positive progeny-specific sequences, and is advantageous for obtaining high-quality progeny-specific sequences.
- the storage medium may include a read only memory, a random access memory, a magnetic disk, or an optical disk.
- the present invention provides a method for detecting a new mutation in a progeny, comprising: obtaining a progeny-specific sequence using the method of one or both of the above-described embodiments of the invention; Specific sequence, detecting new mutations in the progeny.
- new generation of mutations refers to mutations that are not inherited from their parents, that is, mutations that are not carried out on their parents' genomes, that is, the genotypes of the offspring at that location are different from their fathers and their mothers, here Father or mother, referring to their biological parents.
- obtaining the progeny-specific sequence comprises filtering out the reads corresponding to the progeny-specific sub-reads having a frequency less than 2 in the first sub-reading set.
- the read corresponding to the progeny-specific sub-reading segment having a frequency less than 2 in the first sub-reading segment is filtered out, and the number of progeny-specific sub-reading segments appearing thereon is also filtered out. Less than 2 reads.
- the frequency of a sub-read in a set refers to the number of sub-reads in the set.
- the detecting a new mutation of the progeny based on a progeny-specific sequence comprises: aligning the progeny-specific sequence with a reference sequence to obtain an alignment result; For the result, the progeny new mutation is detected, and the progeny new mutation includes at least one of SNP, INDEL, and SV. Comparisons can be made using, but not limited to, Samtools, SOAP, BWA, and TeraMap software. In one embodiment of the invention, the alignment is performed using BWA and Samtools in accordance with their default parameters.
- the reference sequence used is a known sequence and may be any reference template in the biological category to which the target individual belongs in advance.
- the reference sequence may select HG19 provided by the National Center for Biotechnology Information (NCBI). Detection of SNP, INDEL or SV may be selected but not limited to using the software SOAPsnp, SOAPindel, GATK and SOAPdenovo.
- an apparatus for detecting a new mutation of a child comprising: a data input unit for inputting data; a data output unit for outputting data; and a storage unit for storing data, including a computer An executable program; a processor coupled to the data input unit, the data output unit, and the storage unit for executing the program, the executing the program comprising completing an aspect or any embodiment of the present invention All or part of the steps of the method of detecting new mutations in the progeny.
- the so-called processor is generally responsible for computing and processing, and a part of the storage unit is memory, mainly for exchanging data.
- the instructions and input data are temporarily stored in memory, transferred to the processor when the processor is idle, and processed by the processor to output the result to an output device such as a display or printer.
- These results are also stored in memory before the output is completed. If the memory is insufficient, the resulting data will be read much slower. Execute the program All or part of the steps of the method of one aspect of the invention can be implemented quickly, and the memory requirements are lower than current methods.
- the detection method/apparatus of the new mutation of the progeny of this aspect of the invention no longer relies on resequencing, but takes a different approach, and uses the Kmer to identify the progeny-specific reads, thereby performing a new generation of denova mutations based on these reads.
- the detection method has a simple principle, which greatly reduces the negative effects of high heterozygosity, high mutation and high repetition area, and the accuracy and sensitivity of the detection result are high.
- FIG. 1 is a schematic view showing the technical principle of a method for detecting a new mutation of a progeny in an embodiment of the present invention.
- FIG. 2 is a flow diagram of the process of acquiring and filtering progeny-specific reads in one embodiment of the present invention.
- FIG. 3 is a schematic diagram showing the frequency distribution of a child Kmer in one embodiment of the present invention.
- FIG. 4 is a schematic diagram showing the frequency distribution of a parent Kmer in one embodiment of the present invention.
- FIG. 5 is a schematic diagram showing the frequency distribution of the progeny-specific Kmer in the progeny Kmer set in one embodiment of the present invention, wherein the three strip columns corresponding to each frequency are 21mer and 25mer from left to right. And 29mer.
- Figure 6 is a graph showing the relationship between the frequency of the progeny-specific Kmer and the detection sensitivity in one embodiment of the present invention.
- Figure 7 is a graphical representation of the relationship between specific filtration conditions and detection sensitivity of progeny-specific reads of the present invention.
- SNPs single base mutations
- INDEL base deletions or insertions
- SV large chromosome structural variations
- the classification of mutations is a relative concept, in the present invention Variation or The mutation, the nucleic acid variation, the genetic variation, the chromosomal variation and the like are common, and the SNP, the indel and the SV in the present invention are generally defined, but the size of the latter two is not particularly limited in the present invention, and thus the different kinds of mutations
- the cross-over of the size of these types of variations does not prevent one of ordinary skill in the art from performing the methods and/or apparatus of the invention described above, and achieve the results described.
- a branch of variation which is different from an inherited mutation, that is, it is inconsistent with the genetic information of the parent, mainly because the fertilized egg changes before the division.
- the DNA fragments are sequenced, and the sequenced objects are all a physically continuous DNA sequence called an insert, the length of which is called the insert size.
- the insert or at least a part of the sequence of the base is read.
- the base sequence on both sides of the insert is read from the edge to the inside, and the reading is measured.
- the sequence is called "reads" and its length can be called read length.
- the ratio of the total Kmer number to the total number of bases is (L - K + 1) / L, and in the above assumption, the ratio is 0.84. Due to the inevitable sequencing errors, there are often many low-frequency Kmers. The total class of these low-frequency Kmers is related to the sequencing error rate.
- the base error rate is 0.1%
- the Kmer size 25
- the genome size is 3G
- the genome sequencing depth is 50X.
- the number of new Kmers due to the wrong base is about 3.5G
- the number of non-repetitive 25Kmers for the human genome is about 2.5G, so that the wrong base causes the memory to increase to 2.4. Times. So in assembly, for example, before building a contig, the low frequency Kmer can be deleted or corrected.
- the reagents, instruments, or software involved in the following examples are conventional commercial products or open source, such as purchasing a sequencing library preparation kit from Illumina, and building a library according to the kit instructions.
- Autism is a serious widespread developmental disorder that occurs in infants and young children. Its main characteristics are social communication disorders, language communication skills, and repeated stereotyped behavior and narrow interest.
- Family studies and twin studies have suggested that the disease has a genetic predisposition. It is a highly genetically related disease with a genetic correlation between 40% and 90%. The agreement between identical twins was 82%, and the agreement between fraternal twins was 10%.
- the types of genomic variants associated with autism have been found to include SNPs, INDELs, and SVs, but the genetic interpretation of autism is still only about 10-20%. With the corresponding research on astronomical genomics and the application of genome-wide sequencing technology, more and more new genetic related sites have been discovered by researchers, and the degree of interpretation of genetic susceptibility of autism can be It is about 50%.
- Illumina Hiseq2000 was used to perform routine database construction and high-depth whole-genome sequencing of the three samples, and double-end sequencing, the obtained parental and parental sequencing data were both 90G (sampling depth of about 30X each).
- the obtained parental and parental sequencing data were both 90G (sampling depth of about 30X each).
- the previous research results of the family Yong-hui Jiang, et al. 2013. Detection of Clinically Relevant Genetic Variants in Autism Spectrum Disorder by Whole-Genome Sequencing
- Denovo mutation was used as the final sensitivity test.
- Jellyfish count/dev/fd/0-m 21-s 3G-Ct 8-Q T--bf-size 10G the meaning of each parameter and the usage of Jellyfish refer to http://www .cbcb.umd.edu/software/jellyfish/.
- the Kmer analysis requirement for a single individual is 50G and the calculation time is 12 hours.
- Our Kmer analysis uses the third-order Kmer lengths—21 bp, 25 bp, and 29 bp—that corresponds to 21 mer, 25 mer, and 29 mer.
- Figures 3 and 4 show the distribution of Kmer frequencies for the offspring and their parents, respectively, of which “21mer-Q "Represents the frequency distribution curve of the 21mer based on the read after the previous step. From the frequency distribution curve of Kmer of the same length, that is, from the "21mer” and “21mer-Q" curves, it can be seen that at lower depths, such as less than 5X, but with higher frequency Kmer, these Kmers are likely due to Sequencing errors include low quality reads.
- Figure 6 shows the relationship between the minimum frequency of the progeny-specific Kmer and the sensitivity, wherein the horizontal scale represents the frequency of the progeny-specific Kmer, the ordinate represents the ratio of false negative (FN), and the lower the false negative FN value, the higher the sensitivity. It can be seen from the figure that the Kmer of three lengths has better sensitivity in the frequency range of 1-2. As the minimum frequency increases, the false negative FN values increase to different degrees, and the sensitivity decreases to some extent.
- each specific reading generally contains a different number of specific Kmers, in order to objectively reflect the sensitivity and accuracy of the method detection, we also use a unique double gradient test, which is based on the specific Kmer frequency and each Specific reads contain the number of specific Kmers, filter specific reads, and then compare the specific reads set under different filter parameters to the reference genome detection denovo mutation, compared with 60 validated variants.
- the sensitivity results are shown in Figure 7.
- Figure 7 shows the relationship between the filter conditions of the progeny-specific reads and the detection sensitivity.
- the filter conditions for the progeny-specific reads include: the frequency of the progeny-specific Kmer (frequency, which is simply expressed as freq) and specific for each progeny.
- Sexual reads contain the number of specific Kmers.
- the method has the following advantages: the principle is simple, and the negative effects of high heterozygosity, high mutation and high repetition area are greatly reduced; the detection sensitivity can reach nearly 100%, that is, the false negative is close to zero.
Landscapes
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Microbiology (AREA)
- Immunology (AREA)
- Physics & Mathematics (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Analytical Chemistry (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
本发明公开一种获取子代特异性序列的方法,包括:获取第一和第二测序结果,第一和第二测序结果分别为子代的和子代的父母的基因组测序结果;分别利用第一和第二测序结果中的读段,获得第一和第二亚读段集,第一和第二亚读段集各自由多个第一和第二亚读段子集构成,一个第一或第二亚读段子集分别由第一或第二测序结果中的一个读段中出现的全部长度为K的亚读段组成;比较第一和第二亚读段集中的亚读段,获得仅包含于第一亚读段集的亚读段,称之为子代特异性亚读段;从第一测序结果中提取子代特异性亚读段对应的读段,提取得的读段为子代特异性序列。本发明还公开检测子代新突变的方法及装置。
Description
本发明涉及生物信息领域,具体的,本发明涉及一种获取子代特异性序列的方法、一种检测子代新突变的方法以及一种检测子代新突变的装置。
人类是二倍体生物,由父亲的精子和母亲的卵子结合发育而来。受精卵的遗传信息绝大多数都是来自于父母,但是不排除极个别DNA序列在受精卵分裂之前就发生了突变,从而在子代中存在既不同于父亲也不同于母亲的DNA突变序列,可称为子代新突变,也就是我们常说的”denovo mutation”(denovo突变)。如果这些序列发生在基因区,尤其是功能蛋白的编码区,那么有可能影响子代的性状和健康状况,比如,自闭症、精神分裂症等。因此,能否精确地、高灵敏度检测denovo mutation对于医学研究,尤其是辅助疾病解读具有非常重大的意义。
高通量测序在近几年内得到极大发展,测序成本急剧下降,无论是欧美发达国家,还是中印等发展中国家都开始大规模的个体全基因组高深度测序获取大量个体基因组数据并分析解读,包括进行疾病研究。目前,检测子代个体denovo突变的思路,都是分别检测出子代及其父母或者其家系成员各自的变异信息,接着再比较判断出只存在于子代的真实变异。目前的这种思路或方法,直接又简单,但是严重依赖于测序获得的读段(reads)的长短及其比对定位时的准确度,对于reads较短,例如只有不到100bp,利用这些reads检测在一些高突变、高杂合或者高重复区域发生的突变,检测精确度一般较低。
发明内容
依据本发明的一方面,本发明提供一种获取子代特异性序列的方法,该方法包括:获取第一测序结果和第二测序结果,第一测序结果为目标子代的基因组测序结果,所述第二
测序结果为该子代的父亲和母亲的基因组测序结果,第一测序结果包括多个长度为K1的读段,第二测序结果包括多个长度为K2的读段;分别利用第一测序结果和第二测序结果中的读段,获得第一亚读段集和第二亚读段集,第一亚读段集由多个第一亚读段子集构成,一个第一亚读段子集由第一测序结果中的一个读段中出现的全部长度为K的亚读段组成,第二亚读段集由多个第二亚读段子集构成,一个第二亚读段子集由第二测序结果中的一个读段中出现的全部长度为K的亚读段组成;将第一亚读段集的亚读段和第二亚读段集中的亚读段进行比较,获得仅包含于第一亚读段集的亚读段,将仅包含于第一亚读段集的亚读段称为子代特异性亚读段;从第一测序结果中提取子代特异性亚读段对应的读段,提取得的读段为所称的子代特异性序列;其中,K<K1,K<K2。图1示意该方法构思。获取基因组测序结果包括对至少一部分基因组序列进行测序文库构建,以及将构建得的测序文库进行上机测序,测序文库的类型和制备方法可根据所选用的测序方法进行,可选用的测序方法根据来自的测序平台包括但不限于CG(Complete Genomics)CGA、Illumina/Solexa、Life Technologies/Ion Torrent和Roche 454,依据所选测序平台进行单端或多端、链状或环状测序文库的制备。所称的子代的父亲和母亲指子代的生物学父母,为子代遗传信息的来源个体。目标子代和其父母的基因组测序结果可以通过共同测序同时获得,也可以分别测序获得,例如,在文库构建时引入标签以区分来自子代和来自父母的核酸,将构建好的文库混合上机测序,再依据标签区分开子代的测序结果和其父母的测序结果。根据本发明的一个实施例,对子代测序文库及其父母测序文库分别进行双末末端测序,获得相应的测序结果,两个测序结果都包含多对读段对,一对读段对的两个读段分别来自一个插入片段的两端,所说的插入片段通常是将来自目标个体的基因组经过打断获得的或者为游离DNA片段,测序文库为适于在某一测序平台上测序的带测序接头的插入片段。所称的亚读段对应的读段指其上出现该亚读段的读段,亚读段与读段通常非唯一对应,即一般一个亚读段对应多个读段。
根据本发明的另一个实施例,所述第一测序结果和所述第二测序结果均不包括符合任一以下描述的读段,含N超过10%、质量值低于10、被接头污染,其中,N指读段中的不确定的碱基,质量值为测序平台赋予读段的值,质量值为-10*lg(p),这里,p为测错的概率,当一条reads某位置出错概率为0.1时,其质量值为10,被接头污染指reads中包含测序接头。去掉这些reads,能使测序结果包含的reads整体质量提高,也利于减少后续快速检测分析对机器内存的需求。
所称一个读段中出现的每个长度为K的亚读段指出现在一个序列中每个亚序列,一个序列中的出现的亚序列的数目为(序列的长度-亚序列的长度+1),例如,长度7bp的读段AGAGAGT中出现的3bp的亚读段为AGA、GAG、AGA、GAG和AGT,共5(即7-3+1)个3bp的亚读段,这5个亚读段即组成一个所说的一个亚读段子集。将长度为K的亚读段称为Kmer。根据本发明的一个实施例,所述分别利用第一测序结果和第二测序结果中的读段,获得第一亚读段集和第二亚读段集,包括,利用第一测序结果中的全部读段,获得全部第一亚读段子集,全部第一亚读段子集构成所述第一亚读段集,利用第二测序结果中的全部读段,获得全部第二亚读段子集,全部第二亚读段子集构成所述第二亚读段集。再具体的,所述分别利用第一测序结果和第二测序结果中的读段,获得第一亚读段集和第二亚读段集,包括,从第一测序结果中的读段i的一端的第j个核苷酸开始,沿另一端方向拷贝该读段的K个连续核苷酸为一个Kmer,i取遍{1,2,…,m-1,m}中的数值,对应每个i,j取遍{1,2,…,(K1-K),(K1-K+1)}中的数值,以获得所述第一亚读段集,i为第一测序结果中的读段的编号,m为第一测序结果包含的读段的总数,第二亚读段集的获取也类似的,从第二测序结果中的读段w的一端的第z个核苷酸开始,沿另一端方向拷贝该读段的K个连续核苷酸为一个Kmer,w取遍{1,2,…,n-1,n}中的数值,对应每个w,z取遍{1,2,…,(K2-K),(K2-K+1)}中的数值,以获得所述第二亚读段集,w为第二测序结果中的读段的编号,n为第二测序结果包含的读段的总数。在机器上进行上述获取亚
读段的过程即包括输入读段,利用哈希算法(散列算法),把每个输入变换成固定长度的输出,该输出为散列值,即亚读段。这种转换是一种压缩映射,即散列值的空间通常远小于输入(预映射)的空间。将读段转换为Kmer,利于快速比较及计算检出特异性Kmer。基于Kmer的分析对内存的需求与kmer的数量和种类有关,如果kmer过长,kmer的种类就很多,利用机器存储分析会占用较多机器内存,而如果Kmer过短,Kmer的数目就很多,利用机器存储分析也会占用较多机器存储内存,kmer的数量和种类与Kmer的长度相关,即与K的大小有关。在本发明的一个实施例中,第一测序结果和第二测序结果包含的读段的长度都不小于50bp,即K1≥50bp、K2≥50bp,较佳的,可设置亚读段的长度为19~31bp,使对机器内存的要求较低,且使后续的比较分析能快速进行。
根据本发明的一个实施例,在提取对应的读段之前,对子代特异性Kmer进行过滤,其中包括,过滤掉频数小于2的子代特异性Kmer。一亚读段在一个集合中的频数指该亚读段在这个集合中的个数。也可以采用以下实施方式达到与该实施例一样的效果,在获得第一亚读段集和第二亚读段集之后,过滤掉各Kmer集中的频数小于2的Kmer,如此,除了能同样降低最终获得的特异性reads的假阳性,还能减少下一步的比较运行所需的时间。
根据本发明的一些实施例,该方法还包括:在获取子代特异性序列之后,对所述子代特异性序列进行过滤,其中包括,基于子代特异性Kmer频数,过滤掉在第一亚读段集中频数小于2的子代特异性亚读段对应的读段,和/或,过滤掉其上出现的子代特异性亚读段的数目少于2的读段。一亚读段在一个集合中的频数指该亚读段在这个集合中的个数。图2示意获取和过滤子代特异性reads的流程。对Kmer的计数可以直接利用哈希表进行,哈希表是一种以空间换取时间的数据结构,将第一测序结果建成一哈希表,哈希表中的关键字即为Kmers,存储值(哈希值)即为Kmer的个数,也可以利用软件Jellyfish[http://www.cbcb.umd.edu/software/jellyfish/]、Tallymer
[http://www.zbh.uni-hamburg.de/?id=211]和Celera中的Meryl模块[Miller,J.R.et al.(2008)Aggressive assembly of pyrosequencing reads with mates.Bioinformatics,24,2818–2824.]等进行。所说的过滤掉其上出现的子代特异性亚读段的数目少于2的读段指,若一读段中出现的子代特异性Kmer的数量少于2则将其过滤掉,例如,若一子代特异性序列上出现2种子代特异性Kmer,不论这两种特异性Kmer各自的频数,该子代特异性序列都得以保留;若一子代特异性序列上只出现一种子代特异性Kmer,该子代特异性Kmer在该子代特异性序列上出现不小于2次,该子代特异性序列也得以保留。对获得的子代特异性序列进行上述至少之一的过滤,有利于减少获得假阳性子代特异性序列,有利于获得高质量的子代特异性序列。
本领域普通技术人员可以理解,上述实施例中各种方法的全部或部分步骤可以通过程序指令相关硬件完成,该程序可以存储于一机器可读存储介质中,例如存储于计算机可读存储介质中,存储介质可以包括:只读存储器、随机存储器、磁盘或光盘等。
依据本发明的另一方面,本发明提供一种检测子代新突变的方法,包括,利用上述本发明一方面或者任一实施例中的方法,获取子代特异性序列;基于所述子代特异性序列,检测所述子代新突变。所称的子代新突变指非遗传自其父母的变异,即其父母基因组上不带有的变异,即子代在该位置的基因型既不同于其父亲,也不同于其母亲,这里的父亲或者母亲,指其生物学意义上的父母。前述对本发明一方面的方法的技术特征和优点的描述,同样适用本发明这一方面的子代新突变的检测方法,在此不再赘述。
根据本发明的一个实施例,获取子代特异性序列包括,过滤掉在第一亚读段集中频数小于2的子代特异性亚读段对应的读段。根据本发明的再一个实施例,过滤掉在第一亚读段集中频数小于2的子代特异性亚读段对应的读段,也过滤掉其上出现的子代特异性亚读段的数目少于2的读段。一亚读段在一个集合中的频数指该亚读段在这个集合中的个数。对获得的子代特异性序列进行上述至少之一的过滤,利于减少获得假阳性子代特异性序列,
利于提高后续基于子代特异性序列的新突变检测的灵敏度。
根据本发明的一个实施例,所述基于子代特异性序列,检测所述子代新突变,包括:将所述子代特异性序列与参考序列比对,获得比对结果;基于所述比对结果,检测所述子代新突变,所述子代新突变包括SNP、INDEL和SV中的至少一种。比对可以选择利用但不限于Samtools、SOAP、BWA和TeraMap软件来进行。在本发明的一个实施例中,比对是利用BWA和Samtools,依照其默认参数进行的。所使用的参考序列是已知序列,可以是预先获得的目标个体所属生物类别中的任意的参考模板。例如,若目标子代是人类,参考序列可选择美国国家生物技术信息中心(NCBI,national center for biotechnology information)提供的HG19。SNP、INDEL或SV的检测可选择但不限于利用软件SOAPsnp、SOAPindel、GATK和SOAPdenovo进行。
本领域普通技术人员可以理解,上述实施例中各种方法的全部或部分步骤可以通过程序来指令相关硬件完成,该程序可以存储于一计算机可读存储介质中,存储介质可以包括:只读存储器、随机存储器、磁盘或光盘等。
依据本发明的再一方面,提供一种检测子代新突变的装置,包括:数据输入单元,用于输入数据;数据输出单元,用于输出数据;存储单元,用于存储数据,其中包括计算机可执行程序;处理器,与所述数据输入单元、所述数据输出单元和所述存储单元连接,用于执行所述程序,执行所述程序包括完成本发明一方面或者任一实施例中的检测子代新突变的方法的全部或部分步骤。前述对本发明一方面的子代新突变的检测方法的技术特征和优点的描述,同样适用本发明这一方面的子代新突变的检测装置,在此不再赘述。所称的处理器一般是负责运算和处理,存储单元中的一部分为内存,主要为交换数据。当程序或者操作者对处理器发出指令,这些指令和输入数据暂存在内存里,在处理器空闲时传送给处理器,处理器处理后把结果输出到诸如显示器、打印机等输出设备上。在输出完成之前,这些结果也保存在内存里,如果内存不足,结果数据读取的速度会变慢很多。执行该程序
能快速实现本发明一方面的方法的全部或部分步骤,对内存的需求比目前的方法低。
本发明这一方面的子代新突变的检测方法/装置,不再依赖重测序,而是另辟蹊径,利用分析Kmer找出子代特异性reads,从而基于这些reads进行子代新突变(denovo mutation)的检测,原理简单,大幅降低了高杂合、高突变以及高重复区域的负面影响,检测结果的准确度和灵敏度都很高。
本发明的上述和/或附加的方面和优点从结合下面附图对实施例的描述中将变得明显和容易理解,其中:
图1是本发明的一个实施例中的检测子代新突变的方法的技术原理示意图。
图2是本发明的一个实施例中的获取和过滤子代特异性reads的流程示意图。
图3是本发明的一个实施例中的子代Kmer的频数分布示意图。
图4是本发明的一个实施例中的父母Kmer的频数分布示意图。
图5是本发明的一个实施例中的子代Kmer集中的子代特异性Kmer的频数分布示意图,其中,在每个频数对应的3个条形柱,从左往右的依次为21mer、25mer和29mer。
图6是本发明的一个实施例中的子代特异性Kmer的频数与检测灵敏度的关系的示意图。
图7是本发明的子代特异性reads的具体过滤条件与检测灵敏度的关系的示意图。
术语解释
变异:
基因组DNA序列上的各种变化,包括单碱基突变(SNP),碱基的缺失或者插入(INDEL)以及大的染色体结构变异(SV),变异的分类是一个相对的概念,在本发明中的变异或者
突变,与核酸变异、基因变异、染色体变异等可通用,本发明中的SNP、插入缺失(indel)和SV同通常定义,但本发明对后两者的大小不作特别限定,这样不同种变异之间有的有交叉,比如SNP为单核酸突变,包括单核苷酸的插入和/或缺失,这样与插入缺失变异有交叉;又比如,当发生的插入缺失较大时,也属于SV。这些类型变异的大小交叉并不妨碍本领域普通技术人员通过上述描述执行实现本发明的方法和/或装置,并且达到所描述的结果。
新突变(Denovo mutation):
变异中的一个分支,它不同于通过遗传获得的突变(Inherited mutation),也就是说,它与父母的遗传信息是不一致的,主要由于受精卵在分裂前发生了变化。
读段(Reads):
对DNA片段进行测序,其测序对象都是一段物理连续的DNA序列,该片段称为插入片段,其长度称为插入片段长度(insert size)。测序过程中,是对该插入片段或者其上至少其一部分序列的碱基进行读取,例如双末端测序中,是对插入片段的两侧碱基序列从边缘向内部进行读取,读测得的序列称为“reads”,其长度可称为读长。
Kmer:
Kmer是具有指定长度为K(例如K=17)的DNA序列,一条长度为L的reads产生的Kmer数量为L-K+1。假定K=17,L=100,则该reads产生84条Kmer。总Kmer数与总碱基数的比率为(L‐K+1)/L,在上述假定中,该比率为0.84。由于测序错误不可避免,往往有很多低频的Kmer存在,这些低频Kmer的总类与测序错误率有关,在碱基错误率为0.1%,Kmer大小为25,基因组大小为3G,基因组测序深度为50X的情况下,由于错误碱基带来的新的Kmer数大约为3.5G个,对于人的基因组来说自身的无重复25Kmer数为2.5G左右,这样错误的碱基造成内存增加为原来的2.4倍。所以在组装中,例如在构建重叠群(Contigs)之前,可对低频Kmer进行删除或者校正。
下面详细描述本发明的实施例,所述实施例中的示例图,其自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面描述的实施例是示例性的,仅用于解释本发明,而不能理解为对本发明的限制。需要说明的是在本文中所使用的术语“第一”、“第二”等仅用于方便描述目的,而不能理解为指示或暗示相对重要性,也不能理解为之间有先后顺序关系。在本发明的描述中,除非另有说明,“多个”的含义是两个或两个以上。
除另有交待,以下实施例中涉及的试剂、仪器或者软件,都是常规市售产品或者开源的,比如从Illumina公司购买测序文库制备试剂盒,依照试剂盒说明书进行建库等。
实施例
从医院获得一个自闭症孩童以及其父母的血液细胞。自闭症是起病于婴幼儿期的一种严重的广泛性发育障碍,其主要特征为社会交往障碍、语言交流技巧障碍,以及重复刻板的行为和狭窄的兴趣等。家系研究、双生子研究均提示该病具有遗传倾向。它是一种与遗传高度相关的疾病,其遗传相关性在40%-90%之间。其同卵双胞胎之间的一致性在82%,而异卵双胞胎之间的一致性为10%。目前已经发现与自闭症相关的基因组变异类型包括SNP、INDEL和SV等,但是自闭症的遗传解释度仍然只有10-20%左右。随着自闭症在基因组学方面的相应研究的进行以及全基因组测序技术的运用,越来越多新的遗传相关位点被研究者所发现,相应自闭症的遗传易感性的解释度能达到50%左右。
一般可采用的步骤如下:
1、利用Illumina Hiseq2000对这三个样本的核酸进行常规建库和高深度全基因组测序,双末端测序,获得子代和父母测序数据均为90G(测序深度各约30X)。同时,基于前人对该家系的研究结果(Yong-hui Jiang,et al.2013.Detection of Clinically Relevant Genetic Variants in Autism Spectrum Disorder by Whole-Genome Sequencing),我们挑选研究结果中60个被验证正确的denovo mutation用作最后的灵敏度测试。
2、过滤原始下机数据,对于含N超过10%的序列、质量值低于10的序列、包含测序引物的序列,均予以删除。根据以下图3显示的,不同大小Kmer的频数分布曲线以及基于不同reads过滤条件的同一Kmer的频数分布曲线,可看出该步骤对检测结果的准确度和灵敏度基本不影响。但该步骤减少了数据量,利于提高机器运行检测方法的速度。
3、将父母测序数据看成是一个整体,用工具jellyfish对子代和父母的过滤完的原始测序数据分别进行Kmer分析。Jellyfish参数如下:“jellyfish count/dev/fd/0-m 21-s 3G-C-t 8-Q T--bf-size 10G”,各参数的意思及Jellyfish的用法参照说明请见http://www.cbcb.umd.edu/software/jellyfish/。单个个体的Kmer分析需求为50G,计算时间12小时。我们的Kmer分析采用三档的Kmer长度——21bp、25bp和29bp,即对应21mer、25mer和29mer,图3和图4分别显示子代和其父母的Kmer频数的分布,其中的“21mer-Q”表示基于上步过滤后的reads的21mer的频数分布曲线。从同样长度的Kmer的频数分布曲线,即从“21mer”和“21mer-Q”曲线可看出,在较低深度,如小于5X,却有较高频率的Kmer的,这些Kmer很可能是由于测序错误包括低质量reads带来的。
4、获取、过滤子代特异性Kmer。在上一步的基础上,过滤频数小于2的Kmer,对子代和父母各自构建Kmer集合,逐个遍历,找出仅存在于子代Kmer集的特异性Kmer。图5展示了子代Kmer集中的子代特异性Kmer(UC-Kmer)的频数分布,对应于每个频数的三个条柱,从左往右依次为21mer、25mer和29mer,从图中可看出,三种长度的Kmer集中的子代特异性Kmer频数绝大多数都分布在频数为1的位置。可靠的UC-Kmer需要一定频数的支持,基于频数过低很有可能由于测序错误导致的,我们这边将频数不大于3的子代特异性Kmer过滤掉。
基于已找出的特异性Kmer,追踪它们对应的reads,也就是特异性reads,对于双末端测序,追踪获得成对的reads。随后,利用这些特异性reads检测子代变异,例如,利
用重测序手段将这些reads比对到参考基因组,所选工具为bwa和samtools,均为默认参数,检测出子代特异性变异(子代新突变),包括SNP、INDEL和SV。
理论上,对子代特异性Kmer的频数进行要求,会对特异性reads的数量产生非常大的影响,所以我们采用梯度测试,即设置不同的Kmer频数过滤阈值,得到不同数目的特异性reads集合,从而获取不同的denovo mutations,它们最终与60个已被验证正确的变异进行比较。图6显示子代特异性Kmer的最小频数与灵敏度的关系,其中,横标代表子代特异性Kmer的频数,纵坐标代表假阴性(FN)的比率,假阴性FN值越低表明灵敏度越高。由图可见,三种长度的Kmer在频数为1-2的阶段有较好的灵敏度,随着最小频数的上升,假阴性FN值有不同程度的上升,进而灵敏度有不同程度的下降。
5、由于每一条特异性reads一般都包含不同数量的特异性Kmer,所以为了客观反映该方法检测的灵敏度和准确性,我们又采用独特的双重梯度测试,即基于特异性Kmer的频数以及每条特异性reads包含的特异性Kmer的数量,对特异性reads进行过滤,然后,将不同过滤参数下的特异性reads集合分别比回参考基因组检测denovo mutation,与60个已被验证正确的变异进行比较,灵敏度结果如图7所示。图7显示了子代特异性reads的过滤条件与检测灵敏度的关系,子代特异性reads的过滤条件包括:子代特异性Kmer的频数(frequency,图上简单表示为freq)和每个子代特异性reads包含的特异性Kmer的数量。
综上所述,可看出该方法具有如下优点:原理简单,大幅降低了高杂合、高突变以及高重复区域的负面影响;检测灵敏度可达到接近100%,即假阴性接近为0。
Claims (12)
- 一种获取子代特异性序列的方法,其特征在于,包括,获取第一测序结果和第二测序结果,所述第一测序结果为所述子代的基因组测序结果,所述第二测序结果为所述子代的父亲和母亲的基因组测序结果,所述第一测序结果包括多个长度为K1的读段,所述第二测序结果包括多个长度为K2的读段;分别利用所述第一测序结果和所述第二测序结果中的读段,获得第一亚读段集和第二亚读段集,所述第一亚读段集由多个第一亚读段子集构成,一个所述第一亚读段子集由第一测序结果中的一条读段中出现的每条长度为K的亚读段组成,所述第二亚读段集由多个第二亚读段子集构成,一个所述第二亚读段子集由第二测序结果中的一条读段中出现的每条长度为K的亚读段组成;将所述第一亚读段集的亚读段和所述第二亚读段集中的亚读段进行比较,获得仅包含于第一亚读段集的亚读段,将所述仅包含于第一亚读段集的亚读段称为子代特异性亚读段;从所述第一测序结果中提取所述子代特异性亚读段对应的读段,所述提取得的读段为所述子代特异性序列;其中,K<K1,K<K2。
- 权利要求1的方法,其特征在于,所述分别利用第一测序结果和第二测序结果中的读段,获得第一亚读段集和第二亚读段集,包括,利用第一测序结果中的全部读段,获得全部第一亚读段子集,全部第一亚读段子集构成所述第一亚读段集,利用第二测序结果中的全部读段,获得全部第二亚读段子集,全部第二亚读段子集构成所述第二亚读段集。
- 权利要求2的方法,其特征在于,所述分别利用第一测序结果和第二测序结果中的读段,获得第一亚读段集和第二亚读段集,包括,从第一测序结果中的读段i的一端的第j个核苷酸开始,沿另一端方向拷贝该读段的K个连续核苷酸为一个亚读段,i取遍{1,2,…,m-1,m}中的数值,对于每个i,j取遍{1,2,…,(K1-K),(K1-K+1)}中的数值,以获得所述第一亚读段集,i为第一测序结果中的读段的编号,m为第一测序结果包含的读段的总数,从第二测序结果中的读段w的一端的第z个核苷酸开始,沿另一端方向拷贝该读段的K个连续核苷酸为一个亚读段,w取遍{1,2,…,n-1,n}中的数值,对于每个w,z取遍{1,2,…,(K2-K),(K2-K+1)}中的数值,以获得所述第二亚读段集,w为第二测序结果中的读段的编号,n为第二测序结果包含的读段的总数。
- 权利要求1-3任一方法,其特征在于,在将所述第一亚读段集的亚读段和所述第二亚读段集中的亚读段进行比较之前,分别过滤掉第一亚读段集和第二亚读段集中的频数小于2的亚读段。
- 权利要求1-3任一方法,其特征在于,在获得子代特异性亚读段之后,过滤掉频数小于2的子代特异性亚读段。
- 权利要求1的方法,其特征在于,还包括,对所述子代特异性序列进行过滤,其中包括,过滤掉在第一亚读段集中频数小于2的子代特异性亚读段对应的读段,和/或,过滤掉其上出现的子代特异性亚读段的数目少于2的读段。
- 权利要求1-6任一方法,其特征在于,所述第一测序结果和所述第二测序结果均不包括符合以下描述的读段,合N超过10%,或质量值低于10,或被接头污染。
- 权利要求1-7任一方法,其特征在于,在K1≥50bp、K2≥50bp时,19bp≤K≤ 31bp。
- 一种检测子代新突变的方法,其特征在于,包括,利用权利要求1-8任一方法,获取子代特异性序列;基于所述子代特异性序列,检测所述子代新突变。
- 权利要求9的方法,其特征在于,所述基于子代特异性序列,检测所述子代新突变,包括,将所述子代特异性序列与参考序列比对,获得比对结果;基于所述比对结果,检测所述子代新突变,所述子代新突变包括SNP、INDEL和SV中的至少一种。
- 一种检测子代新突变的装置,其特征在于,包括,数据输入单元,用于输入数据;数据输出单元,用于输出数据;存储单元,用于存储数据,其中包括计算机可执行程序;处理器,与所述数据输入单元、所述数据输出单元和所述存储单元连接,用于执行所述程序,执行所述程序包括完成权利要求9或10的方法。
- 一种计算机可读介质,其特征在于,用于存储供计算机执行的程序,执行所述程序包括完成权利要求9或10的方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2015/073816 WO2016141516A1 (zh) | 2015-03-06 | 2015-03-06 | 获取子代特异性序列、检测子代新突变的方法和装置 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2015/073816 WO2016141516A1 (zh) | 2015-03-06 | 2015-03-06 | 获取子代特异性序列、检测子代新突变的方法和装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2016141516A1 true WO2016141516A1 (zh) | 2016-09-15 |
Family
ID=56878532
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2015/073816 Ceased WO2016141516A1 (zh) | 2015-03-06 | 2015-03-06 | 获取子代特异性序列、检测子代新突变的方法和装置 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2016141516A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113628681A (zh) * | 2021-07-21 | 2021-11-09 | 哈尔滨星云医学检验所有限公司 | 一种基于家系denovo突变的分析方法及其应用 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013177581A2 (en) * | 2012-05-24 | 2013-11-28 | University Of Washington Through Its Center For Commercialization | Whole genome sequencing of a human fetus |
| WO2014019267A1 (en) * | 2012-08-01 | 2014-02-06 | Bgi Shenzhen | Method and system to determine biomarkers related to abnormal condition |
| CN103824001A (zh) * | 2014-02-27 | 2014-05-28 | 北京诺禾致源生物信息科技有限公司 | 染色体的检测方法和装置 |
| CN104160391A (zh) * | 2011-09-16 | 2014-11-19 | 考利达基因组股份有限公司 | 确定异质样本的基因组中的变异 |
-
2015
- 2015-03-06 WO PCT/CN2015/073816 patent/WO2016141516A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104160391A (zh) * | 2011-09-16 | 2014-11-19 | 考利达基因组股份有限公司 | 确定异质样本的基因组中的变异 |
| WO2013177581A2 (en) * | 2012-05-24 | 2013-11-28 | University Of Washington Through Its Center For Commercialization | Whole genome sequencing of a human fetus |
| WO2014019267A1 (en) * | 2012-08-01 | 2014-02-06 | Bgi Shenzhen | Method and system to determine biomarkers related to abnormal condition |
| CN103824001A (zh) * | 2014-02-27 | 2014-05-28 | 北京诺禾致源生物信息科技有限公司 | 染色体的检测方法和装置 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113628681A (zh) * | 2021-07-21 | 2021-11-09 | 哈尔滨星云医学检验所有限公司 | 一种基于家系denovo突变的分析方法及其应用 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7684708B2 (ja) | 母体血漿の無侵襲的出生前分子核型分析 | |
| Zhou et al. | Prevention, diagnosis and treatment of high‐throughput sequencing data pathologies | |
| WO2015051006A2 (en) | Phasing and linking processes to identify variations in a genome | |
| CN115989544A (zh) | 用于在基因组的重复区域中可视化短读段的方法和系统 | |
| US20160154930A1 (en) | Methods for identification of individuals | |
| WO2015043278A1 (zh) | 同时进行单体型分析和染色体非整倍性检测的方法和系统 | |
| CN115631790A (zh) | 单细胞转录组测序数据的体细胞突变提取方法及装置 | |
| WO2013097048A1 (zh) | 基因组单核苷酸多态性位点的标记方法和装置 | |
| Bonfiglio et al. | Best practices for germline variant and DNA methylation analysis of second-and third-generation sequencing data | |
| CN117238365A (zh) | 基于高通量测序技术的新生儿遗传病早筛方法及装置 | |
| WO2025167582A1 (zh) | 判断样本污染的方法、装置、电子设备和存储设备 | |
| CN114913918A (zh) | 一种针对孤独症的高通量测序数据分析方法及装置 | |
| WO2016141516A1 (zh) | 获取子代特异性序列、检测子代新突变的方法和装置 | |
| D’Agaro | New advances in NGS technologies | |
| Veeramachaneni | Data Analysis in Rare Disease Diagnostics: CV Veeramachaneni | |
| Harris et al. | Whole-genome sequencing for rapid and accurate identification of bacterial transmission pathways | |
| Hu et al. | panHiTE: a comprehensive and accurate pipeline for TE detection in large-scale population genomes | |
| Smith et al. | Considerations of depth, coverage, and other read quality metrics | |
| US20240412808A1 (en) | Detection of cystic fibrosis transmembrane conductance regulator polytg/polyt variations by an ngs-based method | |
| KR20250092241A (ko) | 핵산 오류 억제 | |
| CN119091955A (zh) | 一种获得基因组变异数据集的方法、装置和存储介质 | |
| HK40080479A (zh) | 母体血浆的无创性产前分子染色体核型分析 | |
| HK40074981A (zh) | 母体血浆的无创性产前分子染色体核型分析 | |
| Hedges | Bioinformatics of Human Genetic Disease Studies | |
| Pu et al. | Benchmarking UMI clustering tools for accurate detection of low-frequency variants from deep sequencing |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 15884199 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 15884199 Country of ref document: EP Kind code of ref document: A1 |