WO2016078067A1 - 个体单核苷酸多态性位点分型方法及装置 - Google Patents
个体单核苷酸多态性位点分型方法及装置 Download PDFInfo
- Publication number
- WO2016078067A1 WO2016078067A1 PCT/CN2014/091824 CN2014091824W WO2016078067A1 WO 2016078067 A1 WO2016078067 A1 WO 2016078067A1 CN 2014091824 W CN2014091824 W CN 2014091824W WO 2016078067 A1 WO2016078067 A1 WO 2016078067A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- site
- base
- ratio
- latitude
- tested
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
Definitions
- the invention relates to the technical field of genomics and bioinformatics, and particularly relates to a single Nucleotide Polymorphisms (SNP) site typing method and device.
- SNP Nucleotide Polymorphisms
- determining the genotype of an individual at a certain locus ie, single nucleotide polymorphism locus typing, referred to as SNP typing
- SNP typing single nucleotide polymorphism locus typing
- the invention provides a method and a device for individual single nucleotide polymorphism locus typing, which can easily and efficiently count SNP typing results.
- an embodiment provides a method for individual single nucleotide polymorphism locus typing, comprising:
- the base information of the site to be tested of the segment sequence; the genotype of the site is typed based on the extracted base information of the site to be tested.
- an embodiment provides an individual single nucleotide polymorphism site typing device, comprising: an acquisition module, a comparison module, an extraction module, and a parting module, wherein the acquisition module is used Obtaining a sequence of reads containing SNP site information; the comparison module is configured to compare the data of the outer linker in the target library containing the sequence of the amplification target, and obtain the sequence of the read sequence after the comparison; The module is configured to extract base information of all aligned read sequence sites; the typing module is configured to classify the locus genotype based on the extracted base information of the detected site.
- the individual single nucleotide polymorphism locus typing method/device of the present invention since the genotype of the locus is typed based on the more specific base information of the extracted locus, rather than the root According to the intrinsic characteristics of the sequence and the statistical probability model, the individual fixed-point typing can be realized simply, effectively and accurately. In particular, when the base information is the base type and its number, the accuracy of the typing is further improved.
- FIG. 1 is a schematic structural diagram of an individual single nucleotide polymorphism site typing device disclosed in an embodiment of the present application
- FIG. 2 is a flow chart of a single single nucleotide polymorphism locus typing method disclosed in an embodiment of the present application
- FIG. 3 is a flow chart of a base information classification based on a site to be tested according to an embodiment of the present application
- FIG. 4 is a schematic diagram of evaluation of the effect of sequencing and typing different depth data based on Proton in the embodiment of the present application.
- the present invention contemplates that the number of bases is counted on a certain SNP site (hereinafter referred to as a fixed point) in a read sequence set containing SNP site information, and then an appropriate threshold is selected, which is implemented according to the selected threshold. A higher accuracy individual SNP locus typing result is achieved.
- the typing method is suitable for all related items of various individual fixed-point SNP detection, such as detection of individual disease sites, identification of individuals, and detection of individual functional genes.
- the main base the largest number of bases at the SNP site of the individual read sequence set; the number of second bases is second; the number of third bases is lower than the number of second bases.
- the latitude is the ratio of the main base to the second base, and the value is not less than 1, and the smaller the latitude, the higher the severity of the heterozygous determination.
- those skilled in the art can evolve and deduct according to the concept, such as the ratio of the second base to the main base, such as the difference between the primary base and the second base.
- the ratio of the main base or the second base, etc. should be considered as equivalent to the latitude equivalent of the concept of the present embodiment.
- sequencing platforms including Roche454, Ion PGM and Ion Proton.
- the embodiments of the present invention are illustrated by the second generation Ion Proton sequencing platform, and other sequencing platforms are equally applicable to the methods provided by the present invention, and the sequencing platform does not constitute a limitation of the present invention. If the specific conditions are not specified in the examples, they are carried out according to the general conditions or the conditions recommended by the manufacturer; the reagents or instruments used are not indicated by the manufacturer, and are conventional products that can be obtained through market purchase.
- FIG. 1 is an apparatus for classifying individual single nucleotide polymorphism sites (SNP sites) according to the embodiment, comprising: acquiring module 1, comparing module 2, extracting module 3, and typing module. 4, among them,
- the acquisition module 1 is configured to obtain a read sequence containing SNP site information
- the comparison module 2 is configured to compare the data of the outer link removed in the target library containing the amplification target sequence, and obtain the read after the comparison.
- a segment sequence set
- the extraction module 3 is configured to extract base information of the read sequence of the read sequence on the alignment
- the typing module 4 is configured to perform the genotype on the site based on the extracted base information of the site to be tested Classification.
- the base information includes: a base type and its number.
- the parting module 4 includes a base ratio calculating unit 41, a ratio determining unit 42, a base number counting unit 43, and a genotyping unit 44.
- the base ratio calculating unit 41 is configured to calculate a bit to be measured.
- the ratio determining unit 42 is configured to determine the ratio of the ratio and the preset value
- the base number counting unit 43 is configured to obtain the The number of the second base and the third base
- the genotyping unit 44 is used to type the genotype of the site to be tested, and if the ratio of the second base to the third base exceeds a preset ratio, it is determined that the The measuring point is the first latitude heterozygous type; otherwise it is the second latitude heterozygous type
- the latitude is a numerical value (scalar) capable of characterizing the severity of the heterozygous determination condition, and the first latitude is greater than the second latitude If the ratio is less than the preset value, it is determined that the to-be-measured site is a third latitude heterozygous type; the third latitude is less than the second latitude.
- the individual SNP site segmentation device further includes: a threshold adjustment module 5, wherein the threshold adjustment module 5 is configured to obtain the supported number of the sequenced read segments to be tested, if the obtained support number is greater than a preset For the number of supports, the values of the preset ratio, the preset value, the first latitude, the second latitude, and the third latitude are adjusted based on the obtained support number.
- this embodiment also discloses an individual single nucleotide polymorphism site typing method, please refer to FIG. 2, which is a flow chart of the method. The steps are as follows:
- Step S100 the read sequence is acquired. Get the read sequence reads containing SNP site information.
- the SNP site is a site that meets the minimum allele frequency MAF > 0.3 and can be amplified by designing primers; and the SNP sites are in harmony with the Hardy-Weinberg equilibrium.
- Step S200 reading the sequence alignment.
- the data of the removed outer linker is aligned in the target library containing the amplification target sequence, and the aligned read sequence set is obtained.
- the acquired read sequence reads are subjected to PCR primer-based data statistics, and then the aligned sequence data sets are obtained in the target sequence set using the comparison software.
- the validity of the data may be statistically compared according to the comparison between the read sequence reads and the PCR primers of each site, including the site corresponding to each read, the length and number of reads corresponding to each site, and the number of effective reads. And percentage.
- Tmap a comparison software
- other alignment schemes may be employed.
- Step S300 base information extraction. Extract all aligned read sequence sites to be tested Base information of (SNP site).
- the sequence of 6-10 bp in front of the site to be tested can be used as a positioning basis to extract base information of a specific position of the corresponding site, wherein the base information includes the base type and its number, that is, statistics per The number of bases.
- the base information includes the base type and its number, that is, statistics per The number of bases.
- Step S400 genotyping.
- the locus genotype is typed based on the extracted base information of the site to be tested. According to the extracted base information, the heterozygous type of the heterozygous type of the site to be tested can be determined more accurately.
- step S400 may include the following steps. It should be noted that, in a preferred embodiment, when the following heterozygous judgment does not meet the heterozygous determination condition, it is defined as Homozygous:
- step S410 the ratio is calculated. Calculate the ratio of the sum of the major base and the second base on the site to be tested. Wherein, the ratio of the sum of the main base and the second base is the ratio of the number of the main base and the second base at the site to be tested to the total number of all the bases at the site, and usually the ratio does not exceed 1.
- step S420 the ratio is judged. Determining the ratio between the ratio of the sum of the main base and the second base on the site to be tested and the preset value, and the preset value may be set based on different read depths according to experience, in the preferred embodiment, different The preset value ranges from 2/3 to 4/5 under the condition of the read depth. If the ratio is greater than the preset value, perform the following steps.
- Step S430 the second and third base numbers are acquired. Obtain the number of second bases and third bases on the site to be tested. In this embodiment, the number of second bases and the number of third bases should be acquired separately.
- step S440 the ratio of the number of second bases to the number of third bases is determined. Determining a ratio of the ratio of the second base number to the third base number; if the ratio of the second base to the third base exceeds a preset ratio, determining that the test target is a first tolerance heterozygous type; If the ratio of the second base to the third base is less than a preset ratio, it is determined that the to-be-tested site is a second latitude heterozygous type.
- the latitude is a value capable of characterizing the severity of the heterozygous determination condition, the first latitude is greater than the second latitude, and the smaller the latitude, the more severe the condition for the heterozygous determination characterizing the site.
- the value of the latitude should be no less than 1.
- the first latitude is in the range of 5 to 15 and the second latitude is in the range of 3 to 10.
- the third latitude is in the range of 3/2. 2.
- the preset ratio may be set based on different read depths according to experience. Preferably, the preset ratio ranges from 8 to 70 under different read depth conditions.
- step S420 if the ratio determined in step S420 is less than the preset value, it is determined that the to-be-tested location is a third latitude heterozygous type, wherein the third latitude is less than the second latitude.
- the preset ratio and the preset value can be set/adjusted in the system. Need to say It is obvious that when the depth of the sequencing data is different, the corresponding preset ratio and the preset value may be different. In addition, the corresponding first latitude, second latitude and the corresponding degree for characterizing the severity of the heterozygous determination are The value of the third latitude may also be different. Therefore, in a preferred embodiment, the preset ratio and the preset value may be adjusted according to the number of supported segments of the sequence to be tested, and the first latitude, The values of the second tolerance and the third tolerance. Specifically, the following steps may be performed before performing step S400:
- step S500 the read support number is acquired and judged. Obtaining the number of supported segments of the sequence to be tested, and if the number of supported supports is greater than the preset number of supports, adjusting the preset ratio, the preset value, the first latitude, the second latitude, and the number based on the obtained support number The value of the three latitudes. Please refer to Table 1 for an example of adjustment of each parameter corresponding to the number of different sequencing read supports tested in this embodiment.
- the "lowest base support number” is the minimum number of bases required for the number of sequencing reads supported; the “accuracy rate” is the accuracy of the above method when used in different sequencing reads. The rate is the result of selecting certain corresponding sites for sanger sequencing).
- Table 1 shows that the values of the corresponding preset ratio, preset value, first latitude, second latitude, and third latitude are different under different number of sequencing read support; when sequencing read support When the number is above 200X, a higher accuracy can be achieved.
- Table 1 is an example of a technical solution for adjusting each parameter according to the number of sequencing reads supported by the site to be tested, and cannot be determined as a limitation on the values of the parameters of the technical solution.
- the values in Table 1 can also be other values.
- the number of support points of the main base of the site to be tested is less than the preset support number (for example, 50X)
- the preset support number for example, 50X
- the number of primary base supports is too low, and it can be determined that the low coverage is insufficient. To accurately classify.
- the number of samples for the first test was 69, and the number of samples with the gold standard control was 23; the number of samples for the second test was 39, including the first test.
- the number of samples passed was 20, and the number of samples containing the gold standard was 9 (included in the first gold standard sample).
- the SNP locus with selected maf>0.3 was used as the analyzed locus, and there was no linkage disequilibrium between the loci.
- the locus was selected from the known database as much as possible, and the SNP loci were consistent with the Hardy-Weinberg equilibrium.
- the number of selected sites is 55, and the number of sites with gold standard is 36. Data and tool preparation are required before annotation, including target genomic sequence to be typed, sequence alignment software and typing program, and a list of selected sequences of known sites.
- the experimental scheme is as follows:
- rs For the naming of SNPs, after the classification of all submitted SNPs in the National Center for Biotechnology Information (NCBI), an "rs" number, also referred to as a reference SNP, is given. SNP specific Information, including context, position information, distribution frequency, etc. For example, “rs11239930” refers to the SNP site numbered rs11239930. One skilled in the art can determine the specific location of the SNP site based on the number in the NCBI database.
- the genotype of the non-gold standard sample is compared with the genotype of the gold standard sample, and the number of sites to be compared is also 36 sites containing the gold standard data, and the first test is found.
- the average agreement rate between the random sample and the corresponding site genotypes of other samples was 37.2%
- the average agreement rate between the random non-gold standard samples of the second test and the corresponding genotypes of the gold standard samples was 38.0%.
- a locus satisfying maf>0.3 was selected as the prediction consistency of the locus to be tested.
- the SNP typing results of the 39 samples in the second batch were compared with the genotypes of all the samples in the first batch.
- the locus was the 55 sites tested to remove 50 sites with low depth sites (rs1908593, rs6022576, rs11239930, rs2292564, rs4786795).
- the samples with the agreement rate >0.8 were identified as the same sample, and 19 samples were identified in the same 20 samples from the same batch.
- the PGM sequencing data of about 1000X and the randomly extracted 5000X, 2500X, 1000X, 500X, 200X, 100X, 50X are also selected in this embodiment.
- the typing effect of Proton sequencing data is shown in Figure 4.
- the average agreement rate of PGM sequencing data of 1000X in 37 samples with the gold standard in 23 samples is 96.2% (95% CI, 0.94-0.98).
- Table 9 the average agreement rate of the Proton sequencing data randomly selected from around 100X in the 36 samples and the gold standard in the 23 samples was 92.5% (95% CI, 0.89-0.96).
- Table 10 The specific results are shown in Table 10.
- Table 10 SNP typing results of 100X Proton sequencing data randomly selected from the gold standard
- the present embodiment also randomly samples the data of 200X four times.
- the accuracy of the four times is 92.5%, 92.2%, 92.6%, and 92.3%, respectively.
- the above experimental results show that the single-nucleotide polymorphism locus typing method based on base statistics disclosed in the present embodiment utilizes the quantitative relationship of the main base, the second base, and the third base, and The corresponding threshold parameters are set in the sequencing depth to determine the genotype of the site to be tested, which can effectively avoid the imbalance of the number of bases of SNP sites caused by experimental factors such as PCR, and can simultaneously output the selected sites at a time. Genotype, easy to operate.
- the high-depth data and relatively low data of NGS can give reliable high-accuracy classification results, the accuracy is higher than the conventional classification software, the classification results can be used for individual differentiation, and can effectively avoid The accuracy of the SNP imbalance caused by the PCR process is greatly improved.
- the method uses fixed-point SNP typing, there is no false positive typing result.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- General Health & Medical Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biotechnology (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Organic Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Microbiology (AREA)
- Molecular Biology (AREA)
- Immunology (AREA)
- Biochemistry (AREA)
- General Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
提供一种个体单核苷酸多态性SNP位点分型方法及装置,包括:获取含有SNP位点信息的读段序列;将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取所有比对后的读段序列待测位点的碱基信息;基于提取的待测位点的碱基信息对该位点基因型进行分型。由于基于提取的待测位点更为具体的碱基信息对该位点基因型进行分型,而非根据序列内在特征和统计概率模型进行分型,因此,可以简单有效、更为准确地实现个体定点分型。
Description
本发明涉及基因组学及生物信息学技术领域,具体涉及一种个体单核苷酸多态性(Single Nucleotide Polymorphisms,SNP)位点分型方法及装置。
目前,在测序序列上确定个体在某个位点的基因型(即单核苷酸多态性位点分型,简称SNP分型)已经是很多工作的一个必需步骤和基本前提,比如基因功能研究、致病基因鉴定、基因型鉴定等。分型结果的好坏直接影响到后续工作的有效性和准确性。
随着基因组学和生物信息学的不断发展,在二代平台上多种多样的SNP分型方法和软件日益涌现,但这些软件总体上都是基于序列内在特征和统计概率模型的de novo分型方法。在较低深度的测序数据中这些软件相对比较准确,但应用到具体的个体定点PCR(Polymerase Chain Reaction,聚合酶链式反应)扩增测序数据的SNP分型项目中就显得繁琐,并且错误率较高。
发明内容
本发明提供一种个体单核苷酸多态性位点分型方法及装置,可简单高效地统计SNP分型结果。
依据本发明的第一方面,一种实施方式提供一种个体单核苷酸多态性位点分型方法,包括:
获取含有SNP位点信息的读段序列;将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取所有比对后的读段序列待测位点的碱基信息;基于提取的待测位点的碱基信息对该位点基因型进行分型。
依据本发明的第二方面,一种实施方式提供一种个体单核苷酸多态性位点分型装置,包括:获取模块、比对模块、提取模块和分型模块,其中,获取模块用于获取含有SNP位点信息的读段序列;比对模块用于将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取模块用于提取所有比对后的读段序列待测位点的碱基信息;分型模块用于基于提取的待测位点的碱基信息对该位点基因型进行分型。
依据本发明的个体单核苷酸多态性位点分型方法/装置,由于基于提取的待测位点更为具体的碱基信息对该位点基因型进行分型,而不是根
据序列内在特征和统计概率模型进行分型,可以简单有效、更为准确地实现个体定点分型。特别地,当碱基信息为碱基类型及其数目时,其分型的准确度会进一步提高。
图1是本申请实施例公开的一种个体单核苷酸多态性位点分型装置结构示意图;
图2是本申请实施例公开的一种个体单核苷酸多态性位点分型方法流程图;
图3是本申请实施例的一种基于待测位点碱基信息分型的流程图;
图4是本申请实施例基于Proton对不同深度数据测序分型效果评测示意图。
下面通过具体实施方式结合附图对本发明作进一步详细说明。
本发明构思为:在含有SNP位点信息的读段序列集(reads)中某一SNP位点(以下称定点)上统计碱基个数,而后选定合适的阈值,根据选定的阈值实现达到一个较高准确度的个体SNP位点分型结果。该分型方法适合各种个体定点SNP检测的所有相关项目,譬如个体疾病位点的检测,个体的身份鉴定以及个体功能基因的检测等。
本申请中用到的术语定义:
(1)主要碱基,个体读段序列集合的SNP位点上,数目最多的碱基;第二碱基的数目次之;第三碱基的数目次于第二碱基的数目。
(2)宽容度,用于表征SNP位点杂合程度的标量。本实施例中,宽容度为主要碱基与第二碱基的比值,其值不小于1,宽容度越小,杂合型判定严苛程度越高。在其它可替代的实施例中,本领域技术人员可以依据该构思进行演变和推演,譬如将其定义为第二碱基与主要碱基的比值,譬如主要碱基和第二碱基的差值与主要碱基或第二碱基之比等,应当认为本实施例构思的宽容度等效替换。
现有的测序平台有多种,包括Roche454,Ion PGM和Ion Proton等。本发明中的实施例以第二代Ion Proton测序平台作说明,其他测序平台亦同样适用本发明所提供的方法,测序平台并不构成本发明的限制。实施例中未注明具体条件的,按照常规条件或制造商建议的条件进行;所用试剂或仪器未注明生产厂商的,均为可以通过市面购买获得的常规产品。
请参考图1,为本实施例公开的一种个体单核苷酸多态性位点(SNP位点)分型装置,包括:获取模块1、比对模块2、提取模块3和分型模块4,其中,
获取模块1用于获取含有SNP位点信息的读段序列;比对模块2用于将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取模块3用于提取比对上的读段序列待测位点的碱基信息;分型模块4用于基于提取的待测位点的碱基信息对该位点基因型进行分型。
在具体实施例中,碱基信息包括:碱基类型及其数目。分型模块4包括:碱基占比计算单元41、占比判断单元42、碱基数目统计单元43和基因型分型单元44,具体地,碱基占比计算单元41用于计算待测位点上主要碱基与第二碱基之和的所占比;占比判断单元42用于判断所占比与预设值的大小;碱基数目统计单元43用于获取待测位点上第二碱基与第三碱基的数目;基因型分型单元44用于对待测位点基因型进行分型,如果第二碱基与第三碱基之比超过预设比值,则判定该待测位点为第一宽容度杂合型;否则为第二宽容度杂合型;宽容度为能够表征杂合型判定条件严苛程度的数值(标量),第一宽容度大于第二宽容度;如果所占比小于预设值,则判定该待测位点为第三宽容度杂合型;第三宽容度小于第二宽容度。
在优选的实施例中,该个体SNP位点分型装置还包括:阈值调整模块5,阈值调整模块5用于获取待测位点测序读段的支持数目,如果获取的支持数目大于预设的支持数目,则基于获取的支持数目调整预设比值、预设值、第一宽容度、第二宽容度和第三宽容度的数值。
基于上述个体单核苷酸多态性位点分型装置,本实施例还公开了一种个体单核苷酸多态性位点分型方法,请参考图2,为该方法流程图,具体包括步骤如下:
步骤S100,读段序列获取。获取含有SNP位点信息的读段序列reads。在具体实施例中,在优选的实施例中,SNP位点为满足最小等位基因频率MAF>0.3,能通过设计引物进行扩增的位点;并且SNP位点之间符合Hardy-Weinberg平衡。
步骤S200,读段序列比对。将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集。在具体实施例中,对获取的读段序列reads进行基于PCR引物的数据统计,然后再利用比对软件在目的序列集中获得比对后的序列数据集。具体地,还可以根据读段序列reads与各个位点PCR引物的比对情况统计数据的有效性,包括每条reads对应的位点、每个位点对应的reads的长度和数量,有效reads数及百分比。在具体实施例中,譬如采用Tmap(一种比对软件)进行比对,在其它实施例中,也可以采用其它的比对方案。
步骤S300,碱基信息提取。提取所有比对后的读段序列待测位点
(SNP位点)的碱基信息。在具体实施例中,可以用待测位点前6-10bp的序列作为定位依据,提取出相应位点特定位置的碱基信息,其中,碱基信息包括碱基类型及其数目,即统计每种碱基的个数。当然,在其它实施例中,也可能只需提取主要碱基或者和第二碱基的个数,以及碱基总数。
步骤S400,基因分型。基于提取的待测位点的碱基信息对该位点基因型进行分型。根据提取的碱基信息可以较为准确地判定待测位点杂合型的杂合型。
请参考图3,在一具体实施例中,步骤S400可以包括如下步骤,需要说明的是,在优选的实施例中,当以下的杂合型判断中不符合杂合型的判定条件则定义为纯合型:
步骤S410,所占比计算。计算待测位点上主要碱基与第二碱基之和的所占比。其中,主要碱基与第二碱基之和的所占比为待测位点上主要碱基与第二碱基的个数与该位点上所有碱基总数之比,通常该比值不超过1。
步骤S420,所占比判断。判断待测位点上主要碱基与第二碱基之和的所占比与预设值之间的大小,预设值大小可以根据经验基于不同读段深度设置,优选实施例中,在不同读段深度的条件下预设值取值范围是2/3~4/5。如果所占比大于预设值,则执行如下步骤。
步骤S430,第二、三碱基数获取。获取待测位点上第二碱基与第三碱基的数目。在本实施例中,应分别获取第二碱基的数目和第三碱基的数目。
步骤S440,第二碱基数与第三碱基数之比判断。判断第二碱基数与第三碱基数之比的大小,如果第二碱基与第三碱基之比超过预设比值,则判定该待测位点为第一宽容度杂合型;如果第二碱基与第三碱基之比小于预设比值,则判定该待测位点为第二宽容度杂合型。其中,宽容度为能够表征杂合型判定条件严苛程度的数值,第一宽容度大于第二宽容度,宽容度越小,则表征该位点的杂合型判定的条件越苛刻。需要说明的是,宽容度的数值应不小于1。本步骤中在不同读段深度的条件下第一宽容度的取值范围是5~15,第二宽容度的取值范围是3~10,第三宽容度的取值范围是3/2~2。需要说明的是,预设比值可以根据经验基于不同读段深度设置,优选地,在不同读段深度的条件下预设比值取值范围是8~70。
在具体实施例中,如果步骤S420判断的所占比小于预设值,则判定该待测位点为第三宽容度杂合型,其中,第三宽容度小于第二宽容度。
上述实施例中,预设比值和预设值可以在系统中设置/调整。需要说
明的是,当测序数据深度不同时,其对应的预设比值和预设值可能会不同,此外,对应的用于表征杂合型判定严苛程度的第一宽容度、第二宽容度和第三宽容度的数值也会有所不同,因此,在优选的实施例中,可以根据待测位点测序读段的支持数目来调整预设比值和预设值,以及第一宽容度、第二宽容度和第三宽容度的数值。具体地,可以在执行步骤S400之前执行如下步骤:
步骤S500,读段支持数目获取并判断。获取待测位点测序读段的支持数目,如果获取的支持数目大于预设的支持数目,则基于获取的支持数目调整预设比值、预设值、第一宽容度、第二宽容度和第三宽容度的数值。请参考表1,为本实施例中测试的不同测序读段支持数目所对应的各参数调整的一种示例。
表1.各参数值调整参照表
表1中,“最低的碱基支持数”为对应测序读段支持数目下,所需要的最少碱基数目;“准确率”为上述方法在不同测序读段支持数目使用时的准确率(准确率是选取某些相应位点进行sanger测序的结果)。表1表明,在不同的测序读段支持数目下,其对应的预设比值、预设值、第一宽容度、第二宽容度和第三宽容度的数值有所不同;当测序读段支持数目在200X以上时,能够达到一个较高的准确度。
需要说明的是,表1为根据待测位点测序读段支持数目调整各参数技术方案的一种示例,不能认定为对该技术方案各参数数值的限定。在其它可替代的实施例中,表1中的数值也可以为其它数值。
在优选的实施例中,当步骤S500获取待测位点主要碱基的支持数目小于预设的支持数目(例如50X)时,则说明主要碱基支持数目过低,可以判定为低覆盖度不足以准确分型。
下文以人血浆中提取的DNA为样本为例进行实验说明。为证明流程的可重复性,分两次测试:第一次测试的样本数为69,其中有金标准对照的样本数为23;第二次测试的样本数为39,其中包含第一次测试过的样本数为20,含有金标准的样本数为9(包含于第一次的金标准样本中)。以选择的maf>0.3的SNP位点为分析的待测位点,位点之间无连锁不平衡,同时位点尽量从已知的数据库中选取,SNP位点之间符合Hardy-Weinberg平衡。所选取的位点数为55,其中有金标准的位点数为36。在进行注释之前需要做好数据及工具准备,包括待分型的目标基因组序列、序列比对软件和分型程序、选取的已知位点的定位序列列表。实验方案如下:
1)实验数据。对获取的读段序列数据进行基于PCR引物的数据统计,并去除扩增较差的位点。第一次所测数据的平均深度为6328X,数据的平均有效率为32.53%,具体每个位点的平均深度如表2所示;第二次所测数据的平均深度为6733X,数据的平均有效率为38.91%,具体每个位点的平均深度如表3所示。
表2 第一次测试基于PCR引物序列的位点平均深度统计结果
表3 第二次测试基于PCR引物序列的位点平均深度统计结果
对于SNP的命名,美国国立生物技术信息中心(National Center for Biotechnology Information,NCBI)里对所有提交的SNP进行分类考证之后,都会给出一个“rs”号,也可称作参考SNP,并给出SNP的具体
信息,包括前后序列,位置信息,分布频率等,例如“rs11239930”是指编号为rs11239930的SNP位点。本领域技术人员可以在NCBI数据库中根据该编号确定该SNP位点的具体位置。
2)分型结果比对。利用本实施例的方法对上述数据进行基因分型,并将该分型结果与其在sanger测序测的相应SNP位点的结果进行比较。统计36个位点(具有金标准对照的位点)每个样本与金标准的一致率,所称一致率为符合的位点数与比较的总位点数的比值。第一次测试得到23个样本与金标准的比较结果,平均一致率为93%(95%CI,0.90-0.96),这里CI即置信区间(confidence interval),第二次测试得到9个样本与金标准的比较结果,平均一致率为94%(95%CI,0.91-0.98),两次具体结果如表4-1、4-2和表5所示,其中具体展示了样本1的分型结果,其他样本比较与此相同。
表4-1 第一次测试SNP分型结果与金标准的一致率(准确度)
表4-2 样本1的具体结果展示
| SNP编号 | 金标准 | SNP分型结果 | 是否一致 |
| rs11239930 | GG | GG | 是 |
| rs10801520 | CT | CT | 是 |
| rs3899750 | GG | GG | 是 |
| rs11714239 | GT | TG | 是 |
| rs1397228 | GA | AG | 是 |
| rs472728 | GG | GG | 是 |
| rs7429010 | AG | AC | 否 |
| rs4478233 | TT | TT | 是 |
| rs2172651 | AG | low | low |
| rs325238 | CC | CC | 是 |
| rs7715674 | CT | CT | 是 |
| rs1337823 | AG | GA | 是 |
| rs574202 | GG | GG | 是 |
| rs7741536 | GG | GG | 是 |
| rs4719491 | AG | AG | 是 |
| rs13438255 | AA | AA | 是 |
| rs7834428 | TT | TT | 是 |
| rs6994603 | AA | AA | 是 |
| rs10124916 | GT | TG | 是 |
| rs4606122 | TT | TT | 是 |
| rs7035090 | CC | CC | 是 |
| rs2038597 | CT | TC | 是 |
| rs1484443 | TT | TT | 是 |
| rs518357 | CT | TC | 是 |
| rs895648 | CC | CC | 是 |
| rs1939904 | AG | GA | 是 |
| rs991718 | CC | CC | 是 |
| rs7306163 | GG | GG | 是 |
| rs10860402 | TT | TT | 是 |
| rs11146962 | AA | AA | 是 |
| rs1147437 | AA | AA | 是 |
| rs4789817 | AA | AA | 是 |
| rs8083190 | TG | GT | 是 |
| rs1908593 | TT | TT | 是 |
| rs2829066 | CC | CC | 是 |
| rs2076039 | CT | CT | 是 |
表5 第二次测试SNP分型结果与金标准的一致率
3)Samtools分型比对。运用Samtools软件将与步骤S200相同的比对方式(例如Tmap)比对后的数据作为输入文件,将SNP的分型结果与相应样品的金标准结果进行比较,比对的总位点数为36,第一次测试得到23个样品与标准的平均一致率为76%(95%CI,0.72-0.80),第二次测试得到9个样品与金标准的平均一致率为75%(95%CI,0.67-0.84)。两次测试的具体比较结果如表6和表7所示:
表6 Samtools第一次测试的SNP分型结果与金标准的比较
表7 Samtools第二次测试的SNP分型结果与金标准的比较
通过对比表4和表6以及表5和表7可知,本实施例公开的SNP分型结果相对于Samtools的SNP分型结果一致率更高。
4)为进一步验证本实施例的方法,将非金标准样本的基因型与金标准样本的基因型进行比较,比较的位点数同样是含有金标准数据的36个位点,发现第一次测试的随机样本与其它样本相应位点基因型的平均一致率为37.2%,第二次测试的随机非金标准样本与金标准样本相应位点基因型的平均一致率为38.0%,这与本实施例选取满足maf>0.3的位点作为待测位点的预测一致。此外,本实施例比较的位点中约有一半的位点在个体中是与参考序列一致的,未发生变异,在第一次测试中所占比例平均为51.4%,第二次测试中所占比例平均为52.7%。以第一批测试的69个样本的SNP分型数据作为数据库,将第二批测试的39个样本的SNP分型结果作为待鉴定样本分别与第一批的所有样本做基因型的比较,所用位点为所测的55个位点去除深度较低位点(rs1908593,rs6022576,rs11239930,rs2292564,rs4786795)共50个位点。将一致率>0.8的样本认定为同一样本,在两批相同的20个样本中,鉴定出了19个样本,另外未鉴定出的一个样本是由于测序深度不足导致的基因型无法判定导致的假阴性,在所鉴定的所有样本中无假阳性现象,本实施例提供的方法已经达到可以进行个体鉴定的准确度,具体结果见表8。
表8 两批样本随机个体鉴定
| 第二批测试的样本编号 | 对应在第一批样本编号 | 鉴定结果 |
| 1 | 无 | 无 |
| 2 | 无 | 无 |
| 3 | 无 | 无 |
| 8 | 无 | 无 |
| 9 | 无 | 无 |
| 14 | 无 | 无 |
| 15 | 无 | 无 |
| 16 | 无 | 无 |
| 17 | cell-11 | CELL-11 |
| 19 | 无 | 无 |
| 21 | cell-13 | CELL-13 |
| 23 | cell-14 | CELL-14 |
| 25 | cell-15 | CELL-15 |
| 27 | cell-16 | CELL-16 |
| 29 | cell-17 | CELL-17 |
| 30 | cell-19 | CELL-19 |
| 31 | cell-20 | CELL-20 |
| 33 | 无 | 无 |
| 34 | cell-3 | CELL-3 |
| 35 | cell-35 | CELL-35 |
| 37 | cell-40 | CELL-40 |
| 39 | cell-42 | CELL-42 |
| 40 | cell-45 | CELL-45 |
| 42 | cell-47 | CELL-47 |
| 44 | cell-54 | CELL-54 |
| 45 | cell-71 | LOW |
| 46 | cell-79 | CELL-79 |
| 47 | cell-91 | CELL-91 |
| 48 | cell-120 | CELL-120 |
| 49 | 无 | 无 |
| 51 | 无 | 无 |
| 53 | 无 | 无 |
| 55 | 无 | 无 |
| 57 | 无 | 无 |
| 59 | 无 | 无 |
| 61 | 无 | 无 |
| 62 | 无 | 无 |
| 63 | 无 | 无 |
| 64 | 无 | 无 |
5)为检测本实施例的方法在较低深度数据上的SNP分型效果,本实施例还在1000X左右的PGM测序数据和随机抽取的5000X,2500X,1000X,500X,200X,100X,50X左右的Proton测序数据的分型效果评测,如图4。1000X左右的PGM测序数据在23个样本37个位点与金标准的平均一致率为96.2%(95%CI,0.94-0.98),具体结果如表9;100X左右随机抽取的Proton测序数据在23个样本36个位点与金标准的平均一致率为92.5%(95%CI,0.89-0.96),具体结果如表10。
表9 PGM测序数据的SNP分型结果与金标准的一致率
表10随机抽取100X Proton测序数据的SNP分型结果与金标准的一致率
同时,为验证抽样数据的随机性,本实施例还对200X的数据又另外随机抽样4次,该四次的准确率分别为92.5%,92.2%,92.6%,92.3%。
上述实验结果表明,本实施例公开的基于碱基统计的个体单核苷酸多态性位点分型方法,利用主要碱基,第二碱基,第三碱基三者的数量关系,以及测序深度等方面设置相应的阈值参数从而确定待测位点的基因型,能够有效地避免因PCR等实验因素产生的SNP位点碱基数的失衡情况,同时能够一次性输出所选位点的基因型,操作简便。在NGS的高深度数据和相对较低数据时都能给出可靠的高准确度的分型结果,准确度高于常规的分型软件,分型结果能够用来做个体区分,同时能够有效避免对PCR过程产生的SNP失衡的误判,准确度大大提升。此外,由于该方法采用的是定点SNP分型,所以不会有假阳性的分型结果。
以上应用了具体个例对本发明进行阐述,只是用于帮助理解本发明并不用以限制本发明。对于本领域的一般技术人员,依据本发明的思想,可以对上述具体实施方式进行变化。
Claims (10)
- 一种个体单核苷酸多态性SNP位点分型方法,其特征在于,包括:获取含有SNP位点信息的读段序列;将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取所有比对后的读段序列待测位点的碱基信息;基于提取的待测位点的碱基信息对该位点基因型进行分型。
- 如权利要求1所述的单核苷酸多态性SNP位点分型方法,其特征在于,所述碱基信息包括:碱基类型及其数目;所述基于提取的待测位点的碱基信息对该位点基因型进行分型包括:计算待测位点上主要碱基与第二碱基之和的所占比;如果所述所占比大于预设值,则获取待测位点上第二碱基与第三碱基的数目;如果第二碱基与第三碱基之比超过预设比值,则判定该待测位点为第一宽容度杂合型;否则为第二宽容度杂合型;所述宽容度为能够表征杂合型判定严苛程度的标量,所述第一宽容度大于第二宽容度。
- 如权利要求2所述的单核苷酸多态性SNP位点分型方法,其特征在于,如果所述所占比小于预设值,则判定该待测位点为第三宽容度杂合型;所述第三宽容度小于第二宽容度。
- 如权利要求3所述的单核苷酸多态性SNP位点分型方法,其特征在于,在执行基于提取的待测位点的碱基信息对该位点基因型进行分型之前,还包括:获取待测位点测序读段的支持数目,如果所述获取的支持数目大于预设的支持数目,则基于获取的支持数目调整预设比值、预设值、第一宽容度、第二宽容度和第三宽容度的数值。
- 如权利要求4所述的单核苷酸多态性SNP位点分型方法,其特征在于,如果所述获取的支持数目小于预设的支持数目,则判定为该待测位点不足以准确分型。
- 如权利要求2-5任意一项所述的单核苷酸多态性SNP位点分型方法,其特征在于,所述宽容度为主要碱基与第二碱基的比值。
- 如权利要求1-5任意一项所述的单核苷酸多态性SNP位点分型方法,其特征在于,所述待测位点为满足最小等位基因频率MAF>0.3,能通过设计引物进行扩增的位点。
- 一种个体单核苷酸多态性SNP位点分型装置,其特征在于,包括:获取模块,用于获取含有SNP位点信息的读段序列;比对模块,用于将去除外接头的数据在含有扩增目的序列的目的库中进行序列的比对,获得比对后的读段序列集;提取模块,用于提取所有比对后的读段序列待测位点的碱基信息;分型模块,用于基于提取的待测位点的碱基信息对该位点基因型进行分型。
- 如权利要求8所述的单核苷酸多态性SNP位点分型装置,其特征在于,所述碱基信息包括:碱基类型及其数目;所述分型模块包括:碱基占比计算单元,用于计算待测位点上主要碱基与第二碱基之和的所占比;占比判断单元,用于判断所占比与预设值的大小;碱基数目统计单元,用于获取待测位点上第二碱基与第三碱基的数目;基因型分型单元,用于对待测位点基因型进行分型,如果第二碱基与第三碱基之比超过预设比值,则判定该待测位点为第一宽容度杂合型;否则为第二宽容度杂合型;所述宽容度为能够表征杂合型判定严苛程度的数值,所述第一宽容度大于第二宽容度;如果所述所占比小于预设值,则判定该待测位点为第三宽容度杂合型;所述第三宽容度小于第二宽容度。
- 如权利要求9所述的单核苷酸多态性SNP位点分型装置,其特征在于,还包括:阈值调整模块,用于获取待测位点测序读段的支持数目,如果所述获取的支持数目大于预设的支持数目,则基于获取的支持数目调整预设比值、预设值、第一宽容度、第二宽容度和第三宽容度的数值。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201480082814.2A CN107075565B (zh) | 2014-11-21 | 2014-11-21 | 个体单核苷酸多态性位点分型方法及装置 |
| PCT/CN2014/091824 WO2016078067A1 (zh) | 2014-11-21 | 2014-11-21 | 个体单核苷酸多态性位点分型方法及装置 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2014/091824 WO2016078067A1 (zh) | 2014-11-21 | 2014-11-21 | 个体单核苷酸多态性位点分型方法及装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2016078067A1 true WO2016078067A1 (zh) | 2016-05-26 |
Family
ID=56013092
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2014/091824 Ceased WO2016078067A1 (zh) | 2014-11-21 | 2014-11-21 | 个体单核苷酸多态性位点分型方法及装置 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN107075565B (zh) |
| WO (1) | WO2016078067A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114220477A (zh) * | 2021-12-29 | 2022-03-22 | 天津华大医学检验所有限公司 | 一种ace基因分型的方法及系统 |
| CN115035950A (zh) * | 2022-06-28 | 2022-09-09 | 广州燃石医学检验所有限公司 | 基因型检测方法、样本污染检测方法、装置、设备及介质 |
| CN116064755A (zh) * | 2023-01-12 | 2023-05-05 | 华中科技大学同济医学院附属同济医院 | 一种基于连锁基因突变检测mrd标志物的装置 |
| CN116935959A (zh) * | 2023-04-25 | 2023-10-24 | 山东省农业科学院畜牧兽医研究所 | Sanger基因测序结果快速判读方法、系统及介质 |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110942806A (zh) * | 2018-09-25 | 2020-03-31 | 深圳华大法医科技有限公司 | 一种血型基因分型方法和装置及存储介质 |
| KR102374615B1 (ko) * | 2019-05-22 | 2022-03-16 | 서울대학교산학협력단 | Ngs 데이터를 이용하여 유전형을 예측하는 방법 및 장치 |
| CN111235239A (zh) * | 2020-03-18 | 2020-06-05 | 浙江大学医学院附属妇产科医院 | 一种多重pcr_snp基因分型检测方法 |
| CN112669903B (zh) * | 2020-12-29 | 2024-04-02 | 北京旌准医疗科技有限公司 | 基于Sanger测序的HLA分型方法及设备 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101539967A (zh) * | 2008-12-12 | 2009-09-23 | 深圳华大基因研究院 | 一种单核苷酸多态性检测方法 |
| US20110301854A1 (en) * | 2010-06-08 | 2011-12-08 | Curry Bo U | Method of Determining Allele-Specific Copy Number of a SNP |
| CN103160937A (zh) * | 2011-12-15 | 2013-06-19 | 深圳华大基因科技有限公司 | 对高等植物复杂基因组基因进行富集建库和snp分析的方法 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| AU2000267888A1 (en) * | 2000-03-28 | 2001-10-08 | Nanogen, Inc. | Methods for determination of single nucleic acid polymorphisms using a bioelectronic microchip |
| JP2002345489A (ja) * | 2000-11-03 | 2002-12-03 | Astrazeneca Ab | 化学物質 |
-
2014
- 2014-11-21 WO PCT/CN2014/091824 patent/WO2016078067A1/zh not_active Ceased
- 2014-11-21 CN CN201480082814.2A patent/CN107075565B/zh active Active
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101539967A (zh) * | 2008-12-12 | 2009-09-23 | 深圳华大基因研究院 | 一种单核苷酸多态性检测方法 |
| US20110301854A1 (en) * | 2010-06-08 | 2011-12-08 | Curry Bo U | Method of Determining Allele-Specific Copy Number of a SNP |
| CN103160937A (zh) * | 2011-12-15 | 2013-06-19 | 深圳华大基因科技有限公司 | 对高等植物复杂基因组基因进行富集建库和snp分析的方法 |
Non-Patent Citations (1)
| Title |
|---|
| WANG YANGKUN ET AL.: "Current Status and Perspective of RAD-seq in Genomic Research", GENETICS, vol. 36, no. 1, 16 October 2013 (2013-10-16), pages 41 - 49 * |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114220477A (zh) * | 2021-12-29 | 2022-03-22 | 天津华大医学检验所有限公司 | 一种ace基因分型的方法及系统 |
| CN115035950A (zh) * | 2022-06-28 | 2022-09-09 | 广州燃石医学检验所有限公司 | 基因型检测方法、样本污染检测方法、装置、设备及介质 |
| CN115035950B (zh) * | 2022-06-28 | 2025-08-12 | 广州燃石医学检验所有限公司 | 基因型检测方法、样本污染检测方法、装置、设备及介质 |
| CN116064755A (zh) * | 2023-01-12 | 2023-05-05 | 华中科技大学同济医学院附属同济医院 | 一种基于连锁基因突变检测mrd标志物的装置 |
| CN116064755B (zh) * | 2023-01-12 | 2023-10-20 | 华中科技大学同济医学院附属同济医院 | 一种基于连锁基因突变检测mrd标志物的装置 |
| CN116935959A (zh) * | 2023-04-25 | 2023-10-24 | 山东省农业科学院畜牧兽医研究所 | Sanger基因测序结果快速判读方法、系统及介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN107075565A (zh) | 2017-08-18 |
| CN107075565B (zh) | 2021-07-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2016078067A1 (zh) | 个体单核苷酸多态性位点分型方法及装置 | |
| Li et al. | SNP detection for massively parallel whole-genome resequencing | |
| CN106462670B (zh) | 超深度测序中的罕见变体召集 | |
| CN105441432B (zh) | 组合物及其在序列测定和变异检测中的用途 | |
| TWI636255B (zh) | 癌症檢測之血漿dna突變分析 | |
| CN106715711B (zh) | 确定探针序列的方法和基因组结构变异的检测方法 | |
| CN103221551B (zh) | Hla基因型别-snp连锁数据库、其构建方法、以及hla分型方法 | |
| JP2014502845A5 (zh) | ||
| CN114530198A (zh) | 一种用于检测样本污染水平的snp位点的筛选方法及样本污染水平的检测方法 | |
| CN104182655B (zh) | 一种判断胎儿基因型的方法 | |
| JP2019500706A5 (zh) | ||
| CN108026576A (zh) | 通过母亲血浆dna的浅深度测序准确定量胎儿dna分数 | |
| CN114258572A (zh) | 用于确定基因组倍性的系统和方法 | |
| Anderson et al. | ReCombine: a suite of programs for detection and analysis of meiotic recombination in whole-genome datasets | |
| Zhao et al. | Assessing linkage disequilibrium in a complex genetic system. I. Overall deviation from random association | |
| JP2021526857A (ja) | 生体試料のフィンガープリンティングのための方法 | |
| US20180247019A1 (en) | Method for determining whether cells or cell groups are derived from same person, or unrelated persons, or parent and child, or persons in blood relationship | |
| EP2971126B1 (en) | Determining fetal genomes for multiple fetus pregnancies | |
| CN108823330A (zh) | 一种大豆hrm-snp分子标记点标记方法及其应用 | |
| CN116888274A (zh) | 一种利用多态性位点和靶位点测序检测胎儿遗传变异的方法 | |
| CN118866116B (zh) | 一种测序样本污染的分析方法、装置、系统及存储介质 | |
| CN110475874A (zh) | 脱靶序列在dna分析中的应用 | |
| JP2023547610A5 (zh) | ||
| CN110993024A (zh) | 建立胎儿浓度校正模型的方法及装置与胎儿浓度定量的方法及装置 | |
| CN111128297B (zh) | 一种基因芯片的制备方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 14906526 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 14906526 Country of ref document: EP Kind code of ref document: A1 |












