WO2011106994A1 - 基于聚合酶链式反应产物测序序列分型的实现方法和系统 - Google Patents
基于聚合酶链式反应产物测序序列分型的实现方法和系统 Download PDFInfo
- Publication number
- WO2011106994A1 WO2011106994A1 PCT/CN2011/000347 CN2011000347W WO2011106994A1 WO 2011106994 A1 WO2011106994 A1 WO 2011106994A1 CN 2011000347 W CN2011000347 W CN 2011000347W WO 2011106994 A1 WO2011106994 A1 WO 2011106994A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- typing
- database
- sequence
- base
- typed
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
Definitions
- the present invention relates to allelic typing techniques, and more particularly to an implementation method and system based on Polymerase Chain Reaction Sequencing-based Typing (PCR-SBT). Background technique
- HLA Human Leucocyte Antigen
- HLA Human Leucocyte Antigen
- PCR-SSP Polymerase Chain Reaction Sequence -Specific Primers, 1 J bad sequence specific primer polymerase chain reaction
- PCR-SSO Polymerase Chain Reaction equence- Speciflc Oligonucleotide Probe Hybridization, polymerization drunk Chain-type anti-oligonucleotide probe hybridization
- PCR-SBT Polymerase Chain Reaction Sequencing-based Typing
- PCR-SSO The principle of PCR-SSO is to design an HLA-specific oligonucleotide sequence as a probe, label the PCR product, and hybridize the PCR product (the DNA to be detected) to the probe, and determine the HLA genotype by detecting the fluorescent signal.
- the disadvantage is that the new allele cannot be detected and the resolution is not high enough; the detection signal is an analog signal.
- the PCR-SSP method designed a set of allele-specific primers to obtain HLA-type specific amplification products by PCR, and directly analyzed the band type to determine the HLA type by electrophoresis.
- the disadvantage is that it is not easy to automate; it cannot detect new alleles.
- PCR-SBT The principle of PCR-SBT is to use a primer to PCR-amplify the region of the HLA gene polymorphism, and then perform DNA sequencing on the amplified product to determine the HLA allelic type under computer aid.
- SBT is a relatively straightforward and accurate method. New alleles were identified by PCR-SSP or PCR-SSO methods and are usually confirmed by sequencing.
- the PCR-SBT method requires high equipment, time and cost, and the prior art performs slow screening and low-visibility screening for a wide range of candidate types.
- the sequence cannot be synchronized and the peak map corresponds, and it cannot be viewed by adjusting the peak map.
- backup and recovery of data results cannot be achieved, which is likely to cause data loss.
- One technical problem to be solved by the present invention is to provide an implementation method and system for sequencing based on sequencing of polymerase chain reaction products, in particular, an implementation method and system for sequencing of HLA based polymerase chain reaction product sequencing sequences , can improve the speed and efficiency of the typing.
- a method for performing sequencing based on sequencing of a polymerase chain reaction product comprising the steps of: interpreting a heterozygous site and a base sequence to be classified according to a sequencing result by a computer program;
- the ligated base sequence of the zygote is aligned to the typing database of the corresponding site, and the positional relationship of the reference sequence of the base sequence to be typed and the typing database is identified; according to the base sequence to be typed and
- the matching positional relationship of the reference sequence of the typing database retrieves the allelic type in the typing database, and obtains the penalty value of the allelic type in the typing database according to the sequencing strategy; according to the allele in the typing database
- the type of penalty value obtains a candidate type combination set.
- the step of retrieving the allelic type in the typing database according to the positional relationship between the base sequence to be typed and the reference sequence of the typing database includes:
- the sequencing strategy accumulates the scores weighted by the different mismatch types in units of DNA sequencing base mass and as a penalty value.
- the method further includes the steps of: graphically displaying the peak map of the sequencing result file, performing peak map shape scaling adjustment and/or sequence peak map linkage viewing, so as to facilitate modification by the typing personnel. And / or confirm the typing results.
- a sequencing system based on a polymerase chain reaction product sequencing sequence, comprising: a base sequence determining subsystem, configured to receive a sequencing result, and to interpret a heterozygous locus according to a sequencing result a typing base sequence; a matching position recognition subsystem for receiving a base sequence to be typed from a base sequence determining subsystem, and matching the base sequence to be typed to a matching database of the corresponding site, identifying a matching positional relationship between the base sequence to be typed and the reference sequence of the typing database; a penalty value determining subsystem for searching the matching positional relationship according to the reference sequence of the base sequence to be typed and the typing database
- the allelic type in the typing database, the penalty value of the allelic type in the typing database is obtained according to the sequencing strategy; the candidate type determining subsystem is used for the punishment according to the allelic type in the typing database Score Select the combination set.
- An embodiment of the analysis system further comprising: an index preprocessing subsystem, configured to pre-establish a reference sequence of the typing database, and the reference sequence and the allelic sequence of the typing database a positional correspondence between the two; a hash array formed by a base symbol sequence on a variable base position in the allelic type of the typing database; a base symbol placed, arranged in order, and then traversing the minute A hash array of the type database, which is scored.
- An embodiment of the parting system according to the present invention further includes: a graphical display subsystem, configured to graphically display and output a peak map of the sequenced crust file, perform peak map shape scaling adjustment, and/or sequence peak map linkage View, so that the typing staff can modify and/or confirm the typing results.
- a graphical display subsystem configured to graphically display and output a peak map of the sequenced crust file, perform peak map shape scaling adjustment, and/or sequence peak map linkage View, so that the typing staff can modify and/or confirm the typing results.
- the system can realize automatic identification of candidate genotypes through computers and other devices, and the processing speed is fast, and the typing efficiency is improved.
- the technical means such as the graphical display interface, it provides convenience for the identification and modification of the type of the personnel, and provides the classification accuracy and the classification efficiency.
- FIG. 1 is a flow chart showing an embodiment of an implementation method of PCR-SBT typing according to the present invention
- Figure 2 is a flow chart showing another embodiment of a method of implementing PCR-SBT typing according to the present invention
- Figure 3 is a cross-sectional view showing a data graphical output interface of an application example of the present invention.
- FIG. 4 shows an embodiment of a PCR-SBT typing system in accordance with the present invention.
- Fig. 5 is a view showing the construction of another embodiment of the PCR-SBT typing system according to the present invention. detailed description
- PCR-SBT PCR-SBT. Type method.
- the heterozygous and the base sequence to be typed are sequenced by computer program sequencing results.
- sequencing results include fluorescence signal intensity readings for a fixed length interval in the time domain, and signal peaks are determined based on fluorescence signal intensity readings to determine bases or heterozygotes.
- a heterozygous base sequence to be typed is aligned to a typing database of the corresponding site, and a positional relationship of the reference sequence of the typing database of the base sequence to be typed and the corresponding site is identified.
- the target site is known in the PCR amplification assay, and the target site corresponding to the base sequence to be typed is also known.
- the target site information according to the base sequence to be typed is compared to the typing database of the corresponding site.
- Each typing database includes a Reference Sequence, and the positional correspondence of each allele in the reference sequence and the typing database.
- a dynamic programming algorithm Dynamic Programmin
- Dot Matrix dot matrix
- step 106 combining the positional relationship of the reference sequence to be typed and the reference sequence of the typing database, and searching for each allelic type in the typing database, according to The ranking strategy obtains the penalty value for each allele in the typing database.
- Traversing the allelic type in the typing database obtaining the base sequence and allele to be typed according to the positional relationship of the base sequence to be typed and the reference sequence, the positional correspondence of the reference sequence and the allelic type The positional correspondence of the type sequence, and then combined with the sequencing strategy to obtain the penalty value of each allele type.
- the sequencing strategy can be based on the type of mismatch at the mismatch position between the base sequence to be typed and the compared allelic type (the allelic type), and the quality of the base sequence to be typed at that position. Accumulation of points and as a basis for sequencing
- a candidate type combination set is obtained based on the penalty value of the allelic type in the typing database. For example, select TopN candidate type combination sets according to the set penalty threshold; or select the first candidate type as the determined allelic type.
- the method for realizing the PCR-SBT typing according to the embodiment of the present invention automatically completes the typing of the gene sequence to be typed by a computing device, etc., especially for a large-scale candidate type screening search, which is fast and efficient.
- the allelic sequence information in the classification database is pre-processed in advance, and the mutated base in the corresponding allelic type is referred to.
- the base symbol at the position corresponding to the variant base is taken out, and the variant base symbols are sequentially arranged and coded to form a hash array.
- the storage format of a hash array is: Key-values for each type in the database.
- the base symbol sequence corresponding to all variable base positions is coded into a binary code.
- the base symbols at the corresponding positions of the variant base are taken out, arranged in order, and then the hash array of the allelic type is traversed according to the key value of the hash array, and the base sequence is treated.
- the allelic sequence in the classification database is scored, the penalty value of each allelic type is obtained according to the sequencing strategy, and then TopN candidate types are obtained according to the penalty value.
- the hash array, the hash array can be resident after the memory is built, and the different base sequences to be typed can be processed repeatedly, the retrieval speed is fast, the retrieval efficiency is greatly improved, and the whole classification of the invention is improved. The speed and efficiency of the method.
- the value of the base of the DNA sequencing sequence and the weighting of the different mismatch types are used as a penalty. Scores are scored and ordered in a small to large rule to obtain a list of candidate genotypes.
- the sequencing strategy uses the following rules: Let the quality value of the mismatched site be q, the base sequence to be typed Q and the target genotype T, Bay, J:
- Non-missing dislocations if Q is heterozygous and is not included in the degenerate base of the corresponding position in T, then +3q.
- the base quality value represents the probability that the base Base Calling result is wrong. The higher the value, the higher the probability of error. Therefore, the base quality is used as the base value (or reference value) of the penalty, which can improve the typing. accuracy.
- the candidate result closest to the real genotype is more effectively placed in the highest priority position of the candidate list, compared with the existing method of ordering only the number of mismatched bases. , thereby improving the efficiency of typing.
- FIG. 2 is a flow chart showing another embodiment of the method of implementing the PCR-SBT typing of the present invention.
- a sequencing result file is input.
- the content and format of the sequencing result file can be seen.
- the heterozygous site and the base sequence are automatically interpreted based on the sequencing result file.
- base and heterozygous interpretations are automatically performed by computer through the Reference-Based Base Calling method.
- the automatic matching and automatic insertion/deletion identification of the reference sequence of the typing database of the target base sequence and the corresponding site are performed.
- various alignment algorithms are introduced. 0 can be automatically linked using a bounded global alignment algorithm (Banded Global Alignment algorithm). Match.
- a classification database in which an index has been established is retrieved, and a candidate allelic type combination set is collected.
- step 210 a list of candidate allele types that are automatically sequenced according to the size of the penalty
- step 212 the data graphically displays the fluorescence signal intensity readings of the fixed length interval in the time domain recorded in the output / sequencing result file, corresponding to the time domain of the sequencing process; the fluorescent signal has four colors, respectively corresponding to four bases, by pressing Specific step size and curve fitting formulas can be used to plot the course of the fluorescence signal over the entire time domain.
- a graph is drawn and displayed by a computer program analyzing the sequencing results.
- the peak map signal waveform diagram in the sequencing result and the corresponding base interpretation and heterozygous recognition result, and the base sequence obtained by automatic interpretation in the sequencing result are matched according to the selected candidate genotype sequence of the target site.
- the positional relationship is displayed overall in the same view form.
- step 214 the peak map shape adjustment and the sequence peak map are linked and viewed.
- the display of the peak map supports single-dimensional zooming and zooming to provide better usability.
- the sequenced base sequence and the candidate genotype sequence are aligned, the sequenced similarities and differences information is collated, the mismatched position is drawn as an overview view of the entire site, and synchronous jumps between the positions on the respective form views are supported. .
- the specific implementation is as follows:
- Sequence peak map linkage view After the peak map corresponding to the multiple sequencing result files is displayed on the panel by the program, the position correspondence between the base sequence to be typed and the reference sequence obtained by each sequencing result file is used as a basis. Thereby, a positional correspondence relationship between each base sequence to be typed is obtained. The positions of the bases on different base sequences to be typed are relative. Each time a base is triggered, the program first searches for the corresponding base position of the other sequence according to the base, plus the base. Offset the value and redraw the complete peak map. After each trigger, make sure that the bases that are triggered this time are aligned, and that the sequence and peak maps at other locations are only guaranteed to be substantially aligned.
- Peak shape adjustment The enlargement and reduction of the peak map can be used to adjust the peak map's parameters each time to adjust the peak map enlargement and reduction ratio to achieve the enlargement and reduction of the peak map.
- the lifting and squeezing of the peak map can be carried out according to the parameters of the set peak map in the lateral direction.
- a typing result and a data backup are obtained.
- the work of the type factor includes the reference peak map to check the false homozygous and confirm or modify the hybrid position automatically recognized by the computer, and modify and final confirm the current sequence to be typed by referring to the rare type list.
- This section currently requires the human eye to identify and grasp the scale of professionally trained and experienced professionals.
- the type factor also needs to confirm when the result of the suspected new gene (the high-quality position does not match the current database) occurs; when the fuzzy result (probably the A genotype, also conforms to the B genotype) appears, GSSP primers are given for additional validation using SSP typing techniques. In addition, the discovery of new genes needs to be confirmed by people.
- the software will copy the type file to the folder specified by the software, and save the type file and generate a temporary file under the folder, mainly record classification.
- Information at the time of classification such as: modification of a certain base at a certain position, the effective range of the captured peak map, the result of the classification, and the like. In this way, when the user wants to view the previously corrected peak image file again, the user can directly open the classification history panel to view the corrected file according to the flood period.
- the software can achieve alignment and linkage between specific bases and sequences according to the algorithm, and when selecting the peak map, the sequence can ensure synchronization and peak map correspondence.
- the automatic alignment of the sequence is realized, and the correction is not performed manually, which shortens the time for identifying the position of the sequence by the human eye, and greatly improves the efficiency of the typing worker.
- the size of the peak image can be adjusted by zooming in and out of the peak image, which is convenient for the typer to view the peak image. , improve the efficiency of typing.
- Figure 3 shows a screenshot of a data graphical output interface of one application of the present invention.
- the main interface of the SBT parting software in this application example is divided into four parts: an upper left part 31, an upper right part 32, a lower left part 33, and a lower right part 34.
- the upper left part 31 is the selected part of the parting file.
- the type file selected by the typeifier will be displayed in the form of a tree structure in this part, which is convenient for the user to select the type file.
- the lower left part 33 identifies the matching point of the typed file according to the file name of the selected type file, selects the matching type database according to the type of the typed point, and divides the penalty according to the candidate type from small to large.
- the ordering rules are sorted, which gives a list of all the types in the lower left corner.
- the upper right portion 32 is the display portion of the selected parting file sequence. Among them, the first line Consensus behavior site comparison library full sequence; the second row Forward behavior selects the forward sequence of the classification file; Reverse is the selection of the classification file The reverse sequence of the pieces; the third line of Patterns matches the result of the matching sequence of the forward and reverse sequences; the bottom two sequences are the selection list from the lower left corner.
- the lower right part 34 is the display part of the peak map and the sequence file. Above the peak map is the sequence corresponding to the peak map. Each peak corresponds to one base. The correspondence between the upper and lower peak maps is based on the comparison result between the classification file and the database. The peaks and peaks of the upper and lower peaks of the peak map, the bases and bases above and below the peak map, All correspond.
- Fig. 4 is a view showing the construction of a PCR-SBT typing system according to an embodiment of the present invention.
- the parting system of this embodiment includes a base sequence determining subsystem 41, a joint position identifying subsystem 42, a penalty value determining subsystem 43, and a candidate type determining subsystem 44.
- the base sequence determining subsystem 41 is configured to receive the sequencing result, and read the heterozygous locus and the base sequence to be typed according to the sequencing result.
- the matching location identification subsystem 42 is configured to receive the base sequence to be typed from the base sequence determination subsystem 41, and compare the base sequences to be typed into a classification database of the corresponding sites to identify the base to be classified.
- the positional relationship between the base sequence and the reference sequence of the typing database of the corresponding site For example, the joint position recognition subsystem 42 identifies the joint positional relationship of the reference sequence of the base sequence to be typed and the classification database by a dynamic programming algorithm or a point matrix method.
- the penalty value determination subsystem 43 is configured to retrieve an allelic type in the classification database based on the joint positional relationship of the base sequence to be typed and the reference sequence of the classification database identified by the joint position recognition subsystem 42, The penalty value for each allelic type in the typing database is obtained according to the sequencing strategy. For example, the sequencing strategy accumulates the scores weighted by the different types of mismatches in units of DNA sequencing bases and as a penalty.
- a candidate type determination subsystem 44 is configured to obtain a candidate type combination set based on the penalty value of each allelic type in the classification database.
- the parting system also optionally includes an index pre-processing subsystem 45.
- the index pre-processing subsystem 45 is configured to pre-establish a reference sequence of the typing database, and a positional correspondence between the reference sequence of the typing database and each allelic sequence of the typing database; further, the index pre-processing subsystem 45
- the base symbol at the position corresponding to the mutated base (abbreviated as the mutated base) in the allelic type in the typing database is taken out, and the mutated base symbols are sequentially arranged and encoded to form a hash array, that is, A hash array formed from the base symbols of the variant bases in the allelic type of the typing database.
- the penalty value determination subsystem 43 takes the base symbols at the corresponding positions of the variant bases from the base sequence to be typed, arranges them in order, and then traverses the hash array of the classification database to perform the matching.
- Fig. 5 is a view showing the construction of another PCR-SBT typing system according to an embodiment of the present invention.
- the parting system of this embodiment includes a base sequence judgment subsystem 51, a joint position recognition subsystem 52, a penalty value determination subsystem 53, a candidate type determination subsystem 54, and an index preprocessor.
- System 55 graphical display subsystem 56 and data backup system 57.
- the base sequence judgment subsystem 51, the joint position recognition subsystem 52, the penalty value determination subsystem 53, the candidate type determination subsystem 54, and the index preprocessing subsystem 55 can refer to the corresponding subsystem in the above embodiment. The description is not described in detail here for the sake of brevity.
- the graphical display subsystem 56 is used to graphically display the peak map of the sequencing result file for peak image shape scaling adjustment and/or sequence peak map linkage viewing, so that the typing personnel can modify and/or confirm the typing result.
- the graphical display subsystem 56 compares the peak map signal waveform pattern in the sequencing result with the corresponding base interpretation and heterozygous recognition result, and the base sequence obtained by automatically interpreting the sequencing result, together with the selected candidate of the target site.
- the genotype sequence is displayed in the same view form as a whole according to the joint positional relationship.
- the data backup system 57 is used for storage and backup confirmation of the results of the classification, as well as modifications made by the type of personnel.
- the software will copy the type file to the folder specified by the software, and save the type file and generate a temporary file under the folder, mainly record classification.
- Information when the person is typing such as: Modifying a certain base at a certain position, intercepting the effective range of the peak map, and the result of the classification.
- the existing software analysis speed is generally 10 / hour
- the solution of the present invention can reach 15 / hour
- the existing software accuracy is generally 90%
- the solution of the present invention can reach 92%.
- the method and system of the present invention can increase the typing speed by 50% per unit time, and the accuracy of base recognition can be increased by 2%.
- the SBT typing method and system provided by the invention can realize the automatic identification of the candidate genotypes by using a computer and the like, thereby improving the typing efficiency; providing the classification confirmation and modification of the typing personnel through the technical means such as a graphical display interface. Convenient, providing typing accuracy and typing efficiency.
- the SBT typing method and system of the present invention can be applied not only to HLA typing, but also to HPV (Human papillomavirus), HBV (hepatitis B virus), and the like.
- HPV Human papillomavirus
- HBV hepatitis B virus
- the invention can be applied to any species with a typing requirement, supported by a typing database.
Landscapes
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Description
基于聚合酶链式反应产物测序序列分型的实现方法和
系统 技术领域
本发明涉及等位基因分型技术, 尤其涉及一种基于聚合酶链 式反应产物测序序列分型 ( Polymerase Chain Reaction Sequencing-based Typing, PCR-SBT ) 的实现方法和系统。 背景技术
HLA ( Human Leucocyte Antigen , 人类白细胞抗原 )是迄 今为止发现的多态性最高的基因系统之一, 是调控人体特异性免 疫应答和决定疾病易感性个体差异的主要基因系统, HLA与同种 异体器官移植的排斥反应密切相关。
目前国际标准的 HLA 分型技术为 PCR-SSP ( Polymerase Chain Reaction Sequence-Specific Primers, 序歹1 J特异引物聚合酶 链式反应 ), PCR-SSO ( Polymerase Chain Reaction equence- Speciflc Oligonucleotide Probe Hybridization , 聚合醉链式反 寡 核苷酸探针杂交) 和 PCR-SBT ( Polymerase Chain Reaction Sequencing-based Typing, 基于聚合酶链式反应产物测序序列分 型)。
PCR-SSO 的原理是设计 HLA型别特异的寡核苷酸序列作 为探针, 把 PCR产物标记, 以 PCR产物 (待检测基因 DNA ) 与探针杂交, 通过检测荧光信号判断 HLA基因型别。 缺点是不 能检测新的等位基因, 分辨率不够高; 检测信号是模拟信号。
PCR-SSP方法通过设计出一整套等位基因组特异性引物, 借 助 PCR技术获得 HLA型别特异的扩增产物, 通过电泳直接分析 带型决定 HLA 型别。 其缺点是不易自动化; 不能检测新的等位
基因; 试剂盒需不断升级; 检测信号是模拟信号。
PCR-SBT 的原理是用引物对 HLA 基因多态性的区域进行 PCR扩增, 然后对扩增产物进行 DNA测序, 在计算机辅助下确 定 HLA 等位基因型别。 对于基因结构的分析, SBT 是比较直 观、 准确的方法。 用 PCR-SSP或 PCR-SSO方法鉴别出新的等位 基因, 通常通过测序加以证实。
SBT技术中, 利用扩增产物对 DNA 测序后, 需要通过软件 将测序所得结杲与国际组织 IMGT 数据库 (http://www.ebi.ac. uk/imgt/)中公布的 HLA 分型中的标准序列进行比对; 通过比对 得出样品序列与标准序列的匹配率; 根据匹配率高低得出样品序 列的分型结论。
但是, PCR-SBT 方法的设备要求高, 时间和费用的消耗 大, 现有技术进行大范围候选型别筛查检索时, 速度慢, 效率 低。 此外, 人工辅助分型阶段, 分型人员进行峰图查看时, 序列 无法同步和峰图对应, 而且无法通过对峰图的调节来进行查看。 同时, 无法实现数据结果的备份和恢复, 容易造成数据的丢失。 发明内容
本发明要解决的一个技术问题是提供一种基于聚合酶链式反 应产物测序序列分型的实现方法和系统, 特别是一种 HLA基于 聚合酶链式反应产物测序序列分型的实现方法和系统, 可以提高 分型速度和效率。
根据本发明的一个方面, 提供一种基于聚合酶链式反应产物 测序序列分型的实现方法, 包括步驟: 通过计算机程序根据测序 结果判读杂合子位点和待分型碱基序列; 将含有杂合子的待分型 碱基序列比对到对应位点的分型数据库, 识别待分型碱基序列和 分型数据库的参考序列的联配位置关系; 根据待分型碱基序列和
分型数据库的参考序列的联配位置关系检索分型数据库中的等位 基因型, 根据定序策略获得分型数据库中的等位基因型的罚分 值; 根据分型数据库中的等位基因型的罚分值获得候选型别组合 集。
进一步, 根据所述待分型碱基序列和所述分型数据库的参考 序列的联配位置关系检索所述分型数据库中的等位基因型的步骤 包括:
从待分型碱基序列中取出变异碱基对应位置上的碱基符号, 顺序排列, 然后遍历预先建立的所述分型数据库中等位基因型中 变异碱基的碱基符号形成的哈希数组, 进行打分, 获得各个等位 基因型的罚分值。
才艮据本发明的方法的一个实施例, 定序策略以 DNA 测序碱 基质量为单位、 按不同错配类型加权后的分值累加和作为罚分 值。
根据本发明的分型方法的一个实施例, 还包括步骤: 将测序 结果文件的峰图图形化显示输出, 进行峰图形态缩放调节和 /或序 列峰图连动查看, 以便于分型人员修改和 /或确认分型结果。
根据本发明的另一个方面, 还提供一种基于聚合酶链式反应 产物测序序列分型系统, 包括: 碱基序列判断子系统, 用于接收 测序结果, 根据测序结果判读杂合子位点和待分型碱基序列; 联 配位置识别子系统, 用于接收来自碱基序列判断子系统的待分型 碱基序列, 将待分型碱基序列比对到对应位点的分型数据库, 识 别待分型碱基序列和分型数据库的参考序列的联配位置关系; 罚 分值确定子系统, 用于根据待分型碱基序列和分型数据库的参考 序列的联配位置关系检索所述分型数据库中的等位基因型, 根据 定序策略获得分型数据库中的等位基因型的罚分值; 候选型别确 定子系统, 用于根据分型数据库中的等位基因型的罚分值获得候
选型别组合集。
根据本发明的分析系统的一个实施例, 还包括: 索引预处理 子系统, 用于预先建立所述分型数据库的参考序列, 以及所述参 考序列和所述分型数据库的等位基因型序列之间的位置对应关 系; 根据所述分型数据库的等位基因型中可变碱基位上的碱基符 号序列形成的哈希数组; 置上的碱基符号, 顺序排列, 然后遍历该分型数据库的哈希数 组, 进行打分。
根据本发明的分型系统的一个实施例, 还包括: 图形化显示 子系统, 用于将测序结杲文件的峰图图形化显示输出, 进行峰图 形态缩放调节和 /或序列峰图连动查看, 以便于分型人员修改和 / 或确认分型结果。 系统, 可以通过计算机等设备实现候选基因型的自动识别, 处理 速度快, 提高了分型效率。
进一步, 通过图形化显示界面等技术手段为分型人员的分型 确认和修改提供方便, 提供了分型准确率以及分型效率。 附图说明
图 1 示出才艮据本发明的 PCR-SBT分型的实现方法的一个实 施例的流程图;
图 2示出根据本发明的 PCR-SBT分型的实现方法的另一个 实施例的流程图;
图 3 示出本发明的一个应用例的数据图形化输出界面的截 图;
图 4示出根据本发明的 PCR-SBT分型系统的一个实施例结
枸图;
图 5示出才 据本发明的 PCR-SBT分型系统的另一个实施例 的结构图。 具体实施方式
下面参照附图对本发明进行更全面的描述, 其中说明本发明 的示例性实施例。 在附图中, 相同的标号表示相同或者相似的组 件或者元素。
图 1 示出本发明基于聚合酶链式反应产物测序序列分型的实 现方法的一个实施例的流程图, 以下将基于聚合酶链式反应产物 测序序列分型的实现方法简称为 PCR-SBT分型方法。
如图 1 所示, 在步骤 102, 通过计算机程序 测序结果判 读杂合子和待分型碱基序列。 例如, 测序结果包括时域下定长间 隔的荧光信号强度读数信息, 根据荧光信号强度读数确定信号峰 值, 从而确定碱基或杂合子。
在步骤 104 , 将含有杂合子的待分型碱基序列比对到对应位 点的分型数据库, 识别待分型碱基序列和对应位点的分型数据库 的参考序列的联配位置关系。 PCR扩增试验中靶位点是已知的, 待分型碱基序列对应的目标位点也是已知的。 根据待分型碱基序 列的目标位点信息比对到对应位点的分型数据库中。 每个分型数 据库包括参考序列 (Reference Sequence )、 以及参考序列和分型 数据库中各个等位基因型的位置对应关系。 例如通过动态规划算 法 ( Dynamic Programmin )或点矩阵 ( Dot Matrix ) 方法实现 待分型碱基序列和对应位点的分型数据库的参考序列的联配位置 关系的识别。
在步骤 106, 结合待分型碱基序列和分型数据库的参考序列 的联配位置关系, 检索分型数据库中的各个等位基因型, 根据定
序策略获得分型数据库中的各个等位基因型的罚分值。 遍历分型 数据库中的等位基因型, 根据待分型碱基序列和参考序列的联配 位置关系、 参考序列和等位基因型的位置对应关系, 获得待分型 碱基序列和等位基因型序列的位置对应关系, 然后结合定序策略 获得各个等位基因型的罚分值。 定序策略可以按待分型碱基序列 与比较的等位基因型 (标的等位基因型)之间的错配位置上的错 配类型、 待分型碱基序列在该位置的质量值罚分的累加和作为定 序依据
在步骤 108, 根据分型数据库中的等位基因型的罚分值获得 候选型别组合集。 例如, 根据设定的罚分阈值选取 TopN 个候选 型别组合集; 或者选取第一个候选型别作为确定的等位基因分 型。
本发明实施例的 PCR-SBT 分型的实现方法, 通过计算设备 等自动完成待分型基因序列的分型, 特别是对于大范围候选型别 筛查搜索时, 速度快, 效率高。
根据本发明的 PCR-SBT 分型的实现方法的一个实施例, 预 先对分型数据库中的等位基因型序列信息进行预处理, 将对应的 等位基因型中的有变异的碱基(简称变异碱基)对应的位置上的 碱基符号取出, 将变异碱基符号顺序排列、 编码, 形成哈希数 组。 例如, 哈希数组的存储格式为: 键值-对数据库中各个型别 在所有可变碱基位置上对应的碱基符号序列按规则编码成 2进制 码后的值。 对于待分型碱基序列, 取出变异碱基对应位置上的碱 基符号, 顺序排列, 然后根据哈希数组的键值 (key )遍历等位 基因型的哈希数组, 对待分型碱基序列和分型数据库中的等位基 因型序列进行打分, 根据定序策略获得各个等位基因型的罚分 值, 然后根据罚分阔值得到 TopN个候选型。
预先进行建库处理形成分型数据库的等位基因型的变异碱基
的哈希数组, 哈希数组建好后可以驻留内存, 可以多次重复对不 同的待分型碱基序列进行处理, 检索速度快, 大大提高了检索效 率, 从而提高了本发明整个分型方法的速度和效率。
根据本发明的 PCR-SBT 分型的实现方法的一个实施例, 在 分型结果的筛选定序中采用以 DNA 测序碱基质量为单位、 按不 同错配类型加权后的分值累加和作为罚分值, 并以该罚分由小到 大的规则定序输出, 获得候选基因型列表。 例如, 定序策略采用 如下规则: 设错配位点的质量值为 q, 待分型碱基序列 Q与标的 基因型 T, 贝, J :
( 1 )缺失位, +1;
( 2 ) 非缺失位错配, 基础罚分为 + q;
( 3 ) 非缺失位错配, 如果 Q为纯合子, 且不为 T中对应位 置的简并碱基所包含, 则 +2q;
( 4 ) 非缺失位错配, 如果 Q为杂合子, 且为 T中对应位置 的碱基所包含, 则 +2q;
( 5 ) 非缺失位错配, 如果 Q为杂合子, 且不被 T中对应位 置的简并碱基所包含, 则 +3q。
本领域的技术人员根据本发明的上述例子, 能够设计出多种 相似或者等同的定序策略, 同样属于本发明的保护范围。
通常碱基质量值代表该碱基 Base Calling结果出错的概率, 该值越高, 出错的概率越高, 所以, 以碱基质量作为罚分的基础 值 (或者参考值), 可以提高分型的准确性。 本发明实施例的分 型方法, 相较于现有的仅以错配碱基数量为考量因子定序的方 法, 更加有效地将最接近真实基因型的候选结果放置在候选列表 的最优先位置, 从而提高了分型效率。
图 2示出本发明 PCR-SBT分型的实现方法的另一个实施例 的流程图。
如图 2 所示, 在步骤 202, 输入测序结果文件。 例如, 测序 结果文件的内容和格式可以参见
【http:〃 www.appliedbiosystems.com/support/software_comm unity/ABIF File— Format.pdf 】;
在步骤 204, 根据测序结果文件自动判读杂合子位点和碱基 序列。 例如, 通过 Reference-Based Base Calling (有参考序列的 碱基识别) 方法通过计算机自动进行碱基及杂合子判读。
在步骤 206, 待分型碱基序列和对应位点的分型数据库的参 考序列的自动联配及自动插入 /删除识别。 例如, 在参考文献中 " Algorithmic Bioinformatics, Daniel Huson, 25, Oktober, 2005" 中介绍了多种比对算法 (Alignment Algorithm )0 可以采 用有界的全局比对算法(Banded Global Alignment algorithm ) 进行自动联配。
在步骤 208, 检索已经建立好索引的分型数据库, 搜集候选 等位基因型别组合集。
在步骤 210, 根据罚分大小自动定序的候选等位基因型列 表;
在步骤 212 , 数据图形化显示输出 ώ 测序结果文件中记录时 域下定长间隔的荧光信号强度读数, 与测序过程时域相对应; 荧 光信号有四种颜色, 分别对应四种碱基, 通过按特定步长和曲线 拟合公式, 可以绘制出荧光信号在整个时域内的变化过程。 这样 的图由计算机程序解析测序结果后绘制并显示。 将测序结果中的 峰图信号波形图与对应的碱基判读和杂合子识别结果, 以及由测 序结果中通过自动判读得到的碱基序列, 连同目标位点的选定候 选基因型序列按照联配位置关系整体显示在同一视图窗体中。
在步骤 214 , 峰图形态缩放调节及序列峰图连动查看。 对峰 图的显示支持单维度放大缩小功能, 提供更佳的可用性。 通过待
分型碱基序列和候选基因型序列比对, 整理得到的序列异同信 息, 将错配位置绘制成位点整体的概览视图, 并支持在这些位置 之间于各个窗体视图上的同步跳转。 具体实现如下:
序列峰图连动查看: 通过程序将多个测序结果文件对应的峰 图在面板上显示出来后, 将根据每个测序结果文件获得的待分型 碱基序列和参考序列的位置对应关系为基础, 从而得到各个待分 型碱基序列之间的位置对应关系。 不同待分型碱基序列上碱基之 间的位置是相对的, 每一次触发一个碱基, 程序首先会根据该碱 基去查找它对应其他序列上对应的碱基位置, 加上碱基的偏移 值, 并且重新画出完整的峰图。 每次触发后, 确保本次触发的碱 基是对齐的, 其他位置上的序列和峰图只保证大体上对齐即可。
峰图形态缩放调节: 峰图的放大和缩小可以根据重绘图时修 改峰图的参数每次来设置峰图放大和缩小的比例以实现峰图的放 大和缩小。 峰图的拉升和挤压可以根据设置峰图在横向的参数来 实现峰图的拉升和挤压。
在步驟 216, 获得分型结果及数据备份。 分型人员的工作包 括参考峰图排查假纯合及确认或修改计算机自动识别出的杂合位 点, 并通过参考罕见型别列表, 对当前待分型序列做修改和最终 确认。 这部分目前需要受过专业培训的有经验的专业人员的人眼 识别和把握尺度。 分型人员还需要在疑似新基因的结果(高质量 位出现与当前数据库无匹配的情况) 出现时, 做出确认; 在模糊 结果(可能是 A基因型, 也符合 B基因型) 出现时, 给出 GSSP 引物, 以备应用 SSP分型技术做附加确认。 另外新基因的发现也 是需要人去整理确认的。 对于数据备份, 在分型人员保存文件的 同时, 软件会将该分型文件拷贝到软件指定的文件夹下, 保存分 型文件的同时在该文件夹下产生一个临时文件, 主要是记录分型
员分型时的信息, 如: 在某一个位置对某一个碱基做了修改, 截取的峰图有效范围, 分型的结果等。 这样当用户下次想再查看 之前校正后的峰图文件时, 可以直接打开分型历史记录面板, 根 据曰期来查看校正后的文件。
多个测序结果对应的峰图序列文件打开后, 软件可以根据算 法实现特定碱基和序列之间的对齐和联动, 并且在选择峰图时, 序列能保证同步和峰图对应。 实现了序列的自动对齐, 不用人工 进行矫正, 缩短了用人眼去辨别序列位置的时间, 大大的提高了 分型工作者的效率。 通过实现峰图的放大, 缩小以及峰值放大和 缩小, 当上下峰图对齐的效果不是很模糊的时候, 可以通过对峰 图的放大和缩小来调整峰图的大小, 便于分型人员查看峰图, 提 高了分型效率。
每一个峰图文件分型结果的备份和恢复机制。 当分型人员通 过修改碱基, 屏蔽序列分型完后, 根据修改后的文件产生新分型 的结果, 将保存在一个临时文件夹里, 主要是方便分型人 对已 阅峰图的核对或是查看等。
图 3 示出本发明的一个应用例的数据图形化输出界面的截 图。 如图 3所示, 该应用例中 SBT分型软件主界面分为四部分: 左上部分 31, 右上部分 32, 左下部分 33, 右下部分 34。 其中, 左上部分 31 为分型文件的选择部分。 分型者选择的分型文件, 都会在该部分以树形结构的形式来展现, 方便用户来选择分型文 件。 左下部分 33 根据选择的分型文件的文件名称, 识别该分型 文件的比对位点, 根据分型的位点选择比对的分型数据库, 并根 据候选型别罚分由小到大的定序规则排序, 即得出左下角的所有 配型列表。 右上部分 32 为选择分型文件序列的显示部分。 其 中, 第一行 Consensus 行为位点的比对库全序列; 第二行 Forward 行为选择分型文件的正向序列; Reverse 为选择分型文
件的反向序列; 第三行 Pattern 行为正反向序列的匹配序列结 果; 最下面的 2条序列为从左下角得来的选择列表。 右下部分 34 为峰图和序列文件的显示部分。 峰图上方为该峰图对应的序列。 每个波峰都对应了一个碱基, 上下峰图的对应是根据分型文件和 数据库的比对结果来对应的, 峰图的上下的波峰和波峰, 峰图的 上下的碱基和碱基, 都是对应的。
图 4 示出本发明实施例的一种 PCR-SBT 分型系统的结构 图。 如图 4 所示, 该实施例的分型系统包括碱基序列判断子系统 41、 联配位置识别子系统 42、 罚分值确定子系统 43 和候选型别 确定子系统 44。 其中, 碱基序列判断子系统 41 用于接收测序结 果, 根据测序结果判读杂合子位点和待分型碱基序列。 联配位置 识别子系统 42用于接收来自碱基序列判断子系统 41的待分型碱 基序列, 将待分型碱基序列比对到对应位点的分型数据库中, 识 别待分型碱基序列和对应位点的分型数据库的参考序列之间的联 配位置关系。 例如, 联配位置识别子系统 42 通过动态规划算法 或者点矩阵方法识别待分型碱基序列和分型数据库的参考序列的 联配位置关系。 罚分值确定子系统 43 用于根据所述联配位置识 别子系统 42 识别的待分型碱基序列和分型数据库的参考序列的 联配位置关系检索分型数据库中的等位基因型, 根据定序策略获 得分型数据库中的各个等位基因型的罚分值。 例如, 定序策略以 DNA测序碱基质量为单位、 按不同错配类型加权后的分值累加和 作为罚分。 候选型别确定子系统 44, 用于才 据分型数据库中的各 个等位基因型的罚分值获得候选型别组合集。
根据本发明的一个实施例, 分型系统还可选地包括索引预处 理子系统 45。 索引预处理子系统 45 用于预先建立分型数据库的 参考序列 , 以及分型数据库的参考序列和分型数据库的各个等位 基因型序列之间的位置对应关系; 此外, 索引预处理子系统 45
还将分型数据库中的等位基因型中的有变异的碱基(简称变异碱 基)对应的位置上的碱基符号取出, 将变异碱基符号顺序排列、 编码, 形成哈希数组, 即根据分型数据库的等位基因型中变异碱 基的碱基符号形成的哈希数组。 罚分值确定子系统 43 从待分型 碱基序列中取出变异碱基对应位置上的碱基符号, 顺序排列, 然 后遍历该分型数据库的哈希数组, 进行联配。
图 5示出本发明实施例的另一种 PCR-SBT分型系统的结构 图。 如图 5 所示, 该实施例的分型系统包括碱基序列判断子系统 51、 联配位置识别子系统 52、 罚分值确定子系统 53、 候选型别 确定子系统 54、 索引预处理子系统 55、 图形化显示子系统 56和 数据备份系统 57。 其中, 碱基序列判断子系统 51、 联配位置识 别子系统 52、 罚分值确定子系统 53、 候选型别确定子系统 54和 索引预处理子系统 55 可以参见上文实施例中对应子系统的描 述, 为简洁起见在此不再详细描述。 图形化显示子系统 56 用于 将测序结果文件的峰图图形化显示输出, 进行峰图形态缩放调节 和 /或序列峰图连动查看, 以便于分型人员修改和 /或确认分型结 果。 图形化显示子系统 56 将测序结果中的峰图信号波形图与对 应的碱基判读和杂合子识别结果, 以及由测序结果中通过自动判 读得到的碱基序列, 连同目标位点的选定候选基因型序列按照联 配位置关系整体显示在同一视图窗体中。 数椐备份系统 57 用于 存储和备份确认的分型结果, 以及分型人员所作的修改等信息。 对于数据备份, 在分型人员保存文件的同时, 软件会将该分型文 件拷贝到软件指定的文件夹下, 保存分型文件的同时在该文件夹 下产生一个临时文件, 主要是记录分型人员分型时的信息, 如: 在某一个位置对某一个碱基做了修改, 截取的峰图有效范围, 分 型的结果等。
需要指出, 本发明实施例中的各个子系统, 可以作为单独的
设备或者装置存在, 通过相互配合和协作一起构成分型系统, 例 如各个子系统以分布式的方式存在; 也可以多个或者所有的子系 统集成在同一设备上。
现有软件分析速度一般为 10 个 /小时, 本发明的方案可以达 到 15个 /小时, 现有软件准确率一般为 90% , 本发明的方案可以 达到 92%。 与现有技术的其他厂家同类产品比较, 本发明的方法 和系统在单位时间内分型速度可以提高 50%, 碱基识别的准确率 可以提高 2%。
本发明提供的 SBT分型方法和系统, 可以通过计算机等设备 实现候选基因型的自动识别, 从而提高了分型效率; 通过图形化 显示界面等技术手段为分型人员的分型确认和修改提供方便, 提 供了分型准确率以及分型效率。
需要指出, 本发明的 SBT分型方法和系统, 不仅可以应用于 HLA 分型, 同样可以应用于 HPV ( Human papillomavirus, 人 乳头瘤病毒)、 HBV ( hepatitis B virus, 乙型肝炎病毒) 等其他 分型的实现。 理论上, 在有分型数据库支持的条件下, 本发明可 以应用到任何有分型需求的物种上的。
本发明的描述是为了示例和描述起见而给出的, 而并不是无 遗漏的或者将本发明限于所公开的形式。 很多修改和变化对于本 领域的普通技术人员而言是显然的。 选择和描述实施例是为了更 好说明本发明的原理和实际应用, 并且使本领域的普通技术人员 能够理解本发明从而设计适于特定用途的带有各种修改的各种实 施例。
Claims
1. 一种基于聚合酶链式反应产物测序序列分型的实现方法, 其特 征在于, 包括:
通过计算机程序根据测序结果判读杂合子位点和待分型碱基序 列;
将含有杂合子的所述待分型碱基序列比对到对应位点的分型数据 库, 识别所述待分型碱基序列和所述分型数据库的参考序列的联 配位置关系; 置关系检索所述分型数据库中的等位基因型, 根据定序策略获得 所述分型数据库中的等位基因型的罚分值;
根据所述分型数据库中的等位基因型的罚分值获得候选型别组合 集。
2. 根据权利要求 1 所述的实现方法, 其特征在于, 根据所述待 分型碱基序列和所述分型数据库的参考序列的联配位置关系检索 所述分型数据库中的等位基因型的步骤包括:
从待分型碱基序列中取出变异碱基对应位置上的碱基符号, 顺序 排列, 然后遍历预先建立的所述分型数据库中等位基因型中变异 碱基的碱基符号形成的哈希数组, 进行打分。
3. 根据权利要求 1 所述的实现方法, 其特征在于, 所述定序策 略以 DNA 测序碱基质量为单位、 按不同错配类型加权后的分值 累加和作为罚分。
4. 根据权利要求 1 所述的实现方法, 其特征在于, 所述定序策 略为:
假定错配位点的廣量值为 q, 待测序列 Q与标的基因型 T, 则: ( 1 )缺失位, +1; ( 2 ) 非缺失位错配, 基础罚分为 + q;
( 3 )非缺失位错配, 如果 Q为纯合子, 且不为 T中对应位置的 简并碱基所包含, 则 +2q;
( 4 ) 非缺失位错配, 如果 Q为杂合子, 且为 T中对应位置的碱 基所包含, 则 +2q;
( 5 ) 非缺失位错配, 如果 Q为杂合子, 且不被 T中对应位置的 简并碱基所包含, 则 +3q。
5. 根据权利要求 1 所述的实现方法, 其特征在于, 通过动态规 划算法或者点矩阵方法识别所述待分型碱基序列和所述分型数据 库的参考序列的联配位置关系。
6. 根据权利要求 1 所述的实现方法, 其特征在于, 还包括步 骤:
将测序结果文件的峰图图形化显示输出, 进行峰图形态缩放调节 和 /或序列峰图连动查看。
7. 根据权利要求 6 所述的实现方法, 其特征在于, 还包括步 骤:
自动存储分型人员的修改和 /或分型结果。
8. —种基于聚合酶链式反应产物测序序列分型系统, 其特征在 于, 包括:
碱基序列判断子系统, 用于接收测序结果, 根据所述测序结果判 读杂合子位点和待分型碱基序列;
联配位置识别子系统, 用于接收来自所述减基序列判断子系统的 待分型碱基序列, 将所述待分型碱基序列比对到对应位点的分型 数据库, 识别所述待分型碱基序列和所述分型数据库的参考序列 的联配位置关系;
罚分值确定子系统, 用于根据所述待分型碱基序列和所述分型数 据库的参考序列的联配位置关系检索所述分型数据库中的等位基 因型, 根据定序策略获得所述分型数据库中的等位基因型的罚分 值;
候选型别确定子系统, 用于根据所述分型数据库中的等位基因型 的罚分值获得候选型别组合集。
9. 根据权利要求 8所述的分型系统, 其特征在于, 还包括: 索引预处理子系统, 用于预先建立所述分型数据库的参考序列, 置对应关系; 根据所述分型数据库的等位基因型中变异碱基的碱 基符号形成的哈希数组;
所述罚分值确定子系统从待分型碱基序列中取出变异碱基对应位 置上的碱基符号, 顺序排列, 然后遍历所述分型数据库的哈希数 組, 进行打分。
10. 根据权利要求 8或 9所述的分型系统, 其特征在于, 所述定 序策略以 DNA 测序碱基质量为单位、 按不同错配类型加权后的 分值累加和作为罚分值。
11. 根据权利要求 8 所述的分型系统, 其特征在于, 联配位置识 别子系统通过动态规划算法或者点矩阵方法识别所述待分型碱基 序列和所述分型数据库的参考序列的联配位置关系。
12. 根据权利要求 8所述的分型系统, 其特征在于, 还包括: 图形化显示子系统, 用于将测序结果文件的峰图图形化显示输 出, 进行峰图形态缩放调节和 /或序列峰图连动查看。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201010117703.4 | 2010-03-04 | ||
| CN2010101177034A CN101984445B (zh) | 2010-03-04 | 2010-03-04 | 一种基于聚合酶链式反应产物测序序列分型的实现方法和系统 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2011106994A1 true WO2011106994A1 (zh) | 2011-09-09 |
Family
ID=43641614
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2011/000347 Ceased WO2011106994A1 (zh) | 2010-03-04 | 2011-03-03 | 基于聚合酶链式反应产物测序序列分型的实现方法和系统 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN101984445B (zh) |
| WO (1) | WO2011106994A1 (zh) |
Families Citing this family (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102321749B (zh) * | 2011-08-11 | 2013-02-06 | 中南大学 | 一种micb基因分型的pcr-sbt方法及试剂盒 |
| CN102750461B (zh) * | 2012-06-14 | 2015-04-22 | 东北大学 | 一种可得到完全解的生物序列局部比对方法 |
| KR101482011B1 (ko) * | 2012-10-29 | 2015-01-14 | 삼성에스디에스 주식회사 | 염기 서열 정렬 시스템 및 방법 |
| KR101508816B1 (ko) * | 2012-10-29 | 2015-04-07 | 삼성에스디에스 주식회사 | 염기 서열 정렬 시스템 및 방법 |
| WO2014152541A1 (en) * | 2013-03-15 | 2014-09-25 | Sherwin Han | Spatial arithmetic method of sequence alignment |
| CN103617375B (zh) * | 2013-12-02 | 2017-08-25 | 深圳华大基因健康科技有限公司 | 聚合酶链式反应产物测序分型的方法及系统 |
| CN104263850B (zh) * | 2014-06-19 | 2017-06-06 | 重庆医科大学 | 基于SNaPshot技术的小鼠肝炎病毒分型检测方法及试剂盒 |
| JP6884143B2 (ja) | 2015-10-21 | 2021-06-09 | コーヒレント・ロジックス・インコーポレーテッド | 階層的転置索引表を使用したdnaアラインメント |
| CN108241792B (zh) * | 2016-12-23 | 2021-03-23 | 深圳华大基因科技服务有限公司 | 一种整合多平台基因分型结果的方法和装置 |
| CN108624671B (zh) * | 2017-03-20 | 2022-02-01 | 深圳华大基因股份有限公司 | 用于hla分型的基因型序列 |
| CN108660198B (zh) * | 2018-05-15 | 2022-02-22 | 广州血液中心 | 一种血小板膜蛋白cd36抗原基因分型的pcr-sbt方法及试剂 |
| CN109753939B (zh) * | 2019-01-11 | 2021-04-20 | 银丰基因科技有限公司 | 一种hla测序峰图识别方法 |
| CN110706746B (zh) * | 2019-11-27 | 2021-09-17 | 北京博安智联科技有限公司 | 一种dna混合分型数据库比对算法 |
| CN112102883B (zh) * | 2020-08-20 | 2023-12-08 | 深圳华大生命科学研究院 | 一种fastq文件压缩中的碱基序列编码方法和系统 |
| CN113380323B (zh) * | 2021-07-19 | 2022-09-23 | 浙江迪谱诊断技术有限公司 | Sanger测序峰图截取标识方法、系统、计算机设备及存储介质 |
| CN114023379B (zh) * | 2021-12-31 | 2022-05-13 | 浙江迪谱诊断技术有限公司 | 一种确定基因型的方法及装置 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO1999007883A1 (en) * | 1997-08-11 | 1999-02-18 | Visible Genetics Inc. | Method and kit for hla class i typing dna |
| WO2008049021A2 (en) * | 2006-10-17 | 2008-04-24 | Life Technologies Corporation | Methods of allele typing |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1680589A (zh) * | 2003-06-06 | 2005-10-12 | 李志广 | 基因芯片用人类白细胞抗原分型探针的筛选及其应用方法 |
| CN1896284B (zh) * | 2006-06-30 | 2013-09-11 | 博奥生物有限公司 | 一种鉴别等位基因类型的方法 |
| CN101654691B (zh) * | 2009-09-23 | 2013-12-04 | 深圳华大基因健康科技有限公司 | Hla基因扩增和基因分型方法及其相关引物 |
-
2010
- 2010-03-04 CN CN2010101177034A patent/CN101984445B/zh active Active
-
2011
- 2011-03-03 WO PCT/CN2011/000347 patent/WO2011106994A1/zh not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO1999007883A1 (en) * | 1997-08-11 | 1999-02-18 | Visible Genetics Inc. | Method and kit for hla class i typing dna |
| WO2008049021A2 (en) * | 2006-10-17 | 2008-04-24 | Life Technologies Corporation | Methods of allele typing |
Non-Patent Citations (4)
| Title |
|---|
| CHENG, LIANGHONG ET AL.: "Application of heterozygous ambiguity resolution primers resolving ambiguous genotyping results of human leukocyte antigen genes.", CHIN J LAB MED., vol. 32, no. 1, January 2009 (2009-01-01), pages 40 - 43 * |
| CHENG, LIANGHONG ET AL.: "Application Value of Allele Frequencies in Direct Identification of Ambiguous HLA Genotypes.", JOURNAL OF EXPERIMENTAL HEMATOLOGY., vol. 17, no. 2, 2009, pages 487 - 492 * |
| ELLEXSON-TURNER, M.E. ET AL.: "Sequence-based typing of HLA class I alleles in Alaskan Yupik Eskimo.", HUMAN IMMUNOLOGY., vol. 62, 2001, pages 639 - 644 * |
| SWELSEN, WENDY T.N. ET AL.: "Sequence-Based Typing of the HLA-A10/A19 Group and Confirmation of a Pseudogene Coamplified With A*3401.", HUMAN IMMUNOLOGY., vol. 66, 2005, pages 535 - 542 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN101984445B (zh) | 2012-03-14 |
| CN101984445A (zh) | 2011-03-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2011106994A1 (zh) | 基于聚合酶链式反应产物测序序列分型的实现方法和系统 | |
| EP2718862B1 (en) | Method for assembly of nucleic acid sequence data | |
| CN103221551B (zh) | Hla基因型别-snp连锁数据库、其构建方法、以及hla分型方法 | |
| CN102682224B (zh) | 检测拷贝数变异的方法和装置 | |
| CN112466395B (zh) | 基于snp多态性位点的样本识别标签筛选方法与样本识别检测方法 | |
| EP2981921A1 (en) | Methods and processes for non-invasive assessment of genetic variations | |
| CN108220403B (zh) | 特定突变位点的检测方法、检测装置、存储介质及处理器 | |
| US20200105370A1 (en) | Genome browser | |
| US12272431B2 (en) | Detecting false positive variant calls in next-generation sequencing | |
| US20140162260A1 (en) | Primers, snp markers and method for genotyping mycobacterium tuberculosis | |
| CN112669903A (zh) | 基于Sanger测序的HLA分型方法及设备 | |
| CN109524060B (zh) | 一种遗传病风险提示的基因测序数据处理系统与处理方法 | |
| WO2017139945A1 (zh) | 分型方法和装置 | |
| KR101539737B1 (ko) | 유전체 정보와 분자마커를 이용한 여교잡 선발의 효율성 증진 기술 | |
| EP1635276B1 (en) | Display method and display apparatus of gene information | |
| CN117542410A (zh) | 肺癌基因组多类型变异的知识图谱致癌性表示预测方法 | |
| CN116209777B (zh) | 基于无创产前基因检测数据的亲缘关系判定方法和装置 | |
| WO2023049558A1 (en) | A graph reference genome and base-calling approach using imputed haplotypes | |
| JP4994676B2 (ja) | 遺伝子多型解析支援プログラム、該プログラムを記録した記録媒体、遺伝子多型解析支援装置、および遺伝子多型解析支援方法 | |
| CN111584003A (zh) | 病毒序列整合的优化检测方法 | |
| CN111128297B (zh) | 一种基因芯片的制备方法 | |
| EP1634964B1 (en) | Method for determining protein binding sites in genomic DNA | |
| CN103617375B (zh) | 聚合酶链式反应产物测序分型的方法及系统 | |
| KR20250020489A (ko) | 유전자 변이체를 식별하기 위한 방법 및 시스템 | |
| CN120452533A (zh) | 基于ngs数据识别hla基因遗传变异的方法、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 11750147 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 11750147 Country of ref document: EP Kind code of ref document: A1 |