WO2014075274A2 - 一种主要组织相容性复合体mhc分型方法及其应用 - Google Patents

一种主要组织相容性复合体mhc分型方法及其应用 Download PDF

Info

Publication number
WO2014075274A2
WO2014075274A2 PCT/CN2012/084689 CN2012084689W WO2014075274A2 WO 2014075274 A2 WO2014075274 A2 WO 2014075274A2 CN 2012084689 W CN2012084689 W CN 2012084689W WO 2014075274 A2 WO2014075274 A2 WO 2014075274A2
Authority
WO
WIPO (PCT)
Prior art keywords
sequence
module
mhc
snp
indel
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2012/084689
Other languages
English (en)
French (fr)
Inventor
张涛
曹红志
王煜
仝欣
刘小敏
王俊
汪建
杨焕明
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BGI Shenzhen Co Ltd
Original Assignee
BGI Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BGI Shenzhen Co Ltd filed Critical BGI Shenzhen Co Ltd
Priority to PCT/CN2012/084689 priority Critical patent/WO2014075274A2/zh
Priority to CN201280076912.6A priority patent/CN104769129B/zh
Publication of WO2014075274A2 publication Critical patent/WO2014075274A2/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6881Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for tissue or cell typing, e.g. human leukocyte antigen [HLA] probes
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2600/00Oligonucleotides characterized by their use
    • C12Q2600/156Polymorphic or mutational markers

Definitions

  • the present invention is in the field of bioinformatics, and in particular, the present invention relates to a major histocompatibility complex MHC typing method and its use. Background technique
  • the major histocompatibility complex is a tightly linked group of genes encoding major histocompatibility antigens on a chromosome of a vertebrate, associated with immune responses, immune regulation, and transplant rejection. . Since human major histocompatibility antigens are first discovered on the surface of leukocytes, they are called human leucocyte antigen (HLA), and human MHC, the gene group encoding HLA, is called HLA complex.
  • HLA human leucocyte antigen
  • Human leukocyte antigen is a genomic region most closely related to immunity in humans. It is located on the short arm of human chromosome 6, and consists of a series of closely linked loci.
  • the HLA gene is one of the most complex genetic systems in humans with the highest polymorphism in the human genome. The HLA gene also plays an important role in recognizing autologous and non-corporeal and regulating immune responses.
  • HLA typing determines the type of allele at each locus of the HLA gene.
  • HLA serological typing There are many methods for HLA typing, the earliest methods for HLA serological typing and cytological typing. Now mainly based on DNA level-based typing methods, including single-strand conformation polymorphism (PCR-SSCP), sequence-specific oligonucleotide probes (PCR-SSO(P)), restriction fragment length polymorphisms (PCR-RFLP, gene chip, sequence-specific primers (PCR-SSP) and sequence-based typing (SBT).
  • PCR-SSCP single-strand conformation polymorphism
  • PCR-SSO(P) sequence-specific oligonucleotide probes
  • PCR-RFLP restriction fragment length polymorphisms
  • gene chip PCR-specific primers
  • SBT sequence-based typing
  • a method for constructing a main histocompatibility complex MHC type database comprising the steps of: (1) aligning the target MHC type sequence with a reference sequence to obtain a difference site of the target MHC type relative to the reference sequence, the difference site being a SNP and/or an InDel site;
  • the reference sequence is derived from hgl 8 or hgl9.
  • the target MHC type sequence is a known sequence from a database (e.g., an IMGT database).
  • a primary histocompatibility complex MHC type database is provided, the database being constructed using the method described in the first aspect.
  • the structure of the database is: the first column indicates different types of MHC, the second column indicates the sequence corresponding to the MHC type, and the third column indicates the MHC type sequence after the reference sequence is aligned.
  • the SNP indicates the InDel after the MHC type is aligned.
  • a unit for constructing a main histocompatibility complex MHC type database comprising a module:
  • the output module outputs the SNP and InDel locus information of the MHC type relative to the reference sequence by using the comparison result of the target MHC type sequence and the reference sequence obtained by the comparison module.
  • the unit further includes: (3) a sequence acquisition module, configured to obtain a target MHC type sequence and a reference sequence.
  • a SNP and InDel detection method comprising the steps of:
  • step (2) (2) re-aligning the SAM file obtained in step (1), obtaining the best type of each reading order, the initial position of the corresponding type, and the number of mismatches with the best type difference;
  • the steps (1M) are compared using BWA software.
  • step (1) and the step (2) further includes: a step of converting the SAM file into an FQ file.
  • step (2) is compared using BWA software.
  • step (2) uses the MHC type sequence corresponding to the target gene as a reference sequence.
  • the MHC database described in the step (2) is prepared by the method described in the first aspect.
  • a SNP and InDel detecting unit comprising a module:
  • comparison module used for comparison reading
  • a SNP and InDel splicing method including the steps of:
  • step (3) splicing the read sequence selected in step (2) with the sequence of step (1) to obtain a longer sequence
  • step (4) aligning and filtering the spliced sequence obtained in step (4) with the MHC type database, and randomly combining and filtering the sequences that do not span the entire exon to obtain a filtered sequence;
  • step (5) The filtered sequence obtained in step (5) is sorted according to the number of read order support.
  • a splicing unit of a SNP and an InDel comprising a module:
  • a sequence acquisition module for selecting and/or obtaining a read sequence
  • a splicing module for splicing the read sequence and the sequence obtained by the sequence acquisition module;
  • aligning and filtering module for comparing and filtering the read order obtained by the splicing module;
  • Output module used to output the stitching information of SNP and InDel.
  • a method for MHC typing of a major histocompatibility complex comprising the steps of: (1) Obtaining the sequence (hap) of the gene of interest and its corresponding MHC type, the sequence is as shown in Formula I:
  • G is the gene type, i is the exon number, j is the corresponding sequence number, n is the number of exons, and i, j and n are positive integers;
  • the combination of the largest total haplotype and the most trusted haplotype is the best combination
  • step (3) Based on the best combination of step (3), the MHC classification information is obtained.
  • a primary histocompatibility complex MHC typing unit comprising a module:
  • a sequence acquisition module for selecting and/or obtaining a read sequence
  • a primary histocompatibility complex MHC typing system comprising:
  • FIG. 1 shows the MHC classification process.
  • the inventors have established a major histocompatibility complex for the first time through extensive and in-depth research.
  • the present invention provides a major histocompatibility complex MHC type database and its construction method and building unit, SNP and InDel detection method and detection unit, SNP and InDel splicing method and splicing unit, and main tissue compatibility Sexual complex MHC typing method and its unit and system.
  • the method and system of the invention have high accuracy, the data to be tested is relatively low, and the typing area is greatly improved compared to the existing typing method.
  • major histocompatibility complex and “MHC” are used interchangeably and refer to a group of closely linked genes encoding major histocompatibility antigens present on a chromosome of a vertebrate. Groups, MHC are associated with immune response, immune regulation, and transplant rejection.
  • HLA Human leukocyte antigen
  • human leucocyte antigen HLA
  • human MHC the gene group encoding HLA
  • HLA complex Human leukocyte antigen HLA is a genomic region most relevant to immunity in humans. It is located on the short arm of human chromosome 6, and consists of a series of closely linked loci.
  • the HLA gene is one of the most complex genetic systems in humans with the highest polymorphism in the human genome. The HLA gene also plays a vital role in recognizing autologous and non-corporeal and regulating immune responses. Gene, exon
  • the term "gene” refers to the basic unit of biological inheritance that exists within the region of the gene on the genome.
  • genes are composed of introns and exons. Genes generally have multiple exons.
  • a gene possesses multiple transcripts, each transcript being a different combination of exons of the gene, even reducing a few bases in the exon of the exon boundary, or extending a few bases to the intron. Base, this is called alternative splicing.
  • a gene can have multiple transcripts. Different transcripts can be obtained at different times in different environments.
  • SNP single nucleotide polymorphism
  • the polymorphisms exhibited by SNPs involve only a single base variation, which can be caused by a single base transition or transversion, or by the insertion or deletion of a base.
  • the gene fragments (including DNA and cDNA) are sequenced, and the sequenced objects are a piece of physically continuous base sequence called an insert, the length of which is called the insert size.
  • double-end sequencing is the sequencing of the two-sided base sequence of the fragment from edge to interior.
  • the sequence measured is called read and the length is called read-length.
  • the read order measured on both sides is from the same insert, and the distance between the ends is insertsize, so the pairing of readings on both sides Relationship is determined. These two readings are called Pair-end reads.
  • High-throughput sequencing of the genome enables humans to detect abnormal changes in disease-associated genes as early as possible, and to facilitate in-depth research into the diagnosis and treatment of individual diseases.
  • Those skilled in the art can typically perform high throughput sequencing using three second generation sequencing platforms: 454 FLX (Roche), Solexa Genome Analyzer (Illumina), Applied Biosystems, SOLID, and the like.
  • the common feature of these platforms is the extremely high sequencing throughput.
  • high-throughput sequencing can read 400,000 to 4 million sequences in one experiment. The read length is from 25bp depending on the platform. Up to 450 bp, so different sequencing platforms can read base numbers ranging from 1G to 14G in one experiment.
  • Solexa high-throughput sequencing includes two steps: DNA cluster formation and on-machine sequencing: a mixture of PCR amplification products is hybridized with a fixed sequencing probe immobilized on a solid phase carrier, and subjected to solid phase bridge PCR amplification to form a sequencing cluster; The sequencing cluster is sequenced by "edge synthesis-edge sequencing” to obtain a sequence of nucleic acid molecules in the sample.
  • the DNA cluster is formed by using a flow cell with a single-stranded primer attached to the surface, and the DNA fragment of the single-stranded state is immobilized on the chip by the principle that the linker sequence and the primer on the surface of the chip are complementary to each other by base complementation.
  • the fixed single-stranded DNA becomes double-stranded DNA
  • the double strand is denatured into a single strand, one end of which is anchored on the sequencing chip, and the other end is randomly and adjacent to another primer to be anchored, Forming a "bridge"; on the sequencing chip, there are tens of millions of DNA single molecules simultaneously reacting; forming a single-stranded bridge, using the surrounding primers as amplification primers, and amplifying again on the surface of the amplification chip to form a double
  • the strand, the double strand is denatured into a single strand, and becomes a bridge again.
  • the template called the next round of amplification continues to expand; after repeated rounds of 30 rounds of amplification, each single molecule is amplified 1000 times, called a single clone. DNA cluster.
  • DNA clusters were sequenced on a Solexa sequencer. During the sequencing reaction, the four bases were labeled with different fluorescence, and each base was blocked by a protected base. Only one base could be added to a single reaction. After reading the color of the reaction, the protection group is removed, and the next reaction can be continued. Thus, the exact sequence of the base is obtained.
  • Index is used to distinguish the samples and, after routine sequencing is completed, The Index section is additionally sequenced, and by index identification, up to 12 different samples can be distinguished in one sequencing channel.
  • the present invention provides an MHC type database.
  • the contents included in the database are expressed in the form of Table 1.
  • the first column indicates the different types of MHC
  • the second column indicates the sequence corresponding to the MHC type
  • the third column indicates the SNP after the MHC type sequence is aligned with the reference sequence
  • the fourth column indicates the MHC type. After the comparison, InDel.
  • T means 29796327 position SNP is T; 29796435-D-C means 29796436 position missing base C; None means no such type of variation.
  • the present invention also provides an MHC type database construction method.
  • the method includes steps a and b:
  • the contents of the MHC type sequence can be expressed in the form of Table 2.
  • the first column indicates the type
  • the second column indicates the type sequence.
  • Step b) Aligning the MHC type sequence with the reference sequence to obtain a difference site with the reference sequence, that is, obtaining the SNP and InDel relationship of each type with respect to the reference sequence, and constructing an MHC type database.
  • the present invention provides a method for obtaining SNP and InDel information for each read order based on a re-alignment strategy.
  • the method comprises the steps of:
  • the type is used as the header, and the sequence corresponding to the type is used as the reference sequence;
  • the present invention provides a splicing method based on linkage SNP and InDel.
  • the method comprises the steps of:
  • the present invention provides an MHC typing method.
  • the steps are as follows:
  • G*01 :01 :01 :01G*01 :01 :01 :02 G*01 :01 :01 :03 as an example.
  • the first two digits indicate the classification at the species level, the third and fourth places.
  • the upper number indicates the exon non-synonymous mutation, the synonymous mutation indicated by the 5th and 6th positions, and the 7th and 8th positions indicate the type of the mutation on the exon, but due to the intron
  • the study of mutations on the above is of little significance, so generally only the first three parts of the classification study.
  • G1 represents the G gene exon 1
  • -1 represents the number of sorts corresponding to hap
  • b) determines the most reliable type type, defined as type type 1, specific practice: score all type types, select The most trusted type type is type 1 and the scoring rules are as follows:
  • the score is not added; if the number of sorts corresponding to hap is 2, the score is increased by 0.5; if it is other data, the score is incremented by 1; if some of the exons have no sorted number, then the score is added 2; The lowest score is defined as haplotype 1 .
  • the degree of difference indicates that the same gene has the same hap number of the same exon, which means that different haps are taken from different parts of the original reading, so hap
  • the calculation rule if the hap number corresponding to the exon is different, add 2; if the correspondence is the same, add 1; if typel exists and type2 does not exist, then decrease 1; the last score is the best combination.
  • G1-2*G2-1*G3-1*G4-1*G5-1 & G1-2*G3-1*G4-1*G5-1 3.
  • the method and system of the present invention have high accuracy, and the data to be tested is relatively low; (2) The method and system of the present invention have a fast analysis process for information and simple operation;
  • the purpose of this example is to perform MHC typing on the resequencing data or the target region capture data by high-throughput sequencing technology to determine SNP linkage.
  • the illuminla sequencing platform is used to sequence the data, and the available readings are obtained after filtering the low-quality read number and the number of readings polled by the adapter.
  • the alignment information of the target gene is obtained from the bam file, and at the same time, the mismatched reading sequence is combined, and the original alignment information of the target gene is converted into the FQ file.
  • the FQ files are re-aligned, and the reference sequence is changed to the MHC type corresponding to the target gene. Because the MHC polymorphism is high, some of the readings are lost during the comparison, so the unmatched readings are re-run. Compare, solve the problem of polymorphism.
  • the alignment information of each reading is obtained from the SNP and InDel corresponding to the optimal type after the comparison.
  • the spliced sequence is filtered globally using the MHC type database. If the database cannot be matched, all reads are lost. The remaining hap is defined as a trusted hap.
  • G1 represents the exon of G gene
  • -1 represents the number of sorts corresponding to hap.
  • G1-2*G3-1 *G4-1 *G5-1 2 Select G1-2*G2-1*G3-1*G4-1*G5-1 for the corresponding type G*01:01:01:01G*01:01:01:02 G*01 :01 :01 :03 As the most trusted type (type).

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Immunology (AREA)
  • Analytical Chemistry (AREA)
  • Wood Science & Technology (AREA)
  • Engineering & Computer Science (AREA)
  • Molecular Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biotechnology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Cell Biology (AREA)
  • Biochemistry (AREA)
  • Microbiology (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Description

一种主要组织相容性复合体 MHC分型方法及其应用
技术领域
本发明属于生物信息学领域, 具体地, 本发明涉及一种主要组织相容性复 合体 MHC分型方法及其应用。 背景技术
主要组织相容性复合体 ( major histocompatibility complex, MHC ) 是存在 于脊椎动物某一染色体上编码主要组织相容性抗原的一组紧密连锁的基因群, 与免疫应答、 免疫调节和移植排斥等有关。 由于人类主要组织相容性抗原首先 在白细胞表面被发现, 故称其为人类白细胞抗原 (human leucocyte antigen, HLA) , 并将人类的 MHC , 即编码 HLA 的基因群称为 HLA 复合 体。
人类白细胞抗原(Human leukocyte antigen, HLA)是人体内与免疫最相关 的一段基因组区域。 它位于人类 6号染色体的短臂, 由一系列紧密连锁的基因 座构成。 HLA基因是人类基因组中多态性最高, 迄今为止人类最复杂的遗传系 统之一。 HLA基因也在识别自体与非体, 调节免疫应答等方面起至关重要的作 用。
HLA分型即确定 HLA基因每个基因座上的等位基因的型别。 目前 HLA分 型的方法有多种, 最早的为 HLA血清学分型、 细胞学分型的方法。 现在主要为 基于 DNA水平的分型方法, 包括单链构象多态性 (PCR-SSCP ) 、 序列特异性 寡核苷酸探针 (PCR-SSO(P)) 、 限制性片段长度多态性 (PCR-RFLP) 、 基因 芯片、序列特异性引物(PCR-SSP)以及基于序列分型法(sequence-based typing, SBT) 。
综上所述, 本领域迫切需要开发操作简单、 准确性高的 HLA分型方法。 发明内容
本发明的目的就是提供一种主要组织相容性复合体 MHC分型方法及其应 用。 在本发明的第一方面,提供了一种主要组织相容性复合体 MHC型别数据库 的构建方法, 包括步骤: (1) 将目标 MHC型别序列和参考序列进行比对,获得目标 MHC型别相对于 参考序列的差异位点, 所述的差异位点为 SNP和 /或 InDel位点; 和
(2) 对已获得所述的差异位点信息的各个目标 MHC型别进行汇总, 构建得 到主要组织相容性复合体 MHC型别数据库。
在另一优选例中, 所述的参考序列来自于 hgl 8或 hgl9。
在另一优选例中, 所述的目标 MHC型别序列为来自于数据库 (如 IMGT数 据库) 的已知序列。
在本发明的第二方面, 提供了一种主要组织相容性复合体 MHC型别数据 库, 所述的数据库是使用第一方面所述的方法构建的。
在另一优选例中, 所述数据库的结构为: 第一列表示 MHC的不同型别, 第 二列表示 MHC型别所对应的序列, 第三列表示 MHC型别序列比对参考序列后 的 SNP, 第四列表示 MHC型别比对后的 InDel。
在本发明的第三方面,提供了一种构建主要组织相容性复合体 MHC型别数 据库的单元, 所述单元包括模块:
(1) 比对模块, 用于比对目标 MHC型别序列和参考序列; 和
(2) 输出模块,用比对模块获得的目标 MHC型别序列和参考序列的比对结果, 输出 MHC型别相对于参考序列的 SNP和 InDel位点信息。
在另一优选例中,所述单元还包括: (3) 序列获取模块,用于获得目标 MHC 型别序列和参考序列。
在本发明的第四方面, 提供了一种 SNP和 InDel 检测方法, 包括步骤:
(1) 获得经比对的不匹配的(unmap) SAM文件和目标基因区的 SAM文件, 并且合并所述的 SAM文件;
(2) 重新比对步骤 (1)所获得的 SAM文件, 获得每条读序比对的最佳型别、 对 应型别的初始位置、 和与最佳型别差异的不匹配的数目; 和
(3) 过滤步骤 (2)所获得的读序, 结合 MHC数据库获得可信读序的 SNP和 InDel 的信息。
在另一优选例中, 步骤 (1M吏用 BWA软件进行比对。
在另一优选例中, 步骤 (1)和步骤 (2)之间还包括: 将 SAM文件转换为 FQ文件 的步骤。 在另一优选例中, 步骤 (2)使用 BWA软件进行比对。
在另一优选例中, 步骤 (2)使用目标基因对应的 MHC型别序列作为参考序列。 在另一优选例中, 步骤 (2)所述的 MHC数据库是用第一方面所述的方法制备 的。
在本发明的第五方面, 提供了一种 SNP和 InDel检测单元, 所述单元包括模 块:
(1) 比对模块, 用于比对读序;
(2) 文件合并模块, 用于合并经比对模块获得的比对的不匹配 SAM文件和目 标基因的 SAM文件;
(3) 文件转换模块, 用于转换文件合并模块获得的 SAM文件和 FQ文件; 和
(4) 输出模块, 用于输出 SNP和 /或 InDel的信息。
在本发明的第六方面, 提供了一种 SNP和 InDel的拼接方法, 包括步骤:
(1) 将目标基因的一条读序所对应的 SNP和 InDel作为起始序列;
(2) 从目标基因中选择与步骤 (1)获得的序列完全匹配的读序;
(3) 将步骤 (2)选择的读序与步骤 (1)的序列拼接, 获得更长的序列;
(4) 从目标基因中提取与步骤 (3)获得的序列完全匹配的读序,直到没有匹配的 读序, 从而获得拼接的序列;
(5) 将步骤 (4)获得的拼接的序列与 MHC型别数据库进行比对和过滤, 对没有 跨过整个外显子的序列进行随机组合和过滤, 从而获得过滤的序列; 和
(6) 对步骤 (5)获得的过滤的序列按照读序支持数进行排序。
在本发明的第七方面, 提供了一种 SNP和 InDel的拼接单元, 所述单元包括 模块:
(1) 序列获取模块, 用于选择和 /或获得读序;
(2) 拼接模块, 用于对序列获取模块获得的读序进行读序与序列的拼接; (3) 比对和过滤模块, 用于比对和过滤拼接模块获得的读序;
(4) 排序模块, 用于对比对和过滤模块获得的序列进行排序; 和
(5) 输出模块, 用于输出 SNP和 InDel的拼接信息。
在本发明的第八方面, 提供了一种主要组织相容性复合体 MHC分型方法, 包括步骤: (1) 获得目的基因的序列 (hap) 及其对应的 MHC型别, 所述序列如式 I所示:
(Gi-j) n
式 I
其中, G为基因类型, i为外显子编号, j为对应的序列排序数, n为外显子的数 目, i、 j和 n均为正整数;
(2) 对各个外显子序列对应的排序数进行打分, 确定最可信单体型 (type) , 打分规则如下:
当 j=l时, 分数 =0;
当 j=2时, 分数 =0.5;
当」=其他时, 分数 =1 ;
当没有外显子排序数时, 分数 =2;
计算各个外显子对应的分数的总和, 分数最低的为最可信单体型;
(3) 分别比较其余单体型与最可信单体型的差异度, 差异度计算规则如下: 比较其余单体型与最可信单体型对应外显子的序列排序数 (gpj值) , 当 j不同时, 分数 =2;
当 j相同时, 分数 =1 ;
当其余单体型中不存在相应的 j, 分数 = -2;
最终总分最大的单体型与最可信单体型的组合, 为最佳组合; 和
(4) 基于步骤 (3)的最佳组合, 得出 MHC的分型信息。
在本发明的第九方面, 提供了一种主要组织相容性复合体 MHC分型单元, 所述的单元包括模块:
(1) 序列获取模块, 用于选择和 /或获得读序;
(2) 排序模块, 用于从序列获取模块获得的读序中, 确定最可信单体型和最佳 单体型的组合; 和
(3) 输出模块, 用于输出 MHC的分型信息。
在本发明的第十方面, 提供了一种主要组织相容性复合体 MHC分型系统, 所述系统包括单元:
(1) 本发明第九方面所述的主要组织相容性复合体 MHC分型单元;
(2)本发明第三方面所述的构建主要组织相容性复合体 MHC型别数据库的 单元;
(3)本发明第五方面所述的 SNP和 InDel 检测单元; 和
(4)本发明第七方面所述的 SNP和 InDel的拼接单元。 应理解, 在本发明范围内中, 本发明的上述各技术特征和在下文(如实施 例)中具体描述的各技术特征之间都可以互相组合, 从而构成新的或优选的技 术方案。 限于篇幅, 在此不再一一累述。 附图说明
下列附图用于说明本发明的具体实施方案, 而不用于限定由权利要求书所 界定的本发明范围。
图 1显示了 MHC的分型流程。 具体实施方式
本发明人经过广泛而深入的研究, 首次建立了一种主要组织相容性复合体
MHC分型方法。 具体地, 本发明提供了主要组织相容性复合体 MHC型别数据 库及其构建方法和构建单元、 SNP和 InDel检测方法和检测单元、 SNP和 InDel的 拼接方法及拼接单元、 以及主要组织相容性复合体 MHC分型方法及其单元和系 统。 本发明方法和系统准确性高, 对待测的数据要求比较低、 相对于现有的分 型方法, 大大提高了分型区域。 主要组织相容性复合体 (MHC)
如本文所用, 术语"主要组织相容性复合体"与" MHC"可以互换使用, 都是 指存在于脊椎动物某一染色体上的、编码主要组织相容性抗原的一组紧密连锁的基 因群, MHC与免疫应答、 免疫调节和移植排斥等有关。 人类白细胞抗原 (HLA)
由于人类主要组织相容性抗原首先在白细胞表面被发现, 故称其为人类白细 胞抗原 (human leucocyte antigen, HLA) , 并将人类的 MHC , 即编码 HLA 的 基因群称为 HLA 复合体。 人类白细胞抗原 HLA是人体内与免疫最相关的一段基因组区域。 它位于人 类 6号染色体的短臂, 由一系列紧密连锁的基因座构成。 HLA基因是人类基因 组中多态性最高, 迄今为止人类最复杂的遗传系统之一。 HLA基因也在识别自 体与非体, 调节免疫应答等方面起至关重要的作用。 基因、 外显子
如本文所用, 术语"基因"是指是生物遗传的基本单位, 存在于基因组上的 基因区域内。 在真核生物中, 基因由内含子和外显子组成。 基因一般拥有多个 外显子。 在很多情况下, 基因拥有多个转录本, 每个转录本是该基因的外显子 的不同组合, 甚至在外显子边界向外显子内缩减若干碱基, 或者向内含子扩展 若干碱基, 这称为可变剪接。 由于这些原因, 一个基因可以拥有多个的转录本。 生物在不同的环境不同的时间, 可以获得不同的转录本。
SNP
如本文所用, 术语" SNP"或"单核苷酸多态性" 可以互换使用, 是指在基因 组水平上由单个核苷酸的变异所引起的 DNA序列的多态性。 SNP是可遗传变异 中最常见的一种, 占所有已知多态性的绝大多数。
SNP所表现的多态性只涉及到单个碱基的变异,这种变异可由单个碱基的 转换 (transition)或颠换 (transver sion)所引起, 也可由碱基的插入或缺失所致。
InDel
如本文所用, 术语" InDel"和"插入缺失突变 "可以互换使用, 是指涉及核苷 酸插入和 /或缺失的突变。 双末端测序
对基因片段 (包括 DNA和 cDNA)进行测序, 其测序对象都是一段物理连续 的碱基序列片段, 该片段称为插入片段, 其长度称为插入片段长度 (insertsize )。
如本文所用, 术语"双末端测序"是对该片段的两侧碱基序列从边缘向内部 的测序, 测得的序列称为读序 (read) , 长度称为读长 (read-length)。 两侧测得的读 序是来自于同一个插入片段, 并且其末端距离为 insertsize , 故两侧读序的配对 关系确定。 这两个读序被称为配对读序 (Pair-end reads)。 高通量测序
基因组的高通量测序使得人类能够尽早地发现与疾病相关基因的异常变 化, 有助于对个体疾病的诊断和治疗进行深入的研究。 本领域技术人员通常可 以采用三种第二代测序平台进行高通量测序: 454FLX(Roche公司)、 Solexa Genome Analyzer(Illumina公司)禾卩 Applied Biosystems 公司的 SOLID等。 这些平 台共同的特点是极高的测序通量, 相对于传统测序的 96道毛细管测序, 高通量 测序一次实验可以读取 40万到 400万条序列,根据平台的不同,读取长度从 25bp 到 450bp不等, 因此不同的测序平台在一次实验中, 可以读取 1G到 14G不等的碱 基数。
Solexa 高通量测序包括 DNA簇形成和上机测序两个步骤: PCR扩增产物的 混合物与固相载体上固定的测序探针进行杂交, 并进行固相桥式 PCR扩增, 形成测 序簇; 对所述测序簇用"边合成 -边测序法"进行测序, 从而得到样本中核酸分子的 序列。
DNA簇的形成是使用表面连有一层单链引物 (primer)的测序芯片 (flow cell), 单链状态的 DNA片段通过接头序列与芯片表面的引物通过碱基互补配对的原 理被固定在芯片的表面, 通过扩增反应, 固定的单链 DNA变为双链 DNA, 双链 再次变性成为单链, 其一端锚定在测序芯片上, 另一端随机和附近的另一个引 物互补从而被锚定, 形成"桥"; 在测序芯片上同时有上千万个 DNA单分子发生 以上的反应; 形成的单链桥, 以周围的引物为扩增引物, 在扩增芯片的表面再 次扩增, 形成双链, 双链经变性成单链, 再次成为桥, 称为下一轮扩增的模板 继续扩增; 反复进行了 30轮扩增后, 每个单分子得到 1000倍扩增, 称为单克隆 的 DNA簇。
DNA簇在 Solexa测序仪上进行边合成边测序, 测序反应中, 四种碱基分别 标记不同的荧光,每个碱基末端被保护碱基封闭,单次反应只能加入一个碱基, 经过扫描, 读取该次反应的颜色后, 该保护集团被除去, 下一个反应可以继续 进行, 如此反复, 即得到碱基的精确序列。 在 Solexa多重测序 (Multiplexed Sequencing)过程中会使用 Index(标签)来区分样品, 并在常规测序完成后, 针对 Index部分额外进行测序, 通过 Index的识别, 可以在 1条测序甬道中区分多达 12 种不同的样品。
MHC型别数据库及其构建
本发明提供一种 MHC型别数据库。 所述数据库包括的内容用表 1 的形式 表示。
表 1
Figure imgf000009_0001
在表 1中, 第一列表示 MHC的不同型别, 第二列表示 MHC型别所对应的 序列, 第三列表示 MHC型别序列比对参考序列后的 SNP, 第四列表示 MHC型 别比对后的 InDel。
例如, 29796327:T 表示 29796327 位置 SNP 为 T; 29796435-D-C 表示 29796436位置缺失碱基 C; None表示没有这种类型的变异。
本发明还提供了 MHC 型别数据库构建方法, 在本发明的一个优选例中, 所述方法包括步骤 a和 b:
步骤 a)下载已知的 MHC型别序列, 本领域的普通技术人员可以使用常规 得 到 获得 这些数据 , 例 如 , 从 IMGT 数据 库 获得 , 网 址 为 htt i //www .ebi.ac. uk/ imgt/¾la 。
MHC型别序列的内容可以用表 2的形式表示。
表 2
Figure imgf000009_0002
在表 2中, 第一列表示型别, 第二列表示型别的序列。
选取对照样本 hgl9或 hgl8中目标基因坐标对应的序列作为参考序列, 例 如 A基因, 选取 A*03:01 :01 :01对应的序列作为参考序列。 步骤 b )将 MHC型别序列与参考序列进行比对, 获得与参考序列的差异位 点, 即得到每个型别相对于参考序列的 SNP和 InDel关系, 构建 MHC型别数 据库。
SNP&InDel检测方法
本发明提供了一种基于重新比对策略得到每条读序的 SNP和 InDel信息的 方法。 在一个优选例中, 所述方法包括步骤:
a)将 BWA软件包比对得到的不匹配的 SAM文件和目标基因区的 SAM文 件合并, 作为原始比对文件;
b) 将合并后的比对文件 (也就是 SAM文件) 转为 FQ文件;
c) 将合并的 FQ文件重新用 BWA软件包比对, 参考序列选为目标基因对 应的 MHC型别序列, 如表 3所示:
表 3
Figure imgf000010_0001
将型别作为表头, 而型别对应的序列作为参考序列;
d )重新比对后,得到每条读序比对的最佳型别, 以及对应型别的初始位置, 以及与最佳型别差异的 mismatch数目。
通过设置 mismatch数目过滤不可信的读序, 然后结合 MHC型别数据库获 得可信读序的 SNP和 InDel信息。 基于连锁的 SNP和 InDel拼接
本发明提供了一种基于连锁的 SNP和 InDel 的拼接方法,在一个优选例中, 所述方法包括步骤:
a)选取目标基因的一条读序对应的 SNP和 InDel作为起始序列, 用下面形 式表示:
序列 1 29910242*29910331 *29910276:G-29910286:T*None b) 从目标基因中挑选与上述序列完全匹配的读序, 读序 1 29910282*29910371 *29910286:T-29910358 :G*None
读序 2 29910287*29910386*29910286:T-29910358:G-29910378:T*None 读序 3 29910282*29910371 *29910286:T-29910371 :T*None
c ) 将挑选的读序与原始的序列进行拼接成更长的序列, 如下:
Figure imgf000011_0001
d)将得到的序列重新从目标基因中提取与序列完全匹配的读序直到没有匹 配的读序;
e ) 将没有 SNP和 InDel的读序按照同样的方法单独拼接;
f) 将拼接完成的序列比对到型别数据库进行过滤, 将保留下来的序列, 但 是又没有跨过整个外显子的序列进行随机组合后重新过滤数据库, 将过滤得到 的 hap排序, 排序规则: 首先选取读序支持数最多的一条 hap, 然后将剩余的 hap与这条 hap取并集并去重复后, 按照并集的读序支持数排序。
MHC分型方法
本发明提供一种 MHC分型方法, 在一个优选例中, 步骤如下:
a) 将获得的 hap以及 hap对应的 MHC型别转化为如表 4所示的格式:
表 4
Figure imgf000011_0002
以 G*01 :01 :01 :01G*01 :01 :01 :02 G*01 :01 :01 :03为例, 头两位数字表示的是 在物种水平上的分型, 第 3、 4位上的数字表示的是外显子非同义突变, 第 5、 6位表示的同义突变, 第 7、 8位上表示的是内显子上的突变的分型, 但是由于 在内显子上的突变的研究意义不大, 所以一般只做前三部分的分型研究。 其中 Gl表示 G基因 1号外显子, -1表示对应 hap的排序数; b)确定最可信的 type型别, 定义为 type型别 1, 具体做法: 对所有的 type 型别进行打分, 选出最可信 type型别作为 type型别 1, 打分规则如下:
如果 hap对应的排序数为 1, 则分数不加; 如果 hap对应的排序数为 2, 则 分数加 0.5; 如果为其他数据, 则分数加 1; 如果部分外显子没有排序数, 则分 数加 2; 如此分数最低的就定义为单体型 1 。
例如:
G1-1*G2-1*G3-2*G4-1*G5-2=1
G1-1*G2-2*G3-2*G4-1*G5-2=1.5
G1-2*G2-1*G3-1*G4-1*G5-1=0.5
G1-2*G3-1*G4-1*G5-1=2.5
选择分数最低为 0.5 的 G1-2*G2-1*G3-1*G4-1*G5-1 作为最可信 type, 从 表 4可以知道, 其对应的型别为 G*01:01:01:01G*01:01:01:02 G*01:01:01:03。
c) 从剩余的 type中选出 type2, 规则如下:
将剩下的 type与步骤 b)中获得的 typel取差异度,差异度表示同一个基因 同一个 exon他们对应的 hap序数不一样, 这就是不同的 hap取自原始读序的不 同部分, 所以 hap差异数越大, 表示这组型别组合可以得到原始读序的最大部 分读序, 也就越可信。
计算规则, 如果外显子对应的 hap序数不一样, 则加 2; 如果对应一样, 则 加 1; 如果 typel存在而 type2不存在, 则减 1; 最后得分最大的就是最佳组合。
以表 4的数据为例, 差异度结果如下:
Gl-l*G2-l*G3-2*G4-l*G5-2 & G1-2*G2-1*G3-1*G4-1*G5-1 = 8;
G1-2*G2-1*G3-1*G4-1*G5-1 & Gl-l*G2-2*G3-2*G4-l*G5-2 =9;
G1-2*G2-1*G3-1*G4-1*G5-1 & G1-2*G3-1*G4-1*G5-1 =3。
因此, 最后选择 type组合为 G*01:01:01:01G*01:01:01:02 G*01:01:01:03与
G*01:04:01。 本发明的主要优点包括:
(1)本发明方法和系统准确性高, 对待测的数据要求比较低; (2)本发明方法和系统对于信息的分析过程快速, 操作简单;
(3)相对于现有的分型方法, 大大提高了分型区域。 下面结合具体实施例, 进一步阐述本发明。 应理解, 这些实施例仅用于说 明本发明而不用于限制本发明的范围。 下列实施例中未注明具体条件的实验方 法,通常按照常规条件如 Sambrook等人,分子克隆:实验室手册(New York : Cold Spring Harbor Laboratory Press, 1989)中所述的条件, 或按照制造厂商所 建议的条件。 实施例 1
4个样品总计 102个基因的分型
本实施例的目的: 通过高通量测序技术对重测序数据或给予目标区域捕获 数据进行 MHC分型, 确定 SNP连锁性。
1: 构建 MHC型别数据库
构建方法:
从网址 http://www.ebi.ac.uk/imgt/hla下载最新的 MHC型别对应的序列。 选取与对照样本 hgl 9/hgl 8对应坐标的序列作为参考序列,对 MHC型别序 列进行比对, 得到相对于参考序列的差异位点即 InDel和 SNP。
2: 生成比对文件
利用 illuminla 测序平台测序后得到下机数据, 经过过滤低质量读序数和 adapter污染的读序数后得到可利用的读序数。 使用 BWA比对软件和 samtools 软件为例来说明。
通过 BWA比对软件, 将这些序列与 hgl9或 hgl 8作为参考序列进行序列 比对, 经过 aln和 sampe两步, 得到比对结果 *.sam文件后, 利用 samtools工具 包对 *sam文件处理, 包括排序, 去重复, 建立索引等处理得到 *bam文件。
3: 挑出目标区域 SNP&InDel信息
从 bam文件中得到目标基因的比对信息, 同时, 与不匹配的读序合并, 作 为目标基因原始的比对信息, 将比对信息重新转为 FQ文件。 将 FQ文件重新进行比对, 比对参考序列改为目标基因对应的 MHC型别, 因为 MHC 多态性很高, 部分读序在比对时被丢失掉, 所以回收不匹配的读序 重新进行比对, 解决比对多态性问题。
从比对后的最佳型别对应的 SNP和 InDel中得到每条读序的比对信息。
4: 根据读序的 SNP和 InDel信息进行连锁
利用读序之间的 overlap进行连锁, 连锁原则, 将读序上没有 SNP和 InDel 的 reads单独连锁。
5: 根据 MHC型别数据库对连锁的 Hap进行过滤以及连接
利用 MHC 型别数据库对拼接完的序列进行整体过滤, 如果不能匹配数据 库, 则所有的读序都被丢失。 而剩下的 hap被定义为可信的 hap。
6: 对过滤得到的 hap进行排序
首先取出读序支持数最多的一条 hap, 定义为 hapl, 然后将剩余的 hap与 hapl取并集后排序。
7: 结合所有外显子得到最后型别
将之前得到的 hap以及 hap对应的 MHC型别转化为如下表 5所示的格式: 表 5
Figure imgf000014_0001
其中 Gl表示 G基因 1号外显子, -1表示对应 hap的排序数。
首先确定最可信的 type, 定义为 typel ,利用打分对所有的 type进行打分选 出最可信 type作为 typel , 结果如下:
G1-1 *G2-1 *G3-2*G4-1 *G5-2=1
G1-1 *G2-2*G3-2*G4-1 *G5-2=1.5
G1-2*G2-1 *G3-1 *G4-1 *G5-1 =0
G1-2*G3-1 *G4-1 *G5-1=2 选择 G1-2*G2-1*G3-1*G4-1*G5-1 对应的型别 G*01:01:01:01G*01:01:01:02 G*01 :01 :01 :03作为最可信 type (型别) 。
从剩余的 type中选出 type2, 规则是将剩下的 type与 typel取差异度, 差异 度表示同一个基因同一个 exon他们对应的 hap序数不一样, 这就是不同的 hap 取自原始读序的不同部分, 所以 hap差异数越大的表示这组型别组合可以得到 原始读序的最大部分读序, 也就是最可信了, 差异度结果如下:
Gl-l*G2-l*G3-2*G4-l*G5-2 & G1-2*G2-1*G3-1*G4-1*G5-1 =8;
G1-2*G2-1*G3-1*G4-1*G5-1 & Gl-l*G2-2*G3-2*G4-l*G5-2 =9;
G1-2*G2-1*G3-1*G4-1*G5-1 & G1-2*G3-1*G4-1*G5-1 =3;
所以最后选择 type 组合为 G*01:01:01:01G*01:01:01:02 G*01 :01 :01 :03 与 G*01:04:01。
8. 综合上述步骤, 4个样品总计 102个基因的分型结果见表 6。
表 6
Figure imgf000015_0001
02:01 02:01 03:03 03:03:01 01:02 01:02:01
DQA1
03:03 03:03:01 05:05 05:05:01 05:01 05:01:01
01:01 01:05N 01:01 01:01:03 01:01 01:01:01
G
01:05N 01:05N 01:01 01:01:01 01:04 01:04:01 结果表明, 本方法的正确率达到 98%以上。
在本发明提及的所有文献都在本申请中引用作为参考, 就如同每一篇文献 被单独引用作为参考那样。 此外应理解, 在阅读了本发明的上述讲授内容之后, 本领域技术人员可以对本发明作各种改动或修改, 这些等价形式同样落于本申 请所附权利要求书所限定的范围。

Claims

权 利 要 求
1. 一种主要组织相容性复合体 MHC型别数据库的构建方法, 其特征在于, 包括步骤:
(1) 将目标 MHC型别序列和参考序列进行比对, 获得目标 MHC型别相对于参 考序列的差异位点, 所述的差异位点为 SNP和 /或 InDel位点; 和
(2) 对已获得所述的差异位点信息的各个目标 MHC型别进行汇总, 构建得到 主要组织相容性复合体 MHC型别数据库。
2. 一种主要组织相容性复合体 MHC型别数据库,其特征在于,所述的数据 库是使用权利要求 1所述的方法构建的。
3. 一种构建主要组织相容性复合体 MHC型别数据库的单元, 其特征在于, 所述单元包括模块:
(1) 比对模块, 用于比对目标 MHC型别序列和参考序列; 和
(2) 输出模块,用比对模块获得的目标 MHC型别序列和参考序列的比对结果, 输出 MHC型别相对于参考序列的 SNP和 InDel位点信息。
4. 一种 SNP和 InDel检测方法, 其特征在于, 包括步骤:
(1) 获得经比对的不匹配的(unmap) SAM文件和目标基因区的 SAM文件, 并 且合并所述的 SAM文件;
(2) 重新比对步骤 (1)所获得的 SAM文件, 获得每条读序比对的最佳型别、 对 应型别的初始位置、 和与最佳型别差异的不匹配的数目; 和
(3) 过滤步骤 (2)所获得的读序, 结合 MHC数据库获得可信读序的 SNP和 InDel 的信息。
5. 一种 SNP和 InDel检测单元, 其特征在于, 所述单元包括模块:
(1) 比对模块, 用于比对读序;
(2) 文件合并模块, 用于合并经比对模块获得的比对的不匹配 SAM文件和目 标基因的 SAM文件;
(3) 文件转换模块, 用于转换文件合并模块获得的 SAM文件和 FQ文件; 和
(4) 输出模块, 用于输出 SNP和 /或 InDel的信息。
6. 一种 SNP和 InDel的拼接方法, 其特征在于, 包括步骤: (1) 将目标基因的一条读序所对应的 SNP和 InDel作为起始序列;
(2) 从目标基因中选择与步骤 (1)获得的序列完全匹配的读序;
(3) 将步骤 (2)选择的读序与步骤 (1)的序列拼接, 获得更长的序列;
(4) 从目标基因中提取与步骤 (3)获得的序列完全匹配的读序,直到没有匹配的 读序, 从而获得拼接的序列;
(5) 将步骤 (4)获得的拼接的序列与 MHC型别数据库进行比对和过滤, 对没有 跨过整个外显子的序列进行随机组合和过滤, 从而获得过滤的序列; 和
(6) 对步骤 (5)获得的过滤的序列按照读序支持数进行排序。
7. 一种 SNP和 InDel的拼接单元, 其特征在于, 所述单元包括模块: (1) 序列获取模块, 用于选择和 /或获得读序;
(2) 拼接模块, 用于对序列获取模块获得的读序进行读序与序列的拼接;
(3) 比对和过滤模块, 用于比对和过滤拼接模块获得的读序;
(4) 排序模块, 用于对比对和过滤模块获得的序列进行排序; 和
(5) 输出模块, 用于输出 SNP和 InDel的拼接信息。
8. 一种主要组织相容性复合体 MHC分型方法, 其特征在于, 包括步骤:
(1) 获得目的基因的序列及其对应的 MHC型别, 所述序列如式 I所示:
(Gi-j) n
式 I
其中, G为基因类型, i为外显子编号, j为对应的 hap排序数, n为外显子的数 目, i、 j和 n均为正整数;
(2) 对各个外显子序列对应的排序数进行打分, 确定最可信单体型 (type) , 打分规则如下:
当 j=l时, 分数 =0;
当 j=2时, 分数 =0.5;
当」=其他时, 分数 =1 ;
当没有外显子排序数时, 分数 =2;
计算各个外显子对应的分数的总和, 分数最低的为最可信单体型;
(3) 分别比较其余单体型与最可信单体型的差异度, 差异度计算规则如下: 比较其余单体型与最可信单体型对应外显子的序列排序数 (gpj值) , 当 j不同时, 分数 =2;
当 j相同时, 分数 =1 ;
当其余单体型中不存在相应的 j, 分数 = -2;
最终总分最大的单体型与最可信单体型的组合, 为最佳组合; 和
(4) 基于步骤 (3)的最佳组合, 得出 MHC的分型信息。
9. 一种主要组织相容性复合体 MHC分型单元, 其特征在于, 所述的单元包 括模块:
(1) 序列获取模块, 用于选择和 /或获得读序;
(2) 排序模块, 从序列获取模块获得的读序中, 确定最可信单体型和最佳单体 型的组合; 和
(3) 输出模块, 用于输出 MHC的分型信息。
10. 一种主要组织相容性复合体 MHC分型系统, 其特征在于, 所述系统包 括单元:
(1) 权利要求 9所述的主要组织相容性复合体 MHC分型单元;
(2) 权利要求 3所述的构建主要组织相容性复合体 MHC型别数据库的单元;
(3) 权利要求 5所述的 SNP和 InDel 检测单元; 和
(4) 权利要求 7所述的 SNP和 InDel的拼接单元。
PCT/CN2012/084689 2012-11-15 2012-11-15 一种主要组织相容性复合体mhc分型方法及其应用 Ceased WO2014075274A2 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/CN2012/084689 WO2014075274A2 (zh) 2012-11-15 2012-11-15 一种主要组织相容性复合体mhc分型方法及其应用
CN201280076912.6A CN104769129B (zh) 2012-11-15 2012-11-15 一种主要组织相容性复合体mhc分型方法及其应用

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2012/084689 WO2014075274A2 (zh) 2012-11-15 2012-11-15 一种主要组织相容性复合体mhc分型方法及其应用

Publications (1)

Publication Number Publication Date
WO2014075274A2 true WO2014075274A2 (zh) 2014-05-22

Family

ID=50731778

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2012/084689 Ceased WO2014075274A2 (zh) 2012-11-15 2012-11-15 一种主要组织相容性复合体mhc分型方法及其应用

Country Status (2)

Country Link
CN (1) CN104769129B (zh)
WO (1) WO2014075274A2 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105512514A (zh) * 2014-09-23 2016-04-20 深圳华大基因股份有限公司 一种mhc补全数据库、其构建方法和应用

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2021534394A (ja) * 2018-08-14 2021-12-09 ボード オブ リージェンツ, ザ ユニバーシティ オブ テキサス システムBoard Of Regents, The University Of Texas System 主要組織適合遺伝子複合体に結合されたペプチドを配列決定する単一分子

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102127819B (zh) * 2010-11-22 2014-08-27 深圳华大基因科技有限公司 Mhc区域核酸文库的构建方法及用途
WO2012068701A2 (zh) * 2010-11-23 2012-05-31 深圳华大基因科技有限公司 Hla基因型别一snp连锁数据库、其构建方法、以及hla分型方法

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105512514A (zh) * 2014-09-23 2016-04-20 深圳华大基因股份有限公司 一种mhc补全数据库、其构建方法和应用

Also Published As

Publication number Publication date
CN104769129A (zh) 2015-07-08
CN104769129B (zh) 2017-07-07

Similar Documents

Publication Publication Date Title
US9562269B2 (en) Haplotying of HLA loci with ultra-deep shotgun sequencing
CN103221551B (zh) Hla基因型别-snp连锁数据库、其构建方法、以及hla分型方法
CN105779280B (zh) 由母本生物样品进行胎儿基因组的分析
CN106103736B (zh) 高分辨率等位基因鉴定
Debladis et al. Detection of active transposable elements in Arabidopsis thaliana using Oxford Nanopore Sequencing technology
WO2018213498A1 (en) Identification of somatic or germline origin for cell-free dna
CN103492588A (zh) 用于单体型测定的方法和系统
WO2015200701A2 (en) Software haplotying of hla loci
Claes et al. Dealing with pseudogenes in molecular diagnostics in the next-generation sequencing era
WO2014023076A1 (zh) 一种地中海贫血的分型方法及其应用
CN116323979A (zh) 用于hla分型的方法、组合物和试剂盒
Claes et al. Dealing with pseudogenes in molecular diagnostics in the next generation sequencing era
Sun et al. Technical strategy for monozygotic twin discrimination by single-nucleotide variants
Yang et al. The next generation of complex lung genetic studies
US20200232033A1 (en) Platform independent haplotype identification and use in ultrasensitive dna detection
JP2016516449A (ja) Hlaマーカーを使用する母体血液中の胎児dna分率の決定方法
WO2013129542A1 (ja) Hla-a*31:01アレルの検出方法
WO2013078684A1 (zh) 一种组装双亲基因组的方法
CN104769129B (zh) 一种主要组织相容性复合体mhc分型方法及其应用
CN103360490B (zh) 一种peb致病基因新突变及其应用
CN104024410A (zh) Hla-a*24组的判定方法
US20240294982A1 (en) Method, kit and cartridge for detecting nucleic acid molecule
EP3596229A1 (en) Method and system for nucleic acid sequencing
Kunkel et al. Molecular methods for human leukocyte antigen typing: current practices and future directions
US20220392568A1 (en) Method for identifying transplant donors for a transplant recipient

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 12888298

Country of ref document: EP

Kind code of ref document: A2

NENP Non-entry into the national phase in:

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 25/09/2015)

122 Ep: pct application non-entry in european phase

Ref document number: 12888298

Country of ref document: EP

Kind code of ref document: A2