WO2013016864A1 - 目标区域捕获方法及其生物信息处理方法和系统 - Google Patents

目标区域捕获方法及其生物信息处理方法和系统 Download PDF

Info

Publication number
WO2013016864A1
WO2013016864A1 PCT/CN2011/077861 CN2011077861W WO2013016864A1 WO 2013016864 A1 WO2013016864 A1 WO 2013016864A1 CN 2011077861 W CN2011077861 W CN 2011077861W WO 2013016864 A1 WO2013016864 A1 WO 2013016864A1
Authority
WO
WIPO (PCT)
Prior art keywords
capture
species
sequence
region
source
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2011/077861
Other languages
English (en)
French (fr)
Inventor
朱昱其
金鑫
李莹
李英睿
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BGI Shenzhen Co Ltd
Original Assignee
BGI Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BGI Shenzhen Co Ltd filed Critical BGI Shenzhen Co Ltd
Priority to PCT/CN2011/077861 priority Critical patent/WO2013016864A1/zh
Priority to CN201180071091.2A priority patent/CN103547681B/zh
Publication of WO2013016864A1 publication Critical patent/WO2013016864A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6813Hybridisation assays
    • C12Q1/6834Enzymatic or biochemical coupling of nucleic acids to a solid phase
    • C12Q1/6837Enzymatic or biochemical coupling of nucleic acids to a solid phase using probe arrays or probe chips
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q2565/00Nucleic acid analysis characterised by mode or means of detection
    • C12Q2565/50Detection characterised by immobilisation to a surface
    • C12Q2565/501Detection characterised by immobilisation to a surface being an array of oligonucleotides

Definitions

  • Target area capture party and its information processing system
  • the present invention relates to the field of bioinformatics technology, and in particular, to a target area capturing method and a biological information processing method and system thereof. Background technique
  • next-generation sequencing technology has driven the sequencing of model species.
  • the cost of sequencing of whole genomes has dropped, and the genome sequences of more and more species have been solved.
  • the large amount of genetic information that comes with it is in urgent need of large-scale, high-throughput analysis methods and means for further research.
  • Gene chip technology from the early 1990s is an important technology for studying the structure and function of genes.
  • high-throughput sequencing technology specific target regions of a certain species, especially exons, can be captured and sequenced. Find functional genes and perform subsequent bioinformatics analysis.
  • This method has high capture efficiency, low cost compared to whole genome sequencing, and very good accuracy.
  • foreign researchers have used target region capture technology to use exon capture chips to perform exon sequencing on humans, and then carry out disease-related research, which has achieved a series of ⁇ 1 ⁇ 2.
  • One technical problem to be solved in one aspect of ⁇ is to provide a target area.
  • the domain capture method can increase the efficiency of capture of the target area of the species.
  • An aspect of the invention provides a method for capturing a target region, comprising: utilizing a target region capture chip to capture a target region capture chip to specify a gene fragment of a near-source species of the species used;
  • a gene fragment of a near-source species is sequenced to obtain a gene fragment.
  • the method further comprises: acquiring a capture region sequence on the target region capture chip designating the use species reference genome; and acquiring a near source region on the near source reference genome based on the capture region sequence.
  • the method further comprises: filtering the gene slice to remove the unqualified gene slice.
  • the target region capture method provided by the method improves the capture of the target region of the species without the target region capture chip by applying the target region capture chip to the capture of the target region of the near-source species of the designated species of the chip.
  • Another technical problem to be solved is to provide a biological information processing method and system, which can improve processing efficiency.
  • Another aspect of the invention provides a biological information processing method, comprising: utilizing a target region capture chip to capture a target region capture chip to specify a gene fragment of a near-source species of a species used;
  • Sequencing a gene fragment of a near-source species to obtain a gene fragment obtaining a target region capture chip designating a capture region on a species reference genome;
  • the near-source region on the reference genome of the near-source species is obtained according to the sequence of the capture region; the bio-information analysis is performed on the gene fragment sequence according to the reference genome of the near-source species and the near-source region of the near-source species.
  • the biological information analysis comprises: assembly, single nucleotide polymorphism site finding, copy number variation detection, interpolation X/deletion detection, detection of chromosome structural variation, association of SNP loci with disease, or analysis of SNP loci and drug effects.
  • a further aspect of the invention provides a biological information processing system, comprising: a capture region sequence acquisition device, configured to acquire a capture region sequence on a target reference capture genome of a target region capture chip;
  • a near-source region acquisition device configured to acquire a near-source region on a reference genome of a near-source species of a specified use species according to a sequence of capture regions;
  • a gene fragment capture device for capturing a gene fragment of a near-source species using a target region capture chip
  • a sequencing device for sequencing a gene fragment of a near-source species to obtain a gene fragment.
  • the system further comprises filtering means for filtering the sequence of the gene fragments obtained by the sequencing device to remove the unqualified gene fragments.
  • the system further comprises an information analysis device for performing biometric analysis on the gene slice according to the reference genome of the near source species and the near source region of the near source species.
  • FIG. 1 is a flow chart showing one embodiment of a target area capturing method according to the present invention.
  • FIG. 2 shows a flow chart of another embodiment of a target area capture method in accordance with the present invention
  • Figure 3 shows a flow chart of one embodiment of a biological information processing method in accordance with the present invention
  • Figure 4 is a graph showing the distribution of homologous regions of the human genome chip according to the present invention to the monkey genome group;
  • FIG. 5 is a block diagram showing an embodiment of a biological information processing system in accordance with the present invention.
  • Fig. 6 is a block diagram showing another embodiment of a biological information processing system according to the present invention. detailed description
  • Figure 1 is a flow chart showing one embodiment of a method of capturing a genomic target region according to the present invention.
  • the target region capture chip capture target region capture chip specifies the gene fragment of the near-source species of the species used.
  • the genetic distance between near-source species is relatively close, and the similarity of genetic material is high (especially in highly conserved regions such as exons). Sequence similarity will always appear in gene exons with similar functions between near-source species. Very high or even identical conserved regions, making it possible for existing target region capture chips to capture gene fragments from other near-source species.
  • the target region capture chip of the designated species the target region capture is performed on the DNA sample of the near-source species (target species) to be studied according to the chip operation manual, and the gene fragments of the near-source species are captured.
  • the gene fragment of the near-source species is sequenced to obtain a gene fragment sequence.
  • high-throughput sequencing of gene fragments captured from near-source species to obtain gene fragment sequences high-throughput sequencing technology using Illumina GA sequencing technology, or other high-throughput sequencing technologies.
  • the target area capture chip by applying the target area capture chip to the target area capture of the near-source species, the target area capture of the near-source species is realized, and the chips with fewer species are better utilized, which helps to obtain more. Analysis of individual functional genes and their subsequent biological information.
  • Fig. 2 is a flow chart showing another embodiment of a method of capturing a genomic target region according to the present invention.
  • the acquisition target region capture chip specifies the capture region sequence on the species reference genome. For example, according to the capture area file of the target area capture chip provided by the chip manufacturer, the corresponding capture area sequence is intercepted on the reference genome of the designated use species of the chip.
  • a near-source region on the near-source species reference genome is obtained from the capture region sequence.
  • the capture region file of the target region capture chip for the near source species may be generated according to the near source region of the near source species.
  • the resulting chip is designated to use a sequence of species capture regions, and the alignment software (for example, software Blast suitable for local alignment) is used for homology alignment with the reference genome of the near-source species to be studied, according to the alignment result.
  • a near-source region corresponding to the capture region sequence on the reference genome of the near-source species is obtained.
  • the homology alignment process allows mismatches to occur, and the mismatch rate that can be allowed to occur is, for example, 1% to 20%, or 1% to 10%, or about 5%.
  • the mismatch rate can be reduced or increased according to different needs.
  • the range of mismatch ratios is determined according to the genetic distance between the species specified by the chip and the near-source species of the target, and the large mismatch rate is accepted between species with large genetic distances, and between species with small genetic distances. Smaller mismatch rate.
  • the contents of the comparison result include, for example: which fragment of the reference region sequence is homologous to the reference genome of the near source species, the size of the homology, and the specific coordinates of the fragment on the reference genome of the near-source species; according to the coordinate information, A near-source region of the near-source species reference genome that may correspond to the capture region sequence is obtained. Further, the capture region file of the chip for the near source species can be generated from the near source region of the near source species. At step 206, the target region capture chip captures the target region capture chip to specify the gene fragment of the near source species of the species used.
  • the gene fragment of the near-source species is sequenced to obtain a gene fragment sequence.
  • steps 202 and 204 may also be performed after step 208 or in parallel with steps 206 and 208.
  • the genes of different species have very conserved regions, so that the gene chip originally designed for a single species is used for other near
  • the capture of the target species of the source species provides the possibility.
  • the target area capture chip is only suitable for the concept of using species, which limits the feasibility of the general technician to carefully analyze the gene chip designed for a single species for the capture of other near-source target areas, as well as the specific implementation of the technical solution. . So far, no research results published using this idea have been found.
  • the above possibilities do not enable those skilled in the art to naturally use gene chips designed for a single species for the capture of target regions of other near-source species, and to overcome some technical problems for bioinformatics analysis.
  • the capture region file provided by the target region capture chip helps to screen the genome-wide data of the near-source species, thereby obtaining the near-source region of the near-source species, thereby obtaining the capture region sequence of the target region capture chip for the near-source species. , greatly reducing the workload of subsequent comparison work.
  • the capture region of the near-source species is found by using the homology comparison, so that the position of the captured sequence fragment on the reference genome of the near-source species can be finely located, which facilitates subsequent analysis and overcomes the capture of the chip through the target region. Obstacles to subsequent analysis are not available after sequencing sequences of near-source species are obtained.
  • the gene fragment is also filtered to remove the unqualified gene fragment.
  • unqualified gene fragments include: Sequencing quality The number of bases below the low quality threshold exceeds the predetermined percentage of the entire sequence of bases
  • a sequence of 40% or more, 50% or more, the low quality threshold is determined by the specific sequencing technology and the sequencing environment; the number of bases with undefined sequencing results in the sequence (such as N in the IUumina GA sequencing result) exceeds the whole number.
  • a sequence of a predetermined number of bases for example, 10%, 20%
  • a sequence of a gene fragment having a foreign sequence which is aligned with an exogenous sequence introduced by other experiments, such as various linker sequences, if present in the sequence
  • the source sequence is considered to be an unqualified sequence.
  • the linker sequence in the remaining qualified short segment sequence is removed.
  • the bio-information analysis of the gene fragments can be performed based on the reference genome of the near-source species and the near-source region of the near-source species.
  • Fig. 3 is a flow chart showing an embodiment of a biological information processing method according to the present invention.
  • step 302 the chip is captured by the target region of the designated species, and the target region is captured by the DNA sample of the near-source species according to the chip operation manual, and the gene fragment of the near-source species is obtained.
  • the captured gene fragment is subjected to high throughput sequencing to obtain a gene fragment sequence.
  • high-throughput sequencing technology can be IUumina GA sequencing technology, or other existing high-throughput sequencing technologies.
  • sequencing data (gene slice) is received for sequencing data pretreatment. For example, filtering the gene fragments.
  • a corresponding capture region is intercepted on the reference genome of the designated species of the chip according to the capture region file of the target region capture chip provided by the chip manufacturer.
  • step 310 the chip obtained in the above step is designated to use the sequence of the species capture region, and the homologous alignment is performed with the reference genome of the near-source species to be studied by using a software suitable for local alignment, such as Blast, to obtain an alignment. result.
  • a capture zone file of the chip for the near source species is generated.
  • Bioinformatics analysis includes, but is not limited to, assembly, single nucleotide polymorphism (SNP) locus finding, copy number variation (CNV) detection, insertion/deletion (Indel) detection, chromosome structural variation (SV) detection, SNP locus Correlation analysis with disease, correlation analysis between SNP locus and drug effects.
  • SNP single nucleotide polymorphism
  • CNV copy number variation
  • Indel insertion/deletion
  • SV chromosome structural variation
  • Proximal species Target species
  • Sample 2 DNA samples of crab-eating macaque (scientific name: Macaca fascicularis).
  • Target area capture chip Agilent's (38M) human exon chip (trade name of the chip: Agilent 2100 Bioanalyzer, article number / model: G2938A)
  • the design object of the designated species of the chip Human genome (Hgl8, ie NCBI BUILD36), can be accessed at http://www.ncbi.nlm.nih.gov/sites/genome/, download the hgl8 data at the following ftp download address: ftp ://ftp.ncbi.nlm.nih.gOv/genomes/H sapiens/ARCHIVE/
  • the unqualified sequence includes: The number of bases with a sequencing quality value below 20 is more than 50% of the entire sequence, which is considered to be a non-conforming sequence; the base with undefined sequencing results in the sequence (N in IUumina GA sequencing results) A number exceeding 10% of the entire sequence number is considered to be a non-conforming sequence; in addition to the sample linker sequence, it is aligned with other experimentally introduced exogenous sequences, such as various linker sequences. A foreign sequence is considered to be a non-conforming sequence if it exists in the sequence.
  • the corresponding capture area sequence is intercepted on the human reference genome (Hgl8) and written into a fastq file exoncapturcfa in the format:
  • the line where > is located represents which chromosome of the capture region sequence is on Hgl8, and its start and stop coordinates.
  • the lower row is the specific genotype of this capture region sequence.
  • the first and second rows indicate the number of aligned and unaligned upper loci in the homology alignment
  • the tenth column indicates the capture region sequence name
  • the fourteenth to seventeenth columns indicate the homologous to the capture region.
  • the chip design is in the region of 20138 bp to 20258 bp of chromosome 1 of the human reference genome, and is homologous to the region of 113205694 bp to 113205813 bp on chromosome 13 of the cynomolgus reference genome.
  • the total length of chromosome 13 of the cynomolgus reference genome is 137686314 bp.
  • the hgl8 genomic fragment was truncated to obtain a theoretical chip capture fragment, which was homologously analyzed with the cynomolgus reference genome.
  • 97.7% of the fragments were found to find homologous regions on the cynomolgus reference genome, and this homology was expressed as homology of the entire exon fragment.
  • the remaining 2.3% cannot find homologous regions, and it is estimated that it may be a unique fragment produced by humans in evolution.
  • Figure 4 shows the distribution of the homology of the human genome chip capture region to the monkey genome, which visually shows where the chip can capture the gene fragment of the cynomolgus monkey on the monkey chromosome (where the captured place is indicated in gray) ). It is obvious that the gray area is distributed on all the chromosomes of the monkey, and it is relatively uniform, indicating that the human genome exon chip can capture a large number of exon fragments on each chromosome of the monkey.
  • the capture region of the Agilent 38M chip on the reference genome of the cynomolgus monkey was obtained, and the capture region file of the Agilent 38M chip for the reference genome of the cynomolgus monkey was generated and separated according to the chromosome.
  • the file capture_region_chr*.txt the specific format is:
  • the first column indicates the chromosome number
  • the second column indicates the starting coordinates
  • the third column indicates the ending coordinates.
  • the quality of the sequencing data is as follows:
  • the coverage in the above table refers to the percentage of the actual captured segment that covers the theoretical chip capture area. That is, the fragment captured on the cynomolgus genome by the human genome chip can cover 85% of the chip design capture area. Because all the parameters are used when running ECP, the conditions are very high. Only when the comparison is very good and the number of comparisons is enough, it is covered. 85% coverage, the current situation is In the face condition, 85% of the chip design area is a well-recovered region of the relevant gene for cynomolgus monkeys.
  • the exon fragment of a large number of cynomolgus monkeys can be efficiently obtained by using an exome of the human genome, and the sequencing quality is also relatively stable.
  • bioinformatics analysis was performed based on the obtained results against the reference genome of cynomolgus monkeys.
  • the cross-species sequence homology alignment is obtained, and the capture region of the human target region capture chip for other species is obtained, and a method for capturing the cross-species target region is realized, and the target of the human near-source species is achieved.
  • Area capture and high capture quality extend the research methodology of low-cost, high-accuracy target area capture to cross-species levels.
  • Fig. 5 is a block diagram showing an embodiment of a biological information processing system according to the present invention.
  • the bio-information processing system includes: a capture region sequence obtaining device 51, configured to acquire a capture region sequence on a target region capture chip designated use species reference genome; a near-source region acquisition device 52, which is obtained according to the capture region sequence
  • the capture region sequence acquired by the device 51 acquires a near-source region on the reference genome of the near-source species of the designated use species;
  • the gene segment capture device 53 captures the gene fragment of the near-source species using the target region capture chip; and the sequencing device 54 for the near-source species
  • the gene fragment was sequenced to obtain a gene fragment.
  • Figure 6 shows a block diagram of another embodiment of a biological information processing system in accordance with the present invention.
  • the biological information processing system in this embodiment includes an information analysis device 66 in addition to the capture region sequence acquisition means 51, the near source region acquisition means 52, the gene fragment capture means 53, and the sequencing means 54.
  • the information analysis device 66 performs biometric analysis on the sequence of the gene fragments obtained by the sequencing device 54 based on the reference genome of the near-source species and the near-source region of the near-source species acquired by the near-source region acquisition device 52.
  • Bioinformatics analysis includes, for example, assembly, single nucleotide polymorphism search, copy number variation detection, insertion/deletion detection, chromosome structural variation detection, SNP locus and disease association analysis, or SNP locus associated with drug action analysis.
  • the biological information processing system may further include a filtering device 65, The sequencer 54 obtains the sequence of the gene fragment for filtering to remove the unqualified gene fragment sequence, and the filtered gene slice is sent to the transmission analysis device 66 for subsequent processing.
  • the methods and systems of the present invention are currently preferred embodiments of genome-wide data, but without species samples of corresponding gene chips to achieve large-scale target region capture.
  • Figures 5 through 6 can be implemented by separate processing or computing devices, or integrated into a single device implementation. They are shown in boxes in Figures 5 to 6 to illustrate their function.
  • Some of the functional blocks can be implemented in hardware, software, firmware, middleware, microcode, hardware description speech, or any combination thereof.
  • one or both of the functional blocks can be implemented using code running on a microprocessor, digital signal processor (DSP), or any other suitable computing device.
  • a code can represent a procedure, a function, a subroutine, a program, a routine, a subroutine, a module, or any combination of instructions, data structures, or program statements.
  • the code can be located on a computer readable medium.
  • the computer readable medium can include one or more storage devices including, for example, RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, mobile hard disk, CD-ROM, or any other form known in the art. Storage medium.
  • the computer readable medium can also include a carrier wave that encodes the data signal.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Physics & Mathematics (AREA)
  • Molecular Biology (AREA)
  • Biotechnology (AREA)
  • Biophysics (AREA)
  • Analytical Chemistry (AREA)
  • Biochemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

本发明公开了一种目标区域捕获方法及其生物信息处理方法和系统。所述目标区域捕获方法包括利用目标区域捕获芯片捕获目标区域捕获芯片指定使用物种的近源物种的基因片段,并对所述近源物种的基因片段进行测序获得基因片段序列。本发明的方法和系统,将为单一物种设计的基因芯片用于其它近源物种的目标区域捕获,从而帮助筛选近源物种的全基因组数据。

Description

目标区域捕获方 及其 物信息处理方 系统 技术领域
本发明涉及生物信息学技术领域, 尤其涉及一种目标区域捕 获方法及其生物信息处理方法和系统。 背景技术
新一代测序技术的发展, 推动了模式物种的测序。 全基因组 测序成本一降再降, 越来越多物种的基因组序列被破解, 随之得 来的大量基因信息急需大规模、 高通量的分析方法和手段对其进 行进一步研究。
20世纪 90年代初 来的基因芯片技术是大 究基 因结构和功能的一种重要技术, 配合高通量测序技术可以实现对 某一物种特定目标区域, 尤其是外显子进行捕获并测序, 进而寻 找功能基因并对其进行后续生物信息学分析。 这种方法捕获效率 高, 成本相比全基因组测序较低, 准确性也非常好。 目前, 国外 学者利用目标区域捕获技术, 使用外显子捕获芯片对人进行外显 子测序, 进而开展疾病相关研究, 已取得一系列^½。
但是, 由于目前各大芯片公司所推出的芯片都是针对模式物 种的基因组来设计的, 有许多已经有全基因组数据的物种并没有 相应的芯片出售, 因此, 这些物种功能基因的获得仍然需要依赖 传统的 PCR (聚合酶链式反应), 经过大量的引物设计、 合成、 实验及优化, 费时费力, 工作量极大, 全基因组数据也没有被很 好的利用起来。 发明内容
^^开的一个方面要解决的一个技术问题是提供一种目标区 域捕获方法, 能够提高物种的目标区域捕获的效率。
开的一个方面提供一种目标区域捕获方法, 包括: 利用目标区域捕获芯片捕获目标区域捕获芯片指定使用物种 的近源物种的基因片段;
对近源物种的基因片段进行测序获得基因片 列。
根据本公开的目标区域捕获方法的一个实施例, 该方法还包 括: 获取目标区域捕获芯片指定使用物种参考基因组上的捕获区 域序列; 根据捕获区域序列获取近源物种参考基因组上的近源区 域。
根据本公开的目标区域捕获方法的一个实施例, 该方法还包 括: 对基因片 列进行过滤以去除不合格的基因片 列。
开提供的目标区域捕获方法, 通过将目标区域捕获芯片 应用于该芯片指定使用物种的近源物种的目标区域的捕获, 提高 了没有目标区域捕获芯片的物种的目标区域捕获的效率。
开的另一个方面要解决的技术问题是提供一种生物信息 处理方法和系统, 能够提高处理效率。
开的另一个方面提供一种生物信息处理方法, 包括: 利用目标区域捕获芯片捕获目标区域捕获芯片指定使用物种 的近源物种的基因片段;
对近源物种的基因片段进行测序获得基因片 列; 获取目标区域捕获芯片指定使用物种参考基因组上的捕获区 列;
根据捕获区域序列获取近源物种参考基因组上的近源区域; 根据近源物种参考基因组和近源物种的近源区域, 对基因片 段序列进行生物信息分析。
根据本公开的生物信息处理方法的一个实施例, 生物信息分 析包括: 组装、 单核苷酸多态性位点寻找、 拷贝数变异检测、 插 X/缺失检测、 染色体结构变异检测、 SNP位点与疾病关联分析、 或 SNP位点与药物作用相关分析。
φ ^开的又一个方面提供一种生物信息处理系统, 包括: 捕获区域序列获取装置, 用于获取目标区域捕获芯片指定使 用物种参考基因组上的捕获区域序列;
近源区域获取装置, 用于根据捕获区域序列获取指定使用物 种的近源物种参考基因组上的近源区域;
基因片段捕获装置, 用于利用目标区域捕获芯片捕获近源物 种的基因片段;
测序装置, 用于对近源物种的基因片段进行测序获得基因片 列。
根据 开的生物信息处理系统的一个实施例, 该系统还包 括过滤装置, 用于对测序装置获得的基因片段序列进行过滤以去 除不合格的基因片^ ^列。
根据 开的生物信息处理系统的一个实施例, 该系统还包 括信息分析装置, 用于根据近源物种的参考基因组和近源物种的 近源区域对基因片^ ^列进行生物信息分析。
开提供的生物信息处理方法和系统, 通过目标区域捕获 芯片实现近源物种的目标区域捕获, 根据芯片的区域捕获文件获 得近源物种参考基因组上的近源区域, 从而有助于后续的生物信 息分析, 提高了生物信息分析的效率。 附图说明
图 1 示出根据本发明目标区域捕获方法的一个实施例的流程 图;
图 2示出根据本发明目标区域捕获方法的另一个实施例的流 程图; 图 3示出根据本发明生物信息处理方法的一个实施例的流程 图;
图 4示出根据本发明的人基因组芯片捕获区域同源到猴子基 因组上的分布情况图;
图 5示出根据本发明的生物信息处理系统的一个实施例的结 构图;
图 6示出根据本发明的生物信息处理系统的另一个实施例的 结构图。 具体实施方式
下面参照附图对本发明进行更全面的描述, 其中说明本发明 的示例性实施例。
图 1 示出根据本发明基因组目标区域捕获方法的一个实施例 的流程图。
如图 1 所示, 在步骤 102, 利用目标区域捕获芯片捕获目标 区域捕获芯片指定使用物种的近源物种的基因片段。 近源物种之 间的遗传距离较近, 遗传物质的相似度高 (特别是在外显子这样 的高度保守区域上), 近源物种之间功能类似的基因外显子中总 会出现序列相似度极高甚至完全相同的保守区域, 从而使现有的 目标区域捕获芯片捕获其他近源物种的基因片段存在可能。 利用 指定物种的目标区域捕获芯片, 按照芯片操作手册对要研究的近 源物种(目标物种) 的 DNA样品进行目标区域捕获, 捕获近源 物种的基因片段。
在步骤 104, 对近源物种的基因片段进行测序获得基因片段 序列。 例如, 对捕获到近源物种的基因片段进行高通量测序获得 基因片段序列, 高通量测序技术可以采用 Illumina GA 测序技 术, 或者其他高通量测序技术。 上述实施例中, 通过将目标区域捕获芯片应用于近源物种的 目标区域捕获, 实现了近源物种的目标区域捕获, 将物种种类较 少的芯片更好的利用起来, 有助于获得更多物种个体功能基因及 其后续生物信息分析。
图 2示出根据本发明基因组目标区域捕获方法的另一个实施 例的流程图。
在步骤 202, 获取目标区域捕获芯片指定使用物种参考基因 组上的捕获区域序列。 例如, 根据芯片厂商提供的目标区域捕获 芯片的捕获区域文件, 在该芯片的指定使用物种的参考基因组上 截取对应的捕获区域序列。
在步骤 204, 根据捕获区域序列获取近源物种参考基因组上 的近源区域。 进一步, 根据近源物种的近源区域可以生成该目标 区域捕获芯片对该近源物种的捕获区域文件。 例如, 将得到的芯 片指定使用物种捕获区域序列, 利用比对软件 (例如, 适合于局 部比对的软件 Blast )与要研究的近源物种的参考基因组进行同源 性比对, 根据比对结果获取近源物种参考基因组上与捕获区域序 列对应的近源区域。 同源性比对过程允许出现错配, 可以允许出 现的错配率例如为 1 % ~ 20%、 或者 1%~10%, 或者 5%左右。 可根据不同的需求降低或升高错配率。 例如, 根据芯片指定使用 物种与目标的近源物种之间的遗传距离决定错配率的范围, 对于 遗传距离大的物种之间接受较大的错配率, 对于遗传距离小的物 种之间接受较小的错配率。 比对结果的内容例如包括: 每一段捕 获区域序列与近源物种的参考基因组的哪一片段同源、 同源性大 小、 这一片段在近源物种参考基因组上的具体坐标; 根据坐标信 息, 得到该芯片在近源物种参考基因组中的可能与捕获区域序列 对应的近源区域。 进一步, 可以根据近源物种的近源区域生成该 芯片对近源物种的捕获区域文件。 在步骤 206, 利用目标区域捕获芯片捕获目标区域捕获芯片 指定使用物种的近源物种的基因片段。
在步骤 208, 对近源物种的基因片段进行测序获得基因片段 序列。
需要指出, 步骤 202和步骤 204可以也可以在步骤 208后执 行, 或者与步骤 206和步骤 208并行执行。
生物在进化的过程中, 由于编码蛋白质的基因序列通常处于 选择压力之下, 因此, 不同物种的基因中都有基因序列非常保守 的区域, 使得原本是为单一物种设计的基因芯片用于其它近源物 种目标区域的捕获提供了可能。 但是, 目标区域捕获芯片仅适用 于指定使用物种的观念, 限制了普通技术人员去仔细分析为单一 物种设计的基因芯片用于其它近源物种目标区域的捕获的可行 性, 以及具体实现的技术方案。 目前为止, 尚未发现公开发表用 该思路做出来的研究成果。 而且, 上述可能性并不能使本领域的 技术人员自然地将为单一物种设计的基因芯片用于其它近源物种 目标区域的捕获, 还需要克服一些技术问题以应用于生物信息分 析。 上述实施例中, 通过目标区域捕获芯片提供的捕获区域文件 帮助筛选近源物种的全基因组数据, 从而获得近源物种的近源区 域, 进而获得目标区域捕获芯片对于该近源物种的捕获区域序 列, 大大减少了后续比对工作的工作量。
上述实施例中, 利用同源性比对找到近源物种的捕获区域, 从而可以精细定位捕获下来的序列片段在近源物种的参考基因组 上的位置, 方便后续分析, 克服了通过目标区域捕获芯片获得近 源物种的测序序列后无法进行后续分析的障碍。
根据本发明的一个实施例, 在对近源物种的基因片段进行测 序获得基因片段序列后, 还对基因片 列进行过滤以去除不合 格的基因片 列。 不合格的基因片 列例如包括: 测序质量 低于低质量阈值的碱基个数超过整条序列碱基个数预定百分比
(例如, 40 %以上, 50%以上) 的序列, 低质量阈值由具体测序 技术及测序环境而定; 序列中测序结果不确定的碱基 (如 IUumina GA测序结果中的 N )个数超过整条序列碱基个数预定 比例 (例如 10%, 20 % ) 的序列; 存在外源序列的基因片段序 列, 与其它实验引入的外源序列比对, 如各种接头序列, 若序列 中存在外源序列则认为是不合格序列。 将剩余合格的短片段序列 中的接头序列去除。
在获得近源物种的基因片段序列后, 可以根据近源物种的参 考基因组和近源物种的近源区域, 对基因片 列进行生物信息 分析。
图 3示出根据本发明生物信息处理方法的一个实施例的流程 图。
如图 3 所示, 在步骤 302, 利用指定物种的目标区域捕获芯 片, 按照芯片操作手册对近源物种的 DNA样品进行目标区域捕 获, 获得近源物种的基因片段。
在步骤 304, 对捕获到的基因片段进行高通量测序获得基因 片段序列。 其中, 高通量测序技术可以为 IUumina GA 测序技 术, 也可以为现有的其他高通量测序技术。
在步骤 306, 接收测序数据(基因片 列), 进行测序数据 预处理。 例如对基因片 列进行过滤。
在步骤 308, 根据芯片厂商提供的目标区域捕获芯片的捕获 区域文件, 在该芯片的指定使用物种的参考基因组上截取对应的 捕获区 列。
在步骤 310, 将上述步骤中得到的芯片指定使用物种捕获区 域序列, 利用适合于局部比对的软件如 Blast, 与想要研究的近源 物种的参考基因组进行同源性比对, 得到比对结果。 在步骤 312, 生成该芯片对近源物种的捕获区域文件。
在步骤 314, 根据目标物种的参考基因组及获得的目标物种 的捕获区域文件, 对测序序列进行后续的生物信息分析。 生物信 息分析包括但不限于组装、 单核苷酸多态性(SNP )位点寻找、 拷贝数变异(CNV )检测、 插入 /缺失(Indel )检测、 染色体结 构变异(SV )检测、 SNP位点与疾病的关联分析、 SNP位点与 药物作用的相关分析。
下面介绍本发明生物信息处理方法的一个应用例。 在该应用 例中:
近源物种 ( 目标物种) 样本: 2 个食蟹猴 ( crab-eating macaque, 学名: Macaca fascicularis ) 的 DNA样本。
目标区域捕获芯片: 安捷伦公司 (38M ) 人外显子芯片 (该芯片的商品名: Agilent 2100 Bioanalyzer , 货号 /型号: G2938A )
芯片指定使用物种的设计对象: 人基因组 (Hgl8, 即 NCBI BUILD36) , 可 以 访 问 网 页 http://www.ncbi.nlm.nih.gov/sites/genome/ , 下面的 ftp下 载 地 址 下 载 hgl8 数 据 : ftp://ftp.ncbi.nlm.nih.gOv/genomes/H sapiens/ARCHIVE/
近源物种参考基因组: 食蟹猴 (CE )参考基因组(数据下 载地址: http://climb.genomics.cn/Macaca_fascicularis )
实施例具体操作流程:
( 1 )跨物种捕获并测序
利用安捷伦公司(38M)人外显子芯片, 对食蟹猴的 DNA样 品进行目标区域捕获及 Illumina GA测序, 接收测序序列, 即基 因片 列;
( 2 )过滤测序序列 接收到高通量测序序列后, 对测序序列进行过滤, 去除不合 格的序列。 不合格序列包括: 测序质量值低于 20 的碱基个数超 过整条序列 个数的 50%则认为是不合格序列; 序列中测序结 果不确定的碱基 ( IUumina GA测序结果中的 N )个数超过整条 序列 基个数的 10%则认为是不合格序列; 除样本接头序列外, 与其它实验引入的外源序列比对, 如各种接头序列。 若序列中存 在外源序列则认为是不合格序列。
( 3 )获取芯片捕获区域序列
根据芯片厂商安捷伦公司提供的该目标区域捕获芯片的捕获 区域文件, 在人的参考基因组 ( Hgl8 )上截取对应的捕获区域序 列, 将其写入到一个 fastq文件 exoncapturcfa中, 其格式为:
>chrl.txt:20138:20258
CCAAAGTCCAGCAGTTGTCCCTCCTGGAATCCGTTG GCTTGCCTCCGGCATTTTTGGCCCTTGCCTTTTAGGGTTG CCAGATTAAAAGACAGGATGCCCAGCTAGTTTGAATTTTA GATAA
>chrl.txt:58932:59892
AACGAGTGAAACGAATAACTCTATGGTGACTGAATTC ATTTTTCTGGGTCTCTCTGATTCTCAGGAACTCCAGACCT TCCTATTTATGTTGTTTTTTGTATTCTATGGAGGAATCGTG TTTGGAAACCTTCTTATTGTCATAACAGTGGTATCTGACT CCCACCTTCACTCTCCCAT
其中 >所在的一行代表该捕获区域序列在 Hgl8上处于哪条染 色体, 以及它的起始、 终止坐标, 下面的一行是这一段捕获区域 序列的具体基因型。
( 4 ) 芯片捕获区域序列与近源物种参考基因组的同源性比 对 将得到的上述捕获区域序列文件 exoncapturcfa, 利用 Blast 比对软件, 与食蟹猴的参考基因组 CE.fa进行同源性比对, 比对 结果生成文件 out.psl, 其主要内容如表 1所示:
Figure imgf000011_0002
Figure imgf000011_0001
其中第一、 二行表示同源性比对时候比对上与未比对上位点 的个数, 第十列表示捕获区域序列名称, 第十四至十七列表示与 该捕获区域同源的片段在食蟹猴参考基因组 CE.fa上的具体染色 体以及该染色体长度、 同源片段起始、 终止坐标。 总结起来, 表 1 所代表的意义即为: 芯片设计的在人参考基因组的 1 号染色体 的 20138 bp到 20258 bp这一段区域, 与食蟹猴参考基因组 13号 染色体上 113205694 bp到 113205813 bp这一段区域同源, 食蟹 猴参考基因组的 13号染色体的总长为 137686314 bp。
根据芯片的捕获区域, 截取 hgl8基因组片段, 得到理论上 的芯片捕获片段, 将其与食蟹猴参考基因组做同源分析。 结果发 现 97.7%的片段可以在食蟹猴参考基因组上找到同源区域, 这种 同源性表现为整段外显子片段的同源。 剩下的 2.3%找不到同源 区域, 估计可能是人类在进化中产生的特有片段。
图 4示出人基因组芯片捕获区域同源到猴子基因组上的分布 情况, 直观地展示了芯片能抓到食蟹猴的基因片段位于猴子染色 体的哪些地方 (其中, 能抓到的地方用灰色表示)。 很明显灰色 区域分布在猴子所有染色体上, 而且比较均匀, 说明人基因组外 显子芯片是可以捕获到猴子各条染色体上的大量外显子片段。
( 5 ) 生成芯片对近源物种的捕获区域文件
根据同源性比对结果文件 out.psl, 得到安捷伦 38M 芯片在 食蟹猴的参考基因组上的捕获区域, 生成安捷伦 38M 芯片对食 蟹猴的参考基因组的捕获区域文件, 并按照染色体分开, 生成文 件 capture_region_chr*.txt, 具体格式为:
染色体号 起始坐标 终止坐标
chrlO 15525 16380
chrlO 21092 21333
chrlO 23712 23932
chrlO 24665 24786
chrlO 28859 29220
chrlO 50626 50875
chrlO 51197 51438 chrlO 53645 53766
表 2
其中, 第一列表示染色体号, 第二列表示起始坐标, 第三列 表示终止坐标。
( 6 )后续生物信息分析
将第一步得到的测序序列 fastq 文件作为输入文件, 以食蟹 猴的参考基因组 CE.fa 为参考序列, 以第 5 步得到的 capture region chr*.txt 文件作为捕获区域, 利用外显子捕获流 程 ECP ( Exon Capture Pipeline )用短序列比对软件 SOAP进行 比对, 得到结果文件 pe.soap。
根据跑完 ECP流程生成的 samplcreport文件, 得到
Figure imgf000013_0001
子的测序数据质量如下:
Figure imgf000013_0002
表 3
上表中 coverage是指实际捕捉的片段能覆盖理论上芯片捕获 区域的百分数, 即用人基因组芯片捕获到食蟹猴基因组上的片段 能覆盖 85%的芯片设计捕获区域。 因为跑 ECP 时全部用的 是人的参数, 条件要求很高, 只有比对得非常好而且比对上的次 数足够多的才算成是覆盖到了, 85%的覆盖度, 说明以现在的实 臉条件, 有 85%的芯片设计区域是可以很好的捕获下来的食蟹猴 的相关基因区域。
综上分析, 可以确定用针对人基因组的外显子芯片可以高效 的获取大量食蟹猴的外显子片段, 测序质量也较为稳定。
根据得到的结果与食蟹猴的参考基因组比对进行后续各种生 物信息学分析。 上述应用例中, 通过跨物种序列同源性比对, 得到了人的目 标区域捕获芯片对于其他物种的捕获区域, 实现了一种跨物种目 标区域捕获的方法, 对人的近源物种进行目标区域捕获且捕获质 量很高, 从而将低成本、 高准确性的目标区域捕获的研究方法扩 展到了跨物种的层次。
近源物种分析一般限于人-黑猩猩这样遗传距离非常近的物种 间的研究, 但本发明方法的适用范围要宽得多, 上述应用例中用 到的人 -食蟹猴的遗传距离就相对较远, 甚至还可以将本发明的方 法应用到人-小鼠的研究, 是对本领域常规做法的突破。
图 5 示出才据本发明的生物信息处理系统的一个实施例的结 构图。 如图 5所示, 该生物信息处理系统包括: 捕获区域序列获 取装置 51, 用于获取目标区域捕获芯片指定使用物种参考基因组 上的捕获区域序列; 近源区域获取装置 52, 根据捕获区域序列获 取装置 51 获取的捕获区域序列获取指定使用物种的近源物种参 考基因组上的近源区域; 基因片段捕获装置 53, 利用目标区域捕 获芯片捕获近源物种的基因片段; 测序装置 54对近源物种的基 因片段进行测序获得基因片 列。
图 6示出根据本发明的生物信息处理系统的另一个实施例的 结构图。 和图 5相比, 该实施例中的生物信息处理系统除了包括 捕获区域序列获取装置 51、 近源区域获取装置 52、 基因片段捕 获装置 53, 和测序装置 54, 还包括信息分析装置 66。 信息分析 装置 66根据近源物种的参考基因组和近源区域获取装置 52获取 的近源物种的近源区域对测序装置 54获得的基因片段序列进行 生物信息分析。 生物信息分析例如包括: 组装、 单核苷酸多态性 位点寻找、 拷贝数变异检测、 插入 /缺失检测、 染色体结构变异检 测、 SNP位点与疾病关联分析、 或 SNP位点与药物作用相关分 析。 可选地, 该生物信息处理系统还可以包括过滤装置 65, 对测 序装置 54 获得基因片段序列进行过滤以去除不合格的基因片段 序列, 将过滤后的基因片 列发送 息分析装置 66 进行后 续处理。
本发明的方法和系统是目前已有全基因组数据, 但没有相应 基因芯片的物种样品实现大规模目标区域捕获的优选技术方案。
对于图 5至图 6中各个装置或单元的功能, 可以参考上文中 关于本发明方法的实施例中对应部分的说明, 为简洁起见, 在此 不再详述。
本领域的技术人员应当理解, 对于图 5 至图 6 中的各个装 置, 可以通过单独的处理或计算设备实现, 或者将其集成为一个 独立的设备实现。 在图 5至图 6中用框示出以说明它们的功能。 部分功能块可以用硬件、 软件、 固件、 中间件、 微代码、 硬件描 述语音或者它们的任意组合来实现。 举例来说, 一个或者两个功 能块都可以利用运行在微处理器、 数字信号处理器(DSP )或任 何其他适当计算设备上的代码实现。 代码可以表示过程、 功能、 子程序、 程序、 例行程序、 子例行程序、 模块或者指令、 数据结 构或程序语句的任意组合。 代码可以位于计算机可读介质中。 计 算机可读介质可以包括一个或者多个存储设备, 例如, 包括 RAM存储器、 闪存存储器、 ROM存储器、 EPROM存储器、 EEPROM存储器、 寄存器、 硬盘、 移动硬盘、 CD-ROM或本领 域公知的其他任何形式的存储介质。 计算机可读介质还可以包括 编码数据信号的载波。
本领域技术人员将意识到硬件、 固件和软件配置在这些情况 下的可替换性, 以及如何最好地实现每个特定应用地该功能。
本发明的描述是为了示例和描述起见而给出的, 而并不是无 遗漏的或者将本发明限于所公开的形式。 很多修改和变化对于本 领域的普通技术人员而言是显然的。 选择和描述实施例是为了更 好说明本发明的原理和实际应用, 并且使本领域的普通技术人员 能够理解本发明从而设计适于特定用途的带有各种修改的各种实 施例。

Claims

权 利 要 求
1. 一种目标区域捕获方法, 其特征在于, 包括:
利用目标区域捕获芯片捕获所述目标区域捕获芯片指定使用 物种的近源物种的基因片段;
对所述近源物种的基因片段进行测序获得基因片^^列。
2. 根据权利要求 1所述的基因组目标区域捕获方法, 其特征 在于, 还包括:
获取所述目标区域捕获芯片指定使用物种参考基因组上的捕 获区域序列;
根据所述捕获区域序列获取所述近源物种参考基因组上的近 源区域。
3. 根据权利要求 2所述的方法, 其特征在于, 还包括: 对所述基因片段序列进行过滤以去除不合格的基因片段序 列。
4. 根据权利要求 3所述的方法, 其特征在于, 所述对所述基 因片段序列进行以去除不合格的基因片段序列包括:
去除测序质量低于低质量阈值的碱基个数超过基因片段序列 全部 个数预定百分比的基因片 列;
和 /或
去除测序结果不确定的碱基个数超过基因片段序列全部碱基 个数预定比例的基因片 列;
和 /或
去除存在外源序列的基因片 列。
5. 根据权利要求 2所述的方法, 其特征在于, 所述获取目标 区域捕获芯片指定使用物种参考基因组上的捕获区域序列包括: 根据所述目标区域捕获芯片的捕获区域文件在所述目标区域 捕获芯片指定使用物种的参考基因组上获取所述捕获区域序列。
6. 根据权利要求 2所述的方法, 其特征在于, 所述根据所述 捕获区域序列获取所述近源物种参考基因组上的近源区域包括: 将所述捕获区域序列与所述近源物种的参考基因组进行同源 比对, 获得所述近源物种参考基因组上的近源区域。
7. 根据权利要求 2所述的方法, 其特征在于, 还包括: 根据所述近源物种参考基因组上的近源区域生成所述目标区 域捕获芯片对所述近源物种的捕获区域文件。
8. 根据权利要求 1所述的方法, 其特征在于, 所述利用目标 区域捕获芯片捕获所述目标区域捕获芯片指定使用物种的近源物 种的基因片段包括:
按照所述目标区域捕获芯片的操作手册对所述近源物种的 DNA样品进行目标区域捕获, 获得所述近源物种的基因片段。
9. 一种基于目标区域捕获的生物信息处理方法, 其特征在 于, 包括权利要求 2至 8中任意一项所述的方法, 还包括:
根据所述近源物种参考基因组和所述近源物种的近源区域, 对所述基因片 列进行生物信息分析。
10. 根据权利要求 9所述的生物信息处理方法, 其特征在 于, 所述生物信息分析包括: 组装、 单核苷酸多态性位点寻找、 拷贝数变异检测、 插入 /缺失检测、 染色体结构变异检测、 SNP位 点与疾病关联分析、 或 SNP位点与药物作用相关分析。
11. 一种生物信息处理系统, 其特征在于, 包括:
捕获区域序列获取装置, 用于获取所述目标区域捕获芯片指 定使用物种参考基因组上的捕获区域序列;
近源区域获取装置, 用于根据所述捕获区域序列获取所述指 定使用物种的近源物种参考基因组上的近源区域;
基因片段捕获装置, 用于利用目标区域捕获芯片捕获所述近 源物种的基因片段;
测序装置, 用于对所述近源物种的基因片段进行测序获得基 因片 列。
12. 根据权利要求 11所述的系统, 其特征在于, 还包括: 过滤装置, 用于对所述测序装置获得所述基因片段序列进行 过滤以去除不合格的基因片 列。
13. 根据权利要求 11所述的系统, 其特征在于, 还包括: 信息分析装置, 用于根据所述近源物种的参考基因组和所述 近源物种的近源区域对所述基因片段序列进行生物信息分析。
14. 根据权利要求 13所述的系统, 其特征在于, 所述生物信 息分析包括: 组装、 单核苷酸多态性位点寻找、 拷贝数变异检 测、 插 X/缺失检测、 染色体结构变异检测、 SNP位点与疾病关联 分析、 或 SNP位点与药物作用相关分析。
PCT/CN2011/077861 2011-08-01 2011-08-01 目标区域捕获方法及其生物信息处理方法和系统 Ceased WO2013016864A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/CN2011/077861 WO2013016864A1 (zh) 2011-08-01 2011-08-01 目标区域捕获方法及其生物信息处理方法和系统
CN201180071091.2A CN103547681B (zh) 2011-08-01 2011-08-01 目标区域捕获方法及其生物信息处理方法和系统

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2011/077861 WO2013016864A1 (zh) 2011-08-01 2011-08-01 目标区域捕获方法及其生物信息处理方法和系统

Publications (1)

Publication Number Publication Date
WO2013016864A1 true WO2013016864A1 (zh) 2013-02-07

Family

ID=47628620

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2011/077861 Ceased WO2013016864A1 (zh) 2011-08-01 2011-08-01 目标区域捕获方法及其生物信息处理方法和系统

Country Status (2)

Country Link
CN (1) CN103547681B (zh)
WO (1) WO2013016864A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110592208A (zh) * 2019-10-08 2019-12-20 北京诺禾致源科技股份有限公司 地中海贫血症三类亚型的捕获探针组合物及其应用方法和应用装置

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108733974B (zh) * 2017-04-21 2021-12-17 胤安国际(辽宁)基因科技股份有限公司 一种基于高通量测序的线粒体序列拼接及拷贝数测定的方法
CN108048916A (zh) * 2017-12-20 2018-05-18 栾图 用于核酸富集提取的基因芯片及其制备方法
CN110491448B (zh) * 2019-07-15 2023-02-07 广州奇辉生物科技有限公司 一种处理pcr引物的方法、系统、平台及存储介质
CN115323046B (zh) * 2022-09-19 2025-06-13 河南大学 一种麦类d基因组泛基因外显子捕获测序芯片及其设计方法

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
ALKAN, C. ET AL.: "Personalized copy number and segmental duplication maps using next-generation sequencing.", NAT GENET., vol. 41, no. 10, October 2009 (2009-10-01), pages 1061 - 1067 *
BURBANO, H.A. ET AL.: "Targeted investigation of the Neandertal genome by array-based sequence capture.", SCIENCE., vol. 328, no. 5979, May 2010 (2010-05-01), pages 723 - 725 *
GEORGE, R.D. ET AL.: "Trans genomic capture and sequencing of primate exomes reveals new targets of positive selection.", GENOME RES., vol. 21, no. 10, July 2011 (2011-07-01), pages 1686 - 1694 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110592208A (zh) * 2019-10-08 2019-12-20 北京诺禾致源科技股份有限公司 地中海贫血症三类亚型的捕获探针组合物及其应用方法和应用装置
CN110592208B (zh) * 2019-10-08 2022-05-03 北京诺禾致源科技股份有限公司 地中海贫血症三类亚型的捕获探针组合物及其应用方法和应用装置

Also Published As

Publication number Publication date
CN103547681A (zh) 2014-01-29
CN103547681B (zh) 2015-03-11

Similar Documents

Publication Publication Date Title
Lowe et al. Transcriptomics technologies
US12065691B2 (en) Recovering long-range linkage information from preserved samples
US11492656B2 (en) Haplotype resolved genome sequencing
Wang et al. Application of next generation sequencing to human gene fusion detection: computational tools, features and perspectives
Bi et al. Transcriptome-based exon capture enables highly cost-effective comparative genomic data collection at moderate evolutionary scales
Birzele et al. Into the unknown: expression profiling without genome sequence information in CHO by next generation sequencing
JP2020058393A (ja) 母体血漿の無侵襲的出生前分子核型分析
AU2013334958B2 (en) HLA typing using selective amplification and sequencing
JP2018509928A (ja) 環状化メイトペアライブラリーおよびショットガン配列決定を用いて、ゲノム変異を検出するための方法
WO2013016864A1 (zh) 目标区域捕获方法及其生物信息处理方法和系统
JP2021532826A (ja) シーケンスリードの独立したアラインメントおよびペアリングによって高度に相同なシーケンスにおける遺伝的変異を検出するための方法
Bai et al. Improving the genome assembly of rabbits with long-read sequencing
JP2024512372A (ja) オフターゲットポリヌクレオチド配列決定データに基づく腫瘍の存在の検出
CN111696628A (zh) 新生抗原的鉴定方法
Al-Haggar et al. Bioinformatics in high throughput sequencing: application in evolving genetic diseases
WO2021173502A1 (en) Systems and methods for identifying adaptive immune cell clonotypes
Lesur et al. A strategy for studying epigenetic diversity in natural populations: proof of concept in poplar and oak
Ding et al. EAnnot: a genome annotation tool using experimental evidence
WO2021050717A1 (en) Immune cell sequencing methods
Gaur et al. A survey of bioinformatics-based tools in RNA-sequencing (RNA-seq) data analysis
Amr et al. Targeted hybrid capture for inherited disease panels
Cook et al. Long Read Annotation (LoReAn): automated eukaryotic genome annotation based on long-read cDNA sequencing
Jiang et al. High-performance single-chip exon capture allows accurate whole exome sequencing using the Illumina Genome Analyzer
Goya et al. Applications of high-throughput sequencing
Jiang et al. Establishment of an Integrated Computational Workflow for Single Cell RNA-Seq Dataset

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 11870280

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 11870280

Country of ref document: EP

Kind code of ref document: A1