WO2024066461A1 - 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法 - Google Patents

基于宏基因组和宏转录组识别油藏驱油功能微生物的方法 Download PDF

Info

Publication number
WO2024066461A1
WO2024066461A1 PCT/CN2023/099026 CN2023099026W WO2024066461A1 WO 2024066461 A1 WO2024066461 A1 WO 2024066461A1 CN 2023099026 W CN2023099026 W CN 2023099026W WO 2024066461 A1 WO2024066461 A1 WO 2024066461A1
Authority
WO
WIPO (PCT)
Prior art keywords
quality
metagenome
metatranscriptome
mags
sequence
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/099026
Other languages
English (en)
French (fr)
Inventor
牟伯中
刘一凡
寿利斌
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
East China University of Science and Technology
Original Assignee
East China University of Science and Technology
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by East China University of Science and Technology filed Critical East China University of Science and Technology
Publication of WO2024066461A1 publication Critical patent/WO2024066461A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/02Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving viable microorganisms
    • C12Q1/04Determining presence or kind of microorganism; Use of selective media for testing antibiotics or bacteriocides; Compositions containing a chemical indicator therefor
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6869Methods for sequencing
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B10/00ICT specially adapted for evolutionary bioinformatics, e.g. phylogenetic tree construction or analysis
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/20Sequence assembly
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B50/00ICT programming tools or database systems specially adapted for bioinformatics
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B50/00ICT programming tools or database systems specially adapted for bioinformatics
    • G16B50/10Ontologies; Annotations

Definitions

  • the present invention relates to the technical field of microbial detection, and in particular to a method for identifying microorganisms with oil recovery function in oil reservoirs based on metagenome and metatranscriptome.
  • Its main mechanism is to reduce the viscosity of heavy oil and improve the efficiency of water flooding through the five functions of microorganisms in the oil reservoir, such as emulsifying crude oil, producing gas, producing surfactants, producing polysaccharides, and degrading hydrocarbons.
  • sequencing methods can be roughly divided into sequencing technology that relies on PCR amplification and metagenomic and metatranscriptomic sequencing technology that does not rely on PCR amplification.
  • the former is represented by the 16S rRNA gene clone library method and has been widely used in reservoir environmental samples.
  • the relevant gene sequences in the sample can be amplified, thereby explaining the potential metabolic functions of microorganisms at the gene level. The latter does not require PCR amplification and can simultaneously determine the sequence information of all genes in the sample.
  • metagenomic sequencing to analyze reservoir environmental samples can deeply analyze the potential metabolic network in the sample. Further combining metagenomic technology and metatranscriptomic technology can obtain the transcription level of each gene in the metabolic pathway, thereby inferring various microbial metabolic processes in the reservoir environment.
  • the purpose of the present invention is to solve the problem of limited information that can be obtained from the existing reservoir microbial detection methods, and to provide a method for accurately identifying microorganisms and metabolic functions in the reservoir by combining metagenomics and metatranscriptomics.
  • the object of the present invention is to provide a method for identifying functional microorganisms for oil recovery in oil reservoirs based on metagenome and metatranscriptome, comprising the following steps:
  • an inhibitor is added to the sample from which RNA is to be extracted to inhibit RNA degradation by anaerobic microorganisms.
  • S2 also includes: preprocessing the original data of the metagenome and metatranscriptome of the reservoir sample to obtain target data with the connectors and low-quality fragments removed;
  • the pre-treatment process includes:
  • Fastp software was used to perform sliding window quality trimming on the raw data of both ends of each DNA or RNA chain sequence in the metagenome and metatranscriptome of the reservoir samples.
  • cutadapt software was used to remove the primers to obtain the double-end sequence data after quality control.
  • the parameters for sliding window quality trimming are -W 4, -M 20, that is, the sliding window size is 4 and the average quality value is 20.
  • the process of analyzing the metagenome and metatranscriptome results includes:
  • the double-end sequence data after quality control are assembled, binned and quality evaluated, and high-quality MAGs (metagenomic assembled genome) data sets are extracted after removing redundancy;
  • the high-quality MAGs assembly data were annotated according to the constructed reference database, the functional genes and MAGs with different oil recovery functions were identified, and the corresponding sequencing depth and relative abundance were calculated;
  • the short sequence fragments obtained by macrotranscriptome sequencing were quality-controlled and filtered, and then compared with the MAGs data to calculate the transcription level of each gene.
  • the process of assembling, boxing and quality assessment includes:
  • bowtie2 software was used to align the quality-controlled MAGs short sequence data information with the long sequence information to obtain the sequencing coverage information of different long sequences.
  • the three binning methods of Maxbin2, Metabat2, and CONCOCT were further used to separate the genomes of dominant bacteria from MAGs, and then further imported into the DAS_Tool program for evaluation, and finally the quality of genomes generated by different methods was integrated and improved;
  • dRep software was used to remove redundancy from genomes with high similarity in MAGs after quality control, and the CheckM tool was used to estimate the integrity and contamination of the genome based on the presence and number of single-copy marker genes in the genome.
  • the process of annotating high-quality MAGs assembly data includes:
  • the spliced long sequences were translated into coding protein sequences (CDS) using the Prodigal program (predicted open reading frames) and submitted to the KEGG database, and the ghostKOALA tool was used for functional annotation and to obtain KO numbers representing different paralogous subunits;
  • the local software KofamKOALA was used to annotate the KO number to each protein sequence according to the hidden Markov model (HMM) of each KO paralogous family protein and the recommended confidence criteria;
  • EggNOG emapper 2 tool was used to annotate the protein sequences with COG numbers, which were then converted into KO numbers, so that the final KO number annotation of each protein was in the following order: 1) ghostKOALA KO, 2) KofamKOALA KO, 3) EggNOG emapper KO.
  • the oil displacement function includes one or more of alkane degradation, gas production, emulsifier production, surfactant production, and polysaccharide production.
  • the process of identifying functional genes and MAGs with different oil recovery functions and calculating the corresponding sequencing depth and relative abundance by combining the sequencing coverage information of each sequence obtained by comparison with bowtie2 software includes:
  • the functional genes with missing information in the database included one or more of the anaerobic hydrocarbon degradation initial activation genes AssA, EbdA, AhyA, AncA, and bacterial microcompartment protein cluster genes.
  • the process of performing evolutionary relationship analysis on high-quality MAGs datasets includes:
  • the target sequence and similar reference sequences in the database were downloaded and merged into files.
  • the merged sequences were first aligned on MAFFT, and the conservative sites were selected with a threshold of 80%.
  • IQ-tree the algorithm uses the maximum likelihood method to construct phylogenetic trees was used to align the sequences pairwise and generate a maximum likelihood evolutionary tree.
  • the process of quality-control filtering the short fragments of the sequence obtained by macrotranscriptome sequencing and comparing them with the MAGs data to calculate the transcription level of each gene includes:
  • Bowtie2 software was used to align the short cDNA fragments in the high-quality MAGs dataset to the long DNA fragments obtained by assembling the metagenomes after binning in S3, and the gene expression of each gene was calculated.
  • the transcription level is expressed as TPM (Transcripts Per Million).
  • This technical solution is a microbial analysis method based on metagenomics and metatranscriptomics developed specifically for oil reservoir environments.
  • total DNA and total RNA are extracted from water samples produced from the oil reservoir, and sequencing is used to obtain the original data of the metagenomics and metatranscriptomes of the oil reservoir samples.
  • sequencing is used to obtain the original data of the metagenomics and metatranscriptomes of the oil reservoir samples.
  • microorganisms with oil recovery function are identified.
  • the overall process does not rely on traditional, slower means of isolation and identification of single microorganisms, and is suitable for processing samples with a large number of unknown species. It can also detect species with extremely low abundance, and the detection is comprehensive and rapid.
  • the method for identifying oil reservoir oil recovery functional microorganisms of the present invention is based on metagenomic and metatranscriptomic sequencing and is specifically comprised of the following steps:
  • Step s1 Collect environmental samples, add RNALater reagent to the samples to inhibit RNA degradation, and then extract DNA and RNA using DNA and RNA extraction kits respectively.
  • Step s2 Perform metagenomic and metatranscriptomic sequencing.
  • Step s3 Analysis of metagenomic and metatranscriptomic sequencing results, including the following steps:
  • Step s31 Use fastp to perform quality inspection on the short nucleic acid fragments obtained by sequencing and remove the sequences with lower quality and the sequencing adapters involved.
  • Step s32 Use the splicing program SPADes in the ‘Meta’ mode to splice the sample short sequences into ‘contigs’ of different lengths, and then connect the different ‘contigs’ into ‘scaffolds’ with sequencing gaps based on the information of double-end sequencing.
  • Step s33 Use the Prodigal program to translate the spliced long sequence into a coding protein sequence (CDS), and use ghostKOALA, KofamKOALA and EggNOG emapper 2 to annotate the protein sequence with COG numbers, which are then converted into KO numbers.
  • the final KO number annotation for each protein uses the following order: 1) ghostKOALA KO, 2) KofamKOALA KO and 3) EggNOG emapper KO.
  • Step s34 For hydrogen reductase, first use the local HMM model of the hydrogen reductase subgroup Type comparison was used to find out potential functional gene proteins. These protein sequences were then submitted to the online analysis software of the HydDB database to further divide the subtypes of hydrogen reductase. Regarding the potential functions of encoding secondary metabolites in the genome, the genome sequences were submitted to the AntiSMASH website, and different tools were combined to find the genomes that potentially encode secondary metabolites in the genome and predict the types of metabolites.
  • Step s35 Use Maxbin2, Metabat2, and CONCOCT Bining methods to separate genomes from the data, import them into the DAS_Tool program for evaluation, and merge and extract high-quality genomes obtained by different methods.
  • Step s36 First, the target sequence and similar reference sequences in the database are downloaded and merged into files. The merged sequences are first aligned on MAFFT, and conservative sites are selected with a threshold of 80%. Then, IQ-tree is used to align the sequences pairwise and generate a maximum likelihood evolutionary tree.
  • Step s37 For the short mRNA sequence fragments obtained by macrotranscriptome sequencing, after quality control, software such as bowtie2 is used to align the short cDNA fragments to the long DNA fragments obtained by macrogenome splicing, and the transcription level of each gene is calculated, expressed as TPM value (Transcripts Per Million).
  • RNALater 1% of the total volume of RNALater to the barrel of RNA sample in advance, and ensure that the reservoir sample fills the entire barrel volume to exclude air when sampling. Store at low temperature before extracting nucleic acid.
  • DNA and RNA were extracted using PowerSoil Total DNA Kit (QIAGEN, USA) and PowerSoil Total RNA Kit (QIAGEN, USA), respectively, and RNA was purified using RNase-Free DNase set (QIAGEN, USA). After cDNA was synthesized, it was stored at -80°C.
  • NEB Ultra TM DNA Library Prep Kit for The sequencing library with index codes was constructed using reagents from New England Biolabs (USA). The quality of the library was evaluated using Qubit 3.0 Fluorometer (Life Technologies, USA) and Agilent 4200 (Agilent, Canada). Finally, the library was sequenced on the Illumina Hiseq X-ten platform, and a 150 bp double-terminal base sequence.
  • the raw data were quality controlled and filtered by using fastp.
  • Metagenomic data were assembled by SPADes and binned by Maxbin2, Metabat2 and CONCOCT.
  • the obtained genome was imported into DAS_Tool to obtain a high-quality genome.
  • Prodigal was used to translate long fragments into protein coding sequence CDS, and ghostKOALA, KofamKOALA and EggNOG emapper 2 were used to annotate protein sequences.
  • the hydrogen production capacity and secondary metabolite synthesis capacity of the genome were analyzed by AntiSMASH of the HydDB database, and the anaerobic hydrocarbon degradation capacity was analyzed by a self-built local database.
  • the metatranscriptome data were aligned to the long DNA fragments of the metagenome splicing using Bowtie2, and the transcription level TPM of each gene was calculated.
  • Combining the relative abundance of genomes in samples and the corresponding metabolic pathways can accurately identify the main microorganisms and metabolic functions in the reservoir, and combining transcriptome data can obtain the expression activity of these microorganisms.
  • the data in Table 1 are the sequencing coverage and metabolic pathways of each genome. From the table, it can be seen that the main microorganisms in the sample are Kerstersia_gyiorum and Serratia_nematodiphila of the Proteobacteria and Enterococcus of the Firmicutes. Most of the microorganisms in the sample are facultative anaerobic microorganisms with nitrate reduction ability, and aerobic hydrocarbon degrading bacteria for different chain lengths are present in the sample. Nearly half of the microorganisms have a complete metabolic pathway for synthesizing lipopeptides. In addition, most hydrogenotrophic microorganisms also have the ability to produce hydrogen. The above information can provide a theoretical basis for decision-making in oil reservoir recovery and corrosion prevention.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biophysics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Biotechnology (AREA)
  • General Health & Medical Sciences (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Medical Informatics (AREA)
  • Evolutionary Biology (AREA)
  • Organic Chemistry (AREA)
  • Theoretical Computer Science (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Analytical Chemistry (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Genetics & Genomics (AREA)
  • Molecular Biology (AREA)
  • Bioethics (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Animal Behavior & Ethology (AREA)
  • Toxicology (AREA)
  • Physiology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,包括以下步骤:S1:自油藏产出水样提取总DNA和总RNA;S2:对获得的总DNA和总RNA进行测序,获取油藏样品的宏基因组和宏转录组原始数据;S3:通过对宏基因组和宏转录组结果进行分析,识别具有驱油功能的微生物。

Description

基于宏基因组和宏转录组识别油藏驱油功能微生物的方法 技术领域
本发明涉及微生物检测技术领域,尤其是涉及一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法。
背景技术
我国稠油资源量约为198.7亿吨,年产量高达3087万吨(2017年),已占原油总产量的16.2%。由于稠油粘度高(50-10000mPa.s),在油藏中流动性差,一般以蒸汽热采方法为主,而热采法耗能大、成本高、开采效果差。此外,油藏是一个天然生物反应发生器,同时蕴含了具有各种功能的好氧及厌氧微生物。通过利用微生物来采油的技术绿色环保、成本低,可用于稠油开采,其主要机理是通过微生物在油藏中乳化原油、产气、产表面活性剂、产多糖和降解烃等五方面功能降低稠油粘度、提高水驱效率。
近20年来,对油藏环境中微生物的研究已经从最初的纯菌分离培养模式过渡到了依赖测序的分子生物学研究方式。其中,测序手段可以大致分为依赖PCR扩增的测序技术和不依赖PCR扩增的宏基因组与宏转录组测序技术。前者以16S rRNA基因克隆文库方法为代表,在油藏环境样品中已经受到广泛应用,通过设计特异性的引物可以扩增出样品中的相关基因序列,从而在基因的水平上阐述微生物的潜在代谢功能。而后者不需要PCR扩增,可以同时测定样品中所有基因的序列信息。因此,采用宏基因组测序分析油藏环境样品可以深入解析样品中潜在的代谢网络,进一步将宏基因组技术和宏转录组技术结合能够得到代谢途径上各个基因的转录水平,从而推断油藏环境下的各种微生物代谢过程。
现有的宏基因组学分析手段(例如一种基于宏转录组学和宏基因组学的环境中抗生素抗性基因的活性定量及宿主鉴定方法,申请号202110740585.0)已经可以根据需求对一些常规环境样品的目标功能基因和重要微生物进行分析。 但是地下油藏作为一个以厌氧条件为主的特殊环境,如果不对样品的采集和提取过程进行针对性的处理,样品中的微生物组成极易受到干扰而发生变化,RNA也会发生降解,从而导致后续的分析无法获得真正的油藏原位微生物数据。并且油藏中微生物的功能多种多样,其中值得关注研究的种类繁多,单一数据库无法对这些功能进行有效的分析,因此必须结合多个公开数据库以及本地自建数据库才能更全面地注释和分析样品中的这些关键功能。
可见,基于目前的技术空白,亟需开发一种特别针对油藏环境而开发的基于宏基因组学和宏转录组学的微生物识别、分析方法。
发明内容
本发明的目的就是为了解决对现有技术中油藏微生物检测手段可以获得的信息有限的问题,提供了一种联合应用宏基因组和宏转录组准确地识别油藏中微生物和代谢功能的方法。
本发明的目的通过以下技术方案实现:
本发明的目的是提供一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,包括以下步骤:
S1:自油藏产出水样提取总DNA和总RNA;
S2:对获得的总DNA和总RNA进行测序,获取油藏样品的宏基因组和宏转录组原始数据;
S3:通过对宏基因组和宏转录组结果进行分析,识别具有驱油功能的微生物。
进一步地,S1中,在提取总DNA和总RNA之前,在待提取RNA的样品中加入抑制剂以抑制厌氧微生物的RNA降解。
进一步地,S2中,还包括:对油藏样品的宏基因组和宏转录组原始数据进行预处理,得到去除接头和低质量片段的目标数据;
所述预处理的过程包括:
利用fastp软件分别对对油藏样品的宏基因组和宏转录组中各DNA或RNA链序列双端的原始数据进行滑窗质量剪裁,同时,根据序列首尾两端的引物信息,利用cutadapt软件去除引物,得到质控后的双端序列数据。
进一步地,滑窗质量剪裁的参数为-W 4,-M 20,即滑动窗大小为4,平均质量值为20。
进一步地,S3中,对宏基因组和宏转录组结果进行分析的过程包括:
对质控后的双端序列数据进行组装、分箱并评估质量,去除冗余后提取其中高质量的MAGs(宏基因组组装基因组)数据集;
根据构建的参考数据库对高质量的MAGs组装数据进行注释,识别出具有不同驱油功能的功能基因和MAGs,并计算对应的测序深度和相对丰度;
对高质量的MAGs数据集做进化关系分析,进而做出驱油微生物的群落结构分析;
将宏转录组测序得到的序列短片段质控过滤后与MAGs数据对比,计算出各个基因的转录水平。
进一步地,S3中,组装、分箱并评估质量的过程包括:
组装:使用拼接程序SPADes在Meta模式下,将质控后的双端序列数据样品短序列拼接成长度不一的contigs(交替片段产物),然后根据双端测序的信息将不同contigs连接成有测序缺口的scaffolds(骨架序列);
分箱:采用bowtie2软件将质控后的MAGs短序列数据信息比对到长序列信息上,获得不同长序列的测序覆盖度信息,进一步同时采用Maxbin2、Metabat2、CONCOCT三种Binning手段从MAGs中分离出优势菌的基因组,并进一步导入DAS_Tool程序中进行评估,最终整合并提高不同方法生成基因组的质量;
评估质量:利用dRep软件将质控后的MAGs中相似度较高的基因组去除冗余,通过CheckM工具根据基因组中单拷贝标记基因的有无和数量来估算基因组的完整度和污染度。
进一步地,S3中,对高质量的MAGs组装数据进行注释的过程包括:
使用Prodigal程序(预测开放阅读框)将拼接后的长序列翻译成编码蛋白序列(CDS),并提交到KEGG数据库,采用GhostKOALA工具进行功能性注释并获得代表不同旁系同源亚基的KO号;
同时采用本地软件KofamKOALA根据各个KO旁系同源家族蛋白的隐马可夫模型(HMM)和推荐的置信标准给各个蛋白序列注释KO号;
最后采用EggNOG emapper 2工具给蛋白序列注释COG号,再转换成KO号,使得最终每个蛋白质的KO号注释采用以下顺序:1)GhostKOALA KO,2)KofamKOALA KO,3)EggNOG emapper KO。
进一步地,S3中,所述驱油功能包括烷烃降解、产气、产乳化剂、产表面活性剂、产多糖中的一种或多种。
进一步地,S3中,识别出具有不同驱油功能的功能基因和MAGs,并结合bowtie2软件对比得到的各序列的测序覆盖度信息计算对应的测序深度和相对丰度的过程包括:
针对氢气还原酶,首先通过本地的氢气还原酶亚组的HMM模型比对找出潜在的功能基因蛋白;
之后将潜在的功能基因蛋白序列提交到HydDB数据库的在线分析软件进一步划分氢气还原酶的亚型;
针对基因组中潜在的编码次级代谢产物的功能,通过将基因组序列提交至AntiSMASH网站,结合不同的工具找到基因组中潜在编码次级代谢产物的基因组,并预测代谢产物的类型;
针对数据库中信息缺少的功能基因,单独构建本地的蛋白序列数据库,通过BlasP(Blast Protein)比对方法找到最相似的蛋白序列并进一步分析,所述数据库中信息缺少的功能基因包括厌氧烃降解初始活化基因AssA、EbdA、AhyA、AncA,和细菌微室蛋白簇基因中的一种或多种。
进一步地,S3中,对高质量的MAGs数据集做进化关系分析的过程包括:
首先将目标序列以及数据库中的相似参比序列下载并合并文件,合并后的序列首先在MAFFT上排列整齐,并以80%的阈值来选择保守位点,进而使用IQ-tree(该算法采用最大似然法构建系统发育树)来两两比对序列并生成最大似然进化树。
进一步地,S3中,将宏转录组测序得到的序列短片段质控过滤后与MAGs数据对比,计算出各个基因的转录水平的过程包括:
采用bowtie2软件将高质量的MAGs数据集中的cDNA短片段比对到通过S3中组装分箱后的宏基因组拼接得到的DNA长片段上,计算出各个基因 的转录水平,以TPM值(Transcripts Per Million)来表示。
与现有技术相比,本技术方案的优势在于:
本技术方案是一种特别针对油藏环境而开发的基于宏基因组学和宏转录组学的微生物分析方法,整体过程中,自油藏产出水样提取总DNA和总RNA,测序获取油藏样品的宏基因组和宏转录组原始数据,通过对宏基因组和宏转录组结果进行分析,识别具有驱油功能的微生物,整体过程不依赖传统的速度较慢的微生物单菌分离鉴定手段,适合处理未知物种较多的样本,且可检测极低丰度物种,检测全面、快速。
具体实施方式
下面结合具体实施例对本发明进行详细说明,但绝不是对本发明的限制。本技术方案中如未明确说明的软件/程序名称、控制方法、算法等特征,均视为现有技术中公开的常见技术特征。
本发明的油藏驱油功能微生物识别方法基于宏基因组和宏转录组测序经行检测,具体包括如下步骤:
步骤s1:收集环境样品,在样品中加入RNALater试剂抑制RNA降解,然后分别用DNA和RNA提取试剂盒提取DNA和RNA。
步骤s2:进行宏基因组和宏转录组测序。
步骤s3:宏基因组和宏转录组测序结果分析,包括以下步骤:
步骤s31:使用fastp对测序得到的核算短片段进行质检并剔除质量较低的序列和参与的测序接头。
步骤s32:使用拼接程序SPADes在‘Meta’模式下进行样品短序列拼接成长度不一的‘contigs’,然后根据双端测序的信息将不同‘contigs’连接成有测序缺口的‘scaffolds’。
步骤s33:使用Prodigal程序将拼接后的长序列翻译成编码蛋白序列(CDS),采用GhostKOALA,KofamKOALA和EggNOG emapper 2三种工具给蛋白序列注释COG号,再转换成KO号。最终每个蛋白质的KO号注释采用以下顺序:1)GhostKOALA KO,2)KofamKOALA KO以及3)EggNOG emapper KO。
步骤s34:针对氢气还原酶,首先通过本地的氢气还原酶亚组的HMM模 型比对找出潜在的功能基因蛋白。之后将这些蛋白序列提交到HydDB数据库的在线分析软件进一步划分氢气还原酶的亚型。针对基因组中潜在的编码次级代谢产物的功能,主要通过将基因组序列提交至AntiSMASH网站,结合不同的工具找到基因组中潜在编码次级代谢产物的基因组,并预测代谢产物的类型。此外,针对数据库中信息较少的功能基因,如厌氧烃降解初始活化基因AssA,EbdA,AhyA和AncA,以及细菌微室蛋白簇基因,专门构建本地的蛋白序列数据库,通过BlasP比对方法找到最相似的蛋白序列并进一步分析。
步骤s35:使用Maxbin2、Metabat2、CONCOCT三种Bining手段从数据中分离出基因组,并导入DAS_Tool程序中进行评估,合并提取通过不同方法得到的高质量基因组。
步骤s36:首先将目标序列以及数据库中的相似参比序列下载并合并文件,合并后的序列首先在MAFFT上排列整齐,并以80%的阈值来选择保守位点,进而使用IQ-tree来两两比对序列并生成最大似然进化树。
步骤s37:针对通过宏转录组测序得到的mRNA序列短片段。在质控后采用bowtie2等软件将cDNA短片段比对到宏基因组拼接得到的DNA长片段上,计算出各个基因的转录水平,以TPM值(Transcripts Per Million)来表示。
实施例1
(1)样品采集及核酸提取
在RNA样品的桶中提前加入总体积20%的RNALater,取样时保证油藏样品充满全部桶体积以排除空气。在提取核酸前低温保存。
分别使用PowerSoil Total DNA Kit(QIAGEN,美国)和PowerSoil Total RNA Kit(QIAGEN,美国)提取DNA和RNA,并使用RNase-Free DNase set(QIAGEN,美国)提纯RNA,合成cDNA后-80℃保存。
(2)获得微生物的宏基因组和宏转录组数据
使用NEBUltraTMDNA Library Prep Kit for(New England Biolabs,美国)试剂构建了带有索引码的测序文库。使用Qubit 3.0 Fluorometer(Life Technologies,美国)和Agilent 4200(Agilent,加拿大)共同评估了文库的质量。最后在Illumina Hiseq X-ten平台上,对文库进行了测序,得到了150bp 的双端碱基序列。
(3)数据分析
原始数据通过使用fastp进行质控和过滤。宏基因组数据通过SPADes经行组装,Maxbin2、Metabat2和CONCOCT进行分箱。获得的基因组导入DAS_Tool后获得高质量的基因组。使用prodigal将长片段翻译成编码蛋白序列CDS,采用GhostKOALA,KofamKOALA和EggNOG emapper 2三种工具给蛋白序列注释,通过HydDB数据库AntiSMASH分析基因组的产氢气能力和次级代谢产物合成能力,通过自建本地数据库分析厌氧烃降解能力。
宏转录组数据通过Bowtie2比对到宏基因组拼接的DNA长片段上,计算出各个基因的转录水平TPM。
结合样品中基因组的相对丰度和对应的代谢通路可以准确地识别油藏中的主要微生物和代谢功能,再结合转录组数据可以得到这些微生物的表达活性。
以大庆油田的一份产出液样品为例,表1中的数据为各基因组测序覆盖度和代谢通路。从表中可以得知该样品中最主要的微生物为变形菌门的Kerstersia_gyiorum和Serratia_nematodiphila以及厚壁菌门的Enterococcus。样品中的微生物大都为具有硝酸盐还原能力的兼性厌氧微生物,且样品中存在针对不同链长的好氧烃降解菌。近一半的微生物都具有完整的合成脂肽的代谢通路。此外,大多数氢营养型微生物同时具有产氢能力。以上这些信息可以为油藏提高采收率和防腐蚀方面的决策提供理论依据。
基因组信息表1

上述的对实施例的描述是为便于该技术领域的普通技术人员能理解和使用发明。熟悉本领域技术的人员显然可以容易地对这些实施例做出各种修改,并把在此说明的一般原理应用到其他实施例中而不必经过创造性的劳动。因此,本发明不限于上述实施例,本领域技术人员根据本发明的揭示,不脱离本发明范畴所做出的改进和修改都应该在本发明的保护范围之内。

Claims (10)

  1. 一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,包括以下步骤:
    S1:自油藏产出水样提取总DNA和总RNA;
    S2:对获得的总DNA和总RNA进行测序,获取油藏样品的宏基因组和宏转录组原始数据;
    S3:通过对宏基因组和宏转录组结果进行分析,识别具有驱油功能的微生物。
  2. 根据权利要求1所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S1中,在提取总DNA和总RNA之前,在待提取RNA的样品中加入抑制剂以抑制厌氧微生物的RNA降解。
  3. 根据权利要求1所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S2中,还包括:对油藏样品的宏基因组和宏转录组原始数据进行预处理,得到去除接头和低质量片段的目标数据;
    所述预处理的过程包括:
    利用fastp软件分别对对油藏样品的宏基因组和宏转录组中各DNA或RNA链序列双端的原始数据进行滑窗质量剪裁,同时,根据序列首尾两端的引物信息,利用cutadapt软件去除引物,得到质控后的双端序列数据。
  4. 根据权利要求3所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,对宏基因组和宏转录组结果进行分析的过程包括:
    对质控后的双端序列数据进行组装、分箱并评估质量,去除冗余后提取其中高质量的MAGs数据集;
    根据构建的参考数据库对高质量的MAGs组装数据进行注释,识别出具有不同驱油功能的功能基因和MAGs,并计算对应的测序深度和相对丰度;
    对高质量的MAGs数据集做进化关系分析,进而做出驱油微生物的群落结构分析;
    将宏转录组测序得到的序列短片段质控过滤后与MAGs数据对比,计算 出各个基因的转录水平。
  5. 根据权利要求4所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,组装、分箱并评估质量的过程包括:
    组装:使用拼接程序SPADes在Meta模式下,将质控后的双端序列数据样品短序列拼接成长度不一的contigs,然后根据双端测序的信息将不同contigs连接成有测序缺口的scaffolds;
    分箱:采用bowtie2软件将质控后的MAGs短序列数据信息比对到长序列信息上,获得不同长序列的测序覆盖度信息,进一步同时采用Maxbin2、Metabat2、CONCOCT三种Binning手段从MAGs中分离出优势菌的基因组,并进一步导入DAS_Tool程序中进行评估,最终整合并提高不同方法生成基因组的质量;
    评估质量:利用dRep软件将质控后的MAGs中相似度较高的基因组去除冗余,通过CheckM工具根据基因组中单拷贝标记基因的有无和数量来估算基因组的完整度和污染度。
  6. 根据权利要求4所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,对高质量的MAGs组装数据进行注释的过程包括:
    使用Prodigal程序将拼接后的长序列翻译成编码蛋白序列,并提交到KEGG数据库,采用GhostKOALA工具进行功能性注释并获得代表不同旁系同源亚基的KO号;
    同时采用本地软件KofamKOALA根据各个KO旁系同源家族蛋白的隐马可夫模型和推荐的置信标准给各个蛋白序列注释KO号;
    最后采用EggNOG emapper 2工具给蛋白序列注释COG号,再转换成KO号,使得最终每个蛋白质的KO号注释采用以下顺序:1)GhostKOALA KO,2)KofamKOALA KO,3)EggNOG emapper KO。
  7. 根据权利要求4所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,所述驱油功能包括烷烃降解、产气、产乳化剂、产表面活性剂、产多糖中的一种或多种。
  8. 根据权利要求7所述的一种基于宏基因组和宏转录组识别油藏驱油功 能微生物的方法,其特征在于,S3中,识别出具有不同驱油功能的功能基因和MAGs的过程包括:
    针对氢气还原酶,首先通过本地的氢气还原酶亚组的HMM模型比对找出潜在的功能基因蛋白;
    之后将潜在的功能基因蛋白序列提交到HydDB数据库的在线分析软件进一步划分氢气还原酶的亚型;
    针对基因组中潜在的编码次级代谢产物的功能,通过将基因组序列提交至AntiSMASH网站,结合不同的工具找到基因组中潜在编码次级代谢产物的基因组,并预测代谢产物的类型;
    针对数据库中信息缺少的功能基因,单独构建本地的蛋白序列数据库,通过BlastP比对方法找到最相似的蛋白序列并进一步分析,所述数据库中信息缺少的功能基因包括厌氧烃降解初始活化基因AssA、EbdA、AhyA、AncA,和细菌微室蛋白簇基因中的一种或多种。
  9. 根据权利要求4所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,对高质量的MAGs数据集做进化关系分析的过程包括:
    首先将目标序列以及数据库中的相似参比序列下载并合并文件,合并后的序列首先在MAFFT上排列整齐,并以80%的阈值来选择保守位点,进而使用IQ-tree来两两比对序列并生成最大似然进化树。
  10. 根据权利要求4所述的一种基于宏基因组和宏转录组识别油藏驱油功能微生物的方法,其特征在于,S3中,将宏转录组测序得到的序列短片段质控过滤后与MAGs数据对比,计算出各个基因的转录水平的过程包括:
    采用bowtie2软件将高质量的MAGs数据集中的cDNA短片段比对到宏基因组,并拼接得到的DNA长片段上,计算出各个基因的转录水平,以TPM值来表示。
PCT/CN2023/099026 2022-09-26 2023-06-08 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法 Ceased WO2024066461A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202211176719.1 2022-09-26
CN202211176719.1A CN115678978A (zh) 2022-09-26 2022-09-26 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法

Publications (1)

Publication Number Publication Date
WO2024066461A1 true WO2024066461A1 (zh) 2024-04-04

Family

ID=85062111

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/099026 Ceased WO2024066461A1 (zh) 2022-09-26 2023-06-08 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法

Country Status (2)

Country Link
CN (1) CN115678978A (zh)
WO (1) WO2024066461A1 (zh)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118212987A (zh) * 2024-05-21 2024-06-18 中国医学科学院北京协和医院 一种基因数据处理方法、装置、存储介质及电子设备
CN118366541A (zh) * 2024-04-29 2024-07-19 中山大学·深圳 一种基于宏转录组测序分析分节段rna病毒的方法及其应用
CN118588164A (zh) * 2024-06-07 2024-09-03 浙江工业大学 一种生物塑料降解的功能基因在线分析平台
CN119339810A (zh) * 2024-10-12 2025-01-21 宁夏红果生物科技有限公司 基于高通量测序数据的微生物成分检测方法
CN120340631A (zh) * 2025-06-20 2025-07-18 北京大学 一种脱卤微生物功能基因组数据库的构建方法及脱卤微生物的分析方法

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115678978A (zh) * 2022-09-26 2023-02-03 华东理工大学 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法
CN116758995B (zh) * 2023-08-15 2023-12-15 广州诺禾医学检验所有限公司 基因组注释方法和电子装置
CN117116341A (zh) * 2023-08-24 2023-11-24 中山大学 一种同位素示踪与宏基因组相结合的解析氮限制生态系统的方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112599198A (zh) * 2020-12-29 2021-04-02 上海派森诺生物科技股份有限公司 一种用于宏基因组测序数据的微生物物种与功能组成分析方法
CN112786102A (zh) * 2021-01-25 2021-05-11 北京大学 一种基于宏基因组学分析精准识别水体中未知微生物群落的方法
CN113257348A (zh) * 2021-05-26 2021-08-13 南开大学 一种宏转录组测序数据处理方法及系统
CN113337591A (zh) * 2021-06-30 2021-09-03 清华大学深圳国际研究生院 一种基于宏转录组学和宏基因组学的环境中抗生素抗性基因的活性定量及宿主鉴定方法
CN115678978A (zh) * 2022-09-26 2023-02-03 华东理工大学 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112599198A (zh) * 2020-12-29 2021-04-02 上海派森诺生物科技股份有限公司 一种用于宏基因组测序数据的微生物物种与功能组成分析方法
CN112786102A (zh) * 2021-01-25 2021-05-11 北京大学 一种基于宏基因组学分析精准识别水体中未知微生物群落的方法
CN113257348A (zh) * 2021-05-26 2021-08-13 南开大学 一种宏转录组测序数据处理方法及系统
CN113337591A (zh) * 2021-06-30 2021-09-03 清华大学深圳国际研究生院 一种基于宏转录组学和宏基因组学的环境中抗生素抗性基因的活性定量及宿主鉴定方法
CN115678978A (zh) * 2022-09-26 2023-02-03 华东理工大学 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
LI JIAWEI: "Application of metagenomic sequencing in reservoir microbial analysis", COMPUTERS AND APPLIED CHEMISTRY, BEIJING HUAGONG XUEYUAN, BEIJING, CN, vol. 35, no. 7, 28 July 2018 (2018-07-28), Beijing, CN , pages 565 - 572, XP093155305, ISSN: 1001-4160, DOI: 10.16866/j.com.app.chem201807005 *
YI-FAN LIU;DANIELADOMINGOS GALZERANI;SERGEMAURICE MBADINGA;LIVIAS. ZARAMELA;JI-DONG GU;BO-ZHONG MU;KARSTEN ZENGLER: "Metabolic capability and in situ activity of microorganisms in an oil reservoir", MICROBIOME, BIOMED CENTRAL LTD, LONDON, UK, vol. 6, no. 1, 5 January 2018 (2018-01-05), London, UK , pages 1 - 12, XP021252263, DOI: 10.1186/s40168-017-0392-1 *

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118366541A (zh) * 2024-04-29 2024-07-19 中山大学·深圳 一种基于宏转录组测序分析分节段rna病毒的方法及其应用
CN118212987A (zh) * 2024-05-21 2024-06-18 中国医学科学院北京协和医院 一种基因数据处理方法、装置、存储介质及电子设备
CN118588164A (zh) * 2024-06-07 2024-09-03 浙江工业大学 一种生物塑料降解的功能基因在线分析平台
CN119339810A (zh) * 2024-10-12 2025-01-21 宁夏红果生物科技有限公司 基于高通量测序数据的微生物成分检测方法
CN120340631A (zh) * 2025-06-20 2025-07-18 北京大学 一种脱卤微生物功能基因组数据库的构建方法及脱卤微生物的分析方法

Also Published As

Publication number Publication date
CN115678978A (zh) 2023-02-03

Similar Documents

Publication Publication Date Title
WO2024066461A1 (zh) 基于宏基因组和宏转录组识别油藏驱油功能微生物的方法
Gruber-Vodicka et al. phyloFlash: rapid small-subunit rRNA profiling and targeted assembly from metagenomes
Bishara et al. High-quality genome sequences of uncultured microbes by assembly of read clouds
US9593382B2 (en) Compositions and methods for identifying and comparing members of microbial communities using amplicon sequences
Almeida et al. Bioinformatics tools to assess metagenomic data for applied microbiology
Sharon et al. Accurate, multi-kb reads resolve complex populations and detect rare microorganisms
Nagy et al. First large-scale DNA barcoding assessment of reptiles in the biodiversity hotspot of Madagascar, based on newly designed COI primers
Jouffret et al. Increasing the power of interpretation for soil metaproteomics data
Der Sarkissian et al. Shotgun microbial profiling of fossil remains
JP6238069B2 (ja) 微生物の識別方法
AU2018288772B2 (en) Methods and systems for decomposition and quantification of dna mixtures from multiple contributors of known or unknown genotypes
Hebert et al. Interrogating 1000 insect genomes for NUMTs: A risk assessment for estimates of species richness
Ionescu et al. Microbial community analysis using high‐throughput amplicon sequencing
Méndez-García et al. Metagenomic protocols and strategies
Esser et al. A predicted CRISPR-mediated symbiosis between uncultivated archaea
Bickhart et al. Generation of lineage-resolved complete metagenome-assembled genomes by precision phasing
Zhang et al. Anaerobic hydrocarbon biodegradation by alkylotrophic methanogens in deep oil reservoirs
JP2025534929A (ja) 機械学習アーキテクチャを利用した複数の配列決定パイプラインからの変異コールの統合
Zhang et al. Reading the underlying information from massive metagenomic sequencing data
Neelapu et al. Next-generation sequencing and metagenomics
Eckert et al. Bioinformatics for biomonitoring: Species detection and diversity estimates across next-generation sequencing platforms
Collin et al. An open-sourced bioinformatic pipeline for the processing of Next-Generation Sequencing derived nucleotide reads: Identification and authentication of ancient metagenomic DNA
Bornemann et al. Reconstruction of Archaeal Genomes from Short-Read Metagenomes
Newell Bioinformatic methods for genome-centric metagenomics
He et al. Mitag4taxa: Extracting SSU rRNA Illumina reads from metagenomes for taxonomic classification

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23869713

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23869713

Country of ref document: EP

Kind code of ref document: A1