WO2024217010A1 - 基因与疾病关联分析方法、装置、计算机设备及存储介质 - Google Patents

基因与疾病关联分析方法、装置、计算机设备及存储介质 Download PDF

Info

Publication number
WO2024217010A1
WO2024217010A1 PCT/CN2023/137633 CN2023137633W WO2024217010A1 WO 2024217010 A1 WO2024217010 A1 WO 2024217010A1 CN 2023137633 W CN2023137633 W CN 2023137633W WO 2024217010 A1 WO2024217010 A1 WO 2024217010A1
Authority
WO
WIPO (PCT)
Prior art keywords
gene
gene expression
disease
disease phenotype
phenotype data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/137633
Other languages
English (en)
French (fr)
Inventor
王中昊
殷鹏
黄华振
朱木春
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shenzhen Institute of Advanced Technology of CAS
Original Assignee
Shenzhen Institute of Advanced Technology of CAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shenzhen Institute of Advanced Technology of CAS filed Critical Shenzhen Institute of Advanced Technology of CAS
Publication of WO2024217010A1 publication Critical patent/WO2024217010A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6876Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes
    • C12Q1/6883Nucleic acid products used in the analysis of nucleic acids, e.g. primers or probes for diseases caused by alterations of genetic material
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/30Detection of binding sites or motifs
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B5/00ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/70ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for mining of medical data, e.g. analysing previous cases of other patients
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02ATECHNOLOGIES FOR ADAPTATION TO CLIMATE CHANGE
    • Y02A90/00Technologies having an indirect contribution to adaptation to climate change
    • Y02A90/10Information and communication technologies [ICT] supporting adaptation to climate change, e.g. for weather forecasting or climate simulation

Definitions

  • the present application belongs to the technical field of gene data analysis, and in particular, relates to a gene and disease association analysis method, device, computer equipment and storage medium.
  • GWAS Gene-wide association study refers to finding sequence variations, i.e., SNPs (single nucleotide polymorphisms), in the whole human genome, and screening out disease-associated SNPs.
  • SNPs single nucleotide polymorphisms
  • GWAS shows limited statistical power.
  • eQTL Expression quantitative trait loci
  • TWAS transcriptome-wide association study
  • the analysis methods of TWAS include PrediXcan, MultiXcan and UTMOST, among which PrediXcan is a two-stage method, and its prediction effect can be attributed to the accuracy of the gene expression prediction model and the association between gene expression and phenotype.
  • PrediXcan is a two-stage method, and its prediction effect can be attributed to the accuracy of the gene expression prediction model and the association between gene expression and phenotype.
  • the current methods do not fully utilize the multi-tissue nature of transcriptome resources and the comprehensive map with regulatory elements, and ignore the existence of shared genetic regulation of tissues.
  • MultiXcan integrates information from multi-tissue studies and improves the statistical power of association analysis by regressing the principal components of predicted expression data across tissues.
  • MultiXcan is a multi-tissue association analysis method, which is not intended to improve the prediction of gene expression in each tissue, and the effect size and direction of each principal component are not easy to interpret.
  • UTMOST is a cross-tissue TWAS method that uses group penalty terms
  • the present application provides a gene-disease association analysis method, apparatus, computer device and storage medium, aiming to solve at least one of the above-mentioned technical problems in the prior art to a certain extent.
  • a gene-disease association analysis method comprising:
  • association analysis on the gene expression prediction results and the disease phenotype data Performing association analysis on the gene expression prediction results and the disease phenotype data, obtaining significant association information between the gene expression prediction results and the disease phenotype data through univariate statistical tests, and screening out important gene fragments significantly associated with the disease phenotype data according to a set threshold;
  • the technical solution adopted in the embodiment of the present application also includes: the deep learning-based GDASA gene expression prediction model is trained using the genotype and matching gene expression data in the 8th version of the GTEx database.
  • a set number of specific tissues were selected from the GTEx.v8 database according to the disease type, and the genotype and disease phenotype data of the whole transcriptome of each specific tissue were obtained.
  • pi represents the i-th P value of the expression trait effect size, and the P value is used as an evaluation index to screen out important gene fragments that are significantly associated with the disease phenotype.
  • the technical solution adopted in the embodiment of the present application also includes: the use of the deep learning interpretability method to calculate the contribution value of each gene site in the important gene fragment to the gene expression prediction of the GDASA gene expression prediction model is specifically:
  • the gene loci whose contribution value is higher than the contribution threshold are screened out as the gene loci significantly associated with the disease phenotype data.
  • a gene and disease association analysis device comprising:
  • Model building module used to build the GDASA gene expression prediction model based on deep learning
  • Gene expression prediction module used to obtain genotype and disease phenotype data, and input the genotype and disease phenotype data into the trained GDASA gene expression prediction model for gene expression prediction;
  • Association analysis module used to perform association analysis on the gene expression prediction results and disease phenotype data, and obtain the association between the gene expression prediction results and disease phenotype data through univariate statistical test. and screen out important gene fragments that are significantly associated with the disease phenotype data according to a set threshold;
  • Gene locus screening module used to use the interpretable method of deep learning to calculate the contribution value of each gene locus in the important gene fragment to the gene expression prediction of the GDASA gene expression prediction model, and screen out the gene loci whose contribution value is higher than the set contribution threshold as the gene loci significantly associated with the disease phenotype data.
  • a computer device includes a processor and a memory coupled to the processor, wherein:
  • the memory stores program instructions for implementing the gene-disease association analysis method
  • the processor is used to execute the program instructions stored in the memory to control the gene-disease association analysis method.
  • a storage medium storing program instructions executable by a processor, wherein the program instructions are used to execute the gene-disease association analysis method.
  • the beneficial effects of the embodiments of the present application are as follows: the gene-disease association analysis method, device, computer equipment and storage medium of the embodiments of the present application construct a GDASA gene expression prediction model based on deep learning, use the GDASA gene expression prediction model to predict gene expression, perform association analysis on gene expression and disease phenotype, obtain significant association information between gene expression and disease phenotype through univariate statistical test, and screen out important gene fragments significantly associated with disease phenotype according to a set threshold, use the interpretability method of deep learning to calculate the contribution value of each gene site in the important gene fragment to the model gene expression prediction, and screen out gene sites with contribution values higher than the set contribution threshold as gene sites significantly associated with the disease.
  • the GDASA gene expression prediction model of the embodiments of the present application is based on a deep learning framework, which makes the prediction accuracy of gene expression higher and greatly improves the accuracy of the model interpretable analysis.
  • FIG1 is a flow chart of a gene-disease association analysis method according to a first embodiment of the present application
  • FIG2 is a schematic diagram of a GDASA framework according to an embodiment of the present application.
  • FIG3 is a flow chart of a gene-disease association analysis method according to a second embodiment of the present application.
  • FIG4 is a TWAS analysis and a Manhattan plot of the association analysis results between gene expression and disease phenotypes according to an embodiment of the present application
  • FIG5 is a schematic diagram of the structure of a gene-disease association analysis device according to an embodiment of the present application.
  • FIG6 is a schematic diagram of the structure of a computer device according to an embodiment of the present application.
  • FIG. 7 is a schematic diagram of the structure of a storage medium according to an embodiment of the present application.
  • first”, “second”, “third” in this application are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features.
  • the features defined as “first”, “second”, “third” can expressly or implicitly include at least one of the features.
  • the meaning of “multiple” is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
  • all directional indications (such as up, down, left, right, front, back%) are only used to explain the relative position relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication also changes accordingly.
  • Figure 1 is a flow chart of the gene-disease association analysis method of the first embodiment of the present application.
  • the gene-disease association analysis method of the second embodiment of the present application includes the following steps:
  • the present application embodiment designs a set of GDASA (Gene and disease association study with AI, gene and disease correlation study based on artificial intelligence) frameworks for testing gene-disease phenotype associations based on gene expression prediction models.
  • Figure 2 it is a schematic diagram of the GDASA framework of the present application embodiment.
  • the GDASA framework constructs a GDASA gene expression prediction model based on the existing gene expression prediction model to predict gene expression based on the gene sequence of deep learning.
  • the GDASA gene expression prediction model is obtained by training the genotype and matching gene expression data of 49 tissues in the GTEx database (8th edition), predicting gene expression through gene SNP sequences, and then performing association analysis on gene expression and disease phenotypes, obtaining important gene fragments significantly associated with disease phenotypes, and performing model interpretable analysis on important gene fragments, and finally obtaining gene loci significantly associated with disease phenotype data.
  • GTEx database 8th edition
  • the GDASA gene expression prediction model includes four parts: feature control gate (FCG), linear and nonlinear modules, prior knowledge fusion module and multi-task module, among which FCG is used to filter noise and insignificant gene sites to improve the prediction performance of gene expression.
  • FCG feature control gate
  • prior knowledge fusion module fully considers the prior knowledge and integrates it
  • multi-task module uses the principal component analysis (PCA) algorithm to construct a joint organization of secondary task labels, hoping to exert a positive impact on the gene expression prediction task.
  • PCA principal component analysis
  • S110 Acquire genotype and disease phenotype data, and input the genotype and disease phenotype data into the trained GDASA gene expression prediction model to perform gene expression prediction;
  • the embodiment of the present application uses the genotype and disease phenotype data in the GTEx.v8 database to predict gene expression, selects a certain number of specific tissues in the GTEx.v8 database according to the disease type, and obtains the genotype and disease phenotype data of the whole transcriptome of each specific tissue.
  • S120 Perform association analysis on the gene expression prediction results and the disease phenotype data, obtain the significant association between the gene expression prediction results and the disease phenotype data through univariate statistical test, and screen out important gene fragments that are significantly associated with the disease phenotype data according to the set threshold;
  • the significant association information refers to various statistics such as P value, AIC value, BIC value, etc. obtained by univariate test of gene expression and disease phenotype data.
  • P value as an example, the significant association information between gene expression and disease phenotype data is calculated as follows: The correlation between gene expression and disease phenotype data in certain specific tissues is modeled using the univariate least squares regression (OLS) algorithm:
  • D is the disease phenotype data
  • ⁇ i is the intercept
  • ⁇ i is the coefficient of the OLS regression model
  • ⁇ i is the residual. Therefore, a t-statistic representing the effect size of the expression trait can be constructed:
  • t i represents the t statistic of the effect size of the ith expression trait
  • SE represents the sample standard deviation. Therefore, the P value of the significant association calculated based on the t value is: p i ⁇ P(
  • pi represents the ith P value of the expression trait effect size.
  • the P value is used as an evaluation indicator, and the P value ⁇ 0.05 is used as the threshold to screen out important gene fragments that are significantly associated with the disease phenotype.
  • S130 Use deep learning interpretability methods to calculate the contribution of each gene site in the important gene fragment to the prediction of model gene expression
  • SHAP Silicon Additive explanations, an effective model interpretability method
  • SHAP belongs to the post-training model interpretation method. Its basic idea is to calculate the average of the different marginal contributions of each feature in all feature sequence combinations as the SHAP value representing each contribution value.
  • the SHAP algorithm is specifically as follows: for each important gene fragment, the SHAP algorithm is used to calculate its SHAP Value to represent its contribution to the model gene expression prediction.
  • the SHAP algorithm will calculate the SHAP Value for each important gene fragment.
  • the SHAP Value may be greater than 0 or less than 0. When the SHAP Value is greater than 0, it indicates that the contribution of the gene site to the gene expression prediction in the corresponding sample is a positive contribution. When the SHAP Value is less than 0, it indicates that the contribution of the gene site to the gene expression prediction in the corresponding sample is a negative contribution. Whether it is a positive contribution or a negative contribution, it indicates that the gene site has a regulatory effect on the gene expression prediction.
  • the absolute value is taken to represent the contribution value of each gene site.
  • the sample dimension of the calculated SHAP Value matrix is averaged to obtain the contribution value corresponding to each gene site on the important gene fragment, and the gene sites with high contribution on the important gene fragment are extracted according to the contribution value corresponding to each gene site.
  • S140 Calculate a contribution threshold through a relevant algorithm, and select gene loci with contribution values higher than the contribution threshold as gene loci significantly associated with the disease phenotype data;
  • the contribution threshold calculation method in the embodiment of the present application is: use the bootstrap sampling method in statistics to perform sampling with replacement on all gene fragments, for a set number of specific tissues, randomly extract a certain number of gene fragments with replacement from all gene fragments of each tissue, and obtain N gene fragments, and then calculate the SHAP Value distribution of all gene loci in the N gene fragments, take the value of the 95% quantile of the confidence interval of the SHAP Value distribution as the contribution threshold, and screen out the gene loci that are significantly associated with the disease phenotype in the important gene fragments of the corresponding tissue according to the contribution threshold.
  • the final calculated contribution threshold is SHAPValue>0.05642.
  • the gene-disease association analysis method of the first embodiment of the present application constructs a GDASA gene expression prediction model based on deep learning, uses the GDASA gene expression prediction model to predict gene expression, and performs association analysis between gene expression and disease phenotype.
  • the univariate statistical test obtains the significant association information between gene expression and disease phenotype, and screens out the important gene segments significantly associated with the disease phenotype according to the set threshold, and uses the deep learning interpretability method to calculate the contribution value of each gene site in the important gene segment to the model gene expression prediction, and screens out the gene sites with contribution values higher than the set contribution threshold as the gene sites significantly associated with the disease.
  • the GDASA gene expression prediction model of the embodiment of the present application is based on a deep learning framework, which makes the prediction accuracy of gene expression higher and greatly improves the accuracy of the model interpretable analysis, which is of great significance to the target prediction of the disease and drug discovery.
  • Figure 3 is a flow chart of the gene-disease association analysis method of the second embodiment of the present application.
  • the gene-disease association analysis method is applied to the target analysis of depression disease phenotypes, and 539 important gene fragments and corresponding gene loci significantly associated with depression are detected in 13 brain tissues, providing a basis for target discovery and drug development for depression.
  • the gene-disease association analysis method of the second embodiment of the present application includes the following steps:
  • S200 obtaining the raw data of ROSMAP, filtering and quality control the raw data, and obtaining genotype and phenotype data for gene expression prediction;
  • the disease phenotype is a disease phenotype related to depression.
  • the embodiment of the present application uses the genotype and expression data of GTEx.v8 to perform expression prediction of the GDASA gene expression prediction model.
  • the whole transcriptome gene expression data of 13 brain tissues including amygdala tissue, hypothalamus tissue, hippocampus tissue, etc. are selected from the GTEx.v8 data set.
  • S220 Perform association analysis on gene expression and disease phenotype, obtain significant association information between gene expression and disease phenotype through univariate statistical test, and screen out important gene fragments significantly associated with disease phenotype according to the set threshold;
  • P O represents the P value before transformation
  • P N represents the P value after transformation
  • the number of significant genes in the substantia nigra tissue is only 12 gene fragments, and the number of iGenes detected by the GDASA expression prediction model in the substantia nigra tissue and cerebral amygdala tissue on the GETx dataset is also relatively small, with only 2575 and 2581 iGenes respectively.
  • S230 Use deep learning interpretability methods to calculate the contribution value of each gene locus in the important gene fragment, and select the gene loci with contribution values higher than the set contribution threshold as the gene loci significantly associated with the disease phenotype;
  • the present application example uses SHAP Value to perform interpretability analysis on the trained GDASA gene expression prediction model to obtain gene loci with a high contribution to the model gene expression prediction.
  • the GDASA gene expression prediction model was trained on the GTEx.v8 dataset, and the SHAP method was used to calculate its SHAP Value to indicate its contribution to the model gene expression prediction.
  • the SHAP Value may be greater than 0 or less than 0.
  • the SHAP Value indicates that the gene locus contributes positively to the gene expression prediction in the corresponding sample.
  • the SHAP Value is less than 0, it indicates that the gene locus contributes negatively to the gene expression prediction in the corresponding sample. Whether it is a positive contribution or a negative contribution, it indicates that the gene locus has a regulatory effect on the gene expression prediction. Therefore, the absolute value of all calculated SHAP Value values is taken to indicate the contribution value of the gene locus.
  • a total of 15,400 gene loci that may be significantly associated with depression were extracted from 13 brain tissues.
  • the anterior cingulate cortex tissue and the basal ganglia nucleus tissue of the accumbens had the least number of significant gene loci detected, only 845 and 803 respectively, while the number of gene loci detected in the remaining tissues exceeded 1,000.
  • FIG5 is a schematic diagram of the structure of a gene-disease association analysis device according to an embodiment of the present application.
  • the gene-disease association analysis device 40 according to an embodiment of the present application comprises:
  • Model building module 41 used to build a GDASA gene expression prediction model based on deep learning
  • Gene expression prediction module 42 used to obtain genotype and disease phenotype data, and input the genotype and disease phenotype data into the trained GDASA gene expression prediction model to perform gene expression prediction;
  • Association analysis module 43 used to perform association analysis on the gene expression prediction results and the disease phenotype data, obtain significant association information between the gene expression prediction results and the disease phenotype data through univariate statistical tests, and screen out important gene fragments significantly associated with the disease phenotype data according to a set threshold;
  • Gene locus screening module 44 used to calculate the contribution value of each gene locus in the important gene fragment to the gene expression prediction of the GDASA gene expression prediction model using the interpretable method of deep learning, and screen out the gene loci whose contribution value is higher than the set contribution threshold as the gene loci significantly associated with the disease phenotype data.
  • FIG6 is a schematic diagram of the computer device structure of an embodiment of the present application.
  • the computer device 50 includes:
  • a memory 51 storing executable program instructions
  • a processor 52 connected to the memory 51;
  • the processor 52 is used to call the executable program instructions stored in the memory 51 and perform the following steps: construct a GDASA gene expression prediction model based on deep learning; obtain genotype and disease phenotype data, and input the genotype and disease phenotype data into the trained GDASA gene expression prediction model for gene expression prediction; perform association analysis on the gene expression prediction results and the disease phenotype data, obtain significant association information between the gene expression prediction results and the disease phenotype data through univariate statistical tests, and screen out important gene fragments that are significantly associated with the disease phenotype data according to a set threshold; use a deep learning interpretability method to calculate the contribution value of each gene site in the important gene fragment to the gene expression prediction of the GDASA gene expression prediction model, and screen out the gene sites whose contribution values are higher than the set contribution threshold as the gene sites significantly associated with the disease phenotype data.
  • the processor 52 may also be referred to as a CPU (Central Processing Unit).
  • the processor 52 may be an integrated circuit chip having signal processing capabilities.
  • the processor 52 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
  • DSP digital signal processor
  • ASIC application-specific integrated circuit
  • FPGA field-programmable gate array
  • the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
  • FIG. 7 is a schematic diagram of the structure of the storage medium of the embodiment of the present application.
  • the storage medium of the embodiment of the present application stores program instructions 61 that can implement the following steps: construct a GDASA gene expression prediction model based on deep learning; obtain genotype and disease phenotype data, and input the genotype and disease phenotype data into the trained GDASA gene expression prediction model for gene expression prediction; perform association analysis on the gene expression prediction results and the disease phenotype data, obtain significant association information between the gene expression prediction results and the disease phenotype data through univariate statistical tests, and screen out important gene fragments that are significantly associated with the disease phenotype data according to a set threshold; use a deep learning interpretability method to calculate the contribution value of each gene site in the important gene fragment to the gene expression prediction of the GDASA gene expression prediction model, and screen out the gene sites whose contribution value is higher than the set contribution threshold as the gene sites that are significantly associated with the disease phenotype data.
  • the program instruction 61 can be stored in the above-mentioned storage medium in the form of a software product, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or
  • the aforementioned storage medium includes: various media that can store program instructions, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks, or terminal computer devices such as computers, servers, mobile phones, and tablets.
  • the server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
  • cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
  • each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
  • the above integrated unit can be implemented in the form of hardware or in the form of software functional units. The above is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the description and drawings of this application, or directly or indirectly used in other related technical fields, is also included in the patent protection scope of the present application.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Medical Informatics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biotechnology (AREA)
  • Chemical & Material Sciences (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Public Health (AREA)
  • Theoretical Computer Science (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Evolutionary Biology (AREA)
  • Data Mining & Analysis (AREA)
  • Molecular Biology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Genetics & Genomics (AREA)
  • Epidemiology (AREA)
  • Databases & Information Systems (AREA)
  • Analytical Chemistry (AREA)
  • Organic Chemistry (AREA)
  • Pathology (AREA)
  • Bioethics (AREA)
  • Wood Science & Technology (AREA)
  • Artificial Intelligence (AREA)
  • Physiology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Software Systems (AREA)
  • Zoology (AREA)
  • Primary Health Care (AREA)
  • Biomedical Technology (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Biochemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

提供一种基因与疾病关联分析方法、装置、计算机设备以及存储介质。所述方法包括:构建基于深度学习的GDASA基因表达预测模型;将基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;对基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与疾病表型数据显著关联的重要基因片段;使用深度学习的可解释性方法计算重要基因片段中各个基因位点的贡献程度值,并筛选出贡献程度值高于设定贡献度阈值的基因位点作为与疾病表型数据显著关联的基因位点。所述方法使基因表达预测精度更高,并提高了模型可解释分析的精确度。

Description

基因与疾病关联分析方法、装置、计算机设备及存储介质 技术领域
本申请属于基因数据分析技术领域,特别涉及一种基因与疾病关联分析方法、装置、计算机设备以及存储介质。
背景技术
GWAS(Genome-wide association study,全基因组关联研究)是指在人类全基因组范围内找出存在的序列变异,即SNP(single nucleotide polymorphism,单核苷酸多态性),从中筛选出与疾病相关的SNPs。随着高通量基因分型成本的快速下降,GWAS已被广泛用于寻找导致个体罹患复杂疾病或性状的遗传风险因素。然而,由于遗传变异的多基因效应以及较小的效应大小,GWAS显示出有限的统计能力。此外,遗传变异与这些性状之间的潜在分子机制尚不清楚。而遗传变异影响性状的一种方式是通过调节基因表达,表达定量性状基因座(eQTL)研究表明了基因表达调控的重要性,基因表达和性状之间的这种关系可以通过测量同一研究对象的遗传变异和转录组数据来研究。当转录组数据在GWAS样本中不可用时,TWAS(transcriptome-wide association,转录组广泛关联研究)为复杂的疾病机制提供了新的见解,并作为执行基于基因的关联分析的替代方法而受到欢迎。
目前,TWAS的分析方法包括PrediXcan、MultiXcan以及UTMOST等,其中PrediXcan是一种两阶段式方法,它的预测效果可归因于基因表达预测模型的准确性以及基因表达和表型之间的关联。然而,目前的方法并没有充分利用转录组资源的多组织性质和具有调控要素的综合图谱,忽略了组织共享遗传调控的存在。MultiXcan整合了多组织研究中的信息,通过回归预测跨组织的预测表达数据的主成分来提高关联分析的统计能力,但是,MultiXcan是一种多组织关联分析方法,其目的不是改善每个组织中基因表达的预测,且每个主成分的效应大小和方向不容易解释。UTMOST是一种跨组织TWAS方法,通过使用组惩罚项来做变量选择以提高预测性能。但是,它没有利用组织之间的相似性。
日前所述,现有的TWAS研究方法在训练基因表达预测模型时大多选用的是传统的线性回归模型,预测精度不高,同时在之后基因调控网络的解释中也很难做到精确。
发明内容
本申请提供了一种基因与疾病关联分析方法、装置、计算机设备以及存储介质,旨在至少在一定程度上解决现有技术中的上述技术问题之一。
为了解决上述问题,本申请提供了如下技术方案:
一种基因与疾病关联分析方法,包括:
构建基于深度学习的GDASA基因表达预测模型;
获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;
使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
本申请实施例采取的技术方案还包括:所述基于深度学习的GDASA基因表达预测模型使用GTEx数据库第8版中的基因型及匹配基因表达数据进行训练得到。
本申请实施例采取的技术方案还包括:所述GDASA基因表达预测模型包括特征控制门、线性和非线性模块、先验知识融合模块和多任务模块,所述特征控制门用于过滤噪声和不显著的基因位点,所述线性和非线性模块用于集成基因位点之间的线性关联和非线性交互,所述先验知识融合模块用于融合先验知识,所述多任务模块用于利用主成分分析算法构建联合组织二级任务标签,对基因表达预测任务施加影响。
本申请实施例采取的技术方案还包括:所述获取基因型以及疾病表型数据具体为:
根据疾病类型在GTEx.v8数据库中挑选设定数量的特定组织,并获取每个特定组织的全转录组的基因型以及疾病表型数据。
本申请实施例采取的技术方案还包括:所述对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段具体为:
使用单变量最小二乘回归算法对特定组织中基因表达和疾病表型数据的相关性进行建模为:
上述公式中,D是疾病表型数据,δi是截距,是特定组织的第i个基因表达,αi是OLS回归模型的系数,εi是残差;构造用于表示表达性状效应大小的t统计值:
上述公式中,ti表示第i个表达性状效应大小的t统计值,SE表示样本标准差,根据t值计算显著性关联的P值为:
pi≡P(|T|>|ti|)
上述公式中,pi表示表达性状效应大小的第i个P值,将P值作为评价指标筛选出与疾病表型显著关联的重要基因片段。
本申请实施例采取的技术方案还包括:所述使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值具体为:
针对每个重要基因片段,使用SHAP算法计算该重要基因片段的SHAP Value;
对所有计算出来的SHAP Value值,取其绝对值表示对应重要基因片段的贡献程度值;
针对每一个重要基因片段的SHAP Value矩阵的样本维度做平均化处理,得到该重要基因片段上每个基因位点对应的贡献程度值。
本申请实施例采取的技术方案还包括:所述筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点具体为:
针对设定数量的特定组织,使用统计学中的bootstrap抽样方法在每个特定组织的所有基因片段中随机有放回地抽取设定数量的基因片段,得到N个基因片段;
计算N个基因片段中所有基因位点的SHAP Value分布,取所述SHAP Value分布的置信区间95%分位点的值作为贡献度阈值;
筛选出所述贡献程度值高于贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
本申请实施例采取的另一技术方案为:一种基因与疾病关联分析装置,包括:
模型构建模块:用于构建基于深度学习的GDASA基因表达预测模型;
基因表达预测模块:用于获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
关联分析模块:用于对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据 之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;
基因位点筛选模块:用于使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
本申请实施例采取的又一技术方案为:一种计算机设备,所述计算机设备包括处理器、与所述处理器耦接的存储器,其中,
所述存储器存储有用于实现所述基因与疾病关联分析方法的程序指令;
所述处理器用于执行所述存储器存储的所述程序指令以控制基因与疾病关联分析方法。
本申请实施例采取的又一技术方案为:一种存储介质,存储有处理器可运行的程序指令,所述程序指令用于执行所述基因与疾病关联分析方法。
相对于现有技术,本申请实施例产生的有益效果在于:本申请实施例的基因与疾病关联分析方法、装置、计算机设备以及存储介质通过构建基于深度学习的GDASA基因表达预测模型,利用GDASA基因表达预测模型进行基因表达预测,对基因表达与疾病表型进行关联分析,通过单变量统计检验得到基因表达与疾病表型之间的显著关联信息,并根据设定阈值筛选出与疾病表型显著关联的重要基因片段,使用深度学习的可解释性方法计算出重要基因片段中各个基因位点对模型基因表达预测的贡献程度值,并筛选出贡献程度值高于设定贡献度阈值的基因位点作为与疾病显著关联的基因位点。本申请实施例的GDASA基因表达预测模型基于深度学习框架,使基因表达的预测精度更高,并大大提高了模型可解释分析的精确度。
附图说明
图1是本申请第一实施例的基因与疾病关联分析方法的流程图;
图2为本申请实施例的GDASA框架示意图;
图3是本申请第二实施例的基因与疾病关联分析方法的流程图;
图4为本申请实施例针对基因表达与疾病表型的关联分析结果进行的TWAS分析及绘制的曼哈顿图;
图5为本申请实施例的基因与疾病关联分析装置结构示意图;
图6为本申请实施例的计算机设备结构示意图;
图7为本申请实施例的存储介质的结构示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅是本申请的一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本申请中的术语“第一”、“第二”、“第三”仅用于描述目的,而不能理解为指示或暗示相对重要性或者隐含指明所指示的技术特征的数量。由此,限定有“第一”、“第二”、“第三”的特征可以明示或者隐含地包括至少一个该特征。本申请的描述中,“多个”的含义是至少两个,例如两个,三个等,除非另有明确具体的限定。本申请实施例中所有方向性指示(诸如上、下、左、右、前、后……)仅用于解释在某一特定姿态(如附图所示)下各部件之间的相对位置关系、运动情况等,如果该特定姿态发生改变时,则该方向性指示也相应地随之改变。此外,术语“包括”和“具有”以及它们任何变形,意图在于覆盖不排他的包含。例如包含了一系列步骤或单元的过程、方法、系统、产品或计算机设备没有限定于已列出的步骤或单元,而是可选地还包括没有列出的步骤或单元,或可选地还包括对于这些过程、方法、产品或计算机设备固有的其它步骤或单元。
在本文中提及“实施例”意味着,结合实施例描述的特定特征、结构或特性可以包含在本申请的至少一个实施例中。在说明书中的各个位置出现该短语并不一定均是指相同的实施例,也不是与其它实施例互斥的独立的或备选的实施例。本领域技术人员显式地和隐式地理解的是,本文所描述的实施例可以与其它实施例相结合。
请参阅图1,是本申请第一实施例的基因与疾病关联分析方法的流程图。本申请第二实施例的基因与疾病关联分析方法包括以下步骤:
S100:构建基于深度学习的GDASA基因表达预测模型;
本步骤中,本申请实施例基于基因表达预测模型设计了一套检验基因-疾病表型关联关系的GDASA(Gene and disease association study with AI,基于人工智能的基因与疾病相关性研究)框架。具体如图2所示,为本申请实施例的GDASA框架示意图。所述GDASA框架在现有基因表达预测模型的基础上构建了一个基于深度学习的基因序列预测基因表达量的GDASA基因表达预测模型,GDASA基因表达预测模型通过使用GTEx数据库(第8版)中49个组织的基因型和匹配基因表达数据进行训练得到,通过基因SNP序列预测基因表达,再对基因表达与疾病表型进行关联分析,得到与疾病表型显著关联的重要基因片段,并对重要基因片段做模型可解释分析,最终得到与疾病表型数据显著关联的基因位点。本申请实施例中,为了整合多组织信息、先验知识和基因组位点的非线性相互作用,开发了一个带注释的多任务深度学习架构 GDASA,突破了线性可加预测模型的局限性,提升基因表达的预测效果。
进一步地,GDASA基因表达预测模型包括特征控制门(FCG)、线性和非线性模块、先验知识融合模块和多任务模块四个部分,其中,FCG用于过滤噪声和不显著的基因位点,提高基因表达的预测性能。基于神经网络的因子分解机(NFM)思想,本申请实施例通过引入线性和非线性模块的组合架构,可以集成SNP位置之间的线性关联和非线性交互。先验知识融合模块充分考虑先验知识并进行融合,多任务模块利用主成分分析(PCA)算法构建联合组织二级任务标签,期望对基因表达预测任务施加积极影响。
S110:获取基因型以及疾病表型数据,将基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
本步骤中,本申请实施例使用GTEx.v8数据库中的基因型以及疾病表型数据进行基因表达预测,根据疾病类型在GTEx.v8数据库中挑选一定数量的特定组织,并获取每个特定组织的全转录组的基因型以及疾病表型数据。
S120:对基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联,并根据设定阈值筛选出与疾病表型数据显著关联的重要基因片段;
本步骤中,为了提取与疾病表型数据显著关联的重要基因片段,需要预先计算基因表达预测结果与疾病表型数据之间的显著性关联信息,显著性关联信息是指基因表达量与疾病表型数据单变量检验所得到的P值、AIC值、BIC值等各种统计量。具体的,以P值为例,基因表达与疾病表型数据之间的显著性关联信息计算方式为:使用单变量最小二乘回归(OLS)算法对某些特定组织中基因表达和疾病表型数据的相关性进行建模为:
公式(1)中,D是疾病表型数据,δi是截距,是特定组织的第i个基因表达,αi是OLS回归模型的系数,εi是残差。因此,可以构造用于表示表达性状效应大小的t统计值:
公式(2)中,ti表示第i个表达性状效应大小的t统计值,SE表示样本标准差。因此,根据t值计算显著性关联的P值为:
pi≡P(|T|>|ti|)     (3)
公式(3)中,pi表示表达性状效应大小的第i个P值,将P值作为评价指标,使用P值<0.05作为设定阈值筛选出与疾病表型显著关联的重要基因片段。
S130:使用深度学习的可解释性方法计算出重要基因片段中各个基因位点对模型基因表达预测的贡献程度值;
本步骤中,为了获得重要基因片段中的基因位点,本申请实施例使用SHAP(Shapley Additive explanations,即沙普利加和解释,一个有效的模型可解释方法)来提取基因位点对模型基因表达预测的贡献程度值。SHAP属于训练后的模型解释方法,其基本思想是计算所有特征序列组合中每个特征的不同边际贡献的平均值,作为代表每个贡献程度值的SHAP值。
进一步地,SHAP算法具体为:针对每个重要基因片段,用SHAP算法计算其SHAP Value来表示其对模型基因表达预测的贡献程度值,SHAP算法对于每个重要基因片段都会计算出其SHAP Value值,SHAP Value值有大于0或小于0的情况,SHAP Value大于0时表示该基因位点在对应样本下对基因表达预测的贡献为正面贡献,SHAP Value小于0则表示该基因位点在对应样本下对基因表达预测的贡献为负面贡献,无论是正面贡献还是负面贡献,都表示该基因位点对基因表达预测具有调控作用,因此对所有计算出来的SHAP Value值,取其绝对值来表示每个基因位点的贡献程度值,同时针对每一个重要基因片段,对计算出的SHAP Value矩阵的样本维度做平均化处理,以得到该重要基因片段上每个基因位点对应的贡献程度值,根据每个基因位点对应的贡献程度值提取出重要基因片段上具有高贡献程度的基因位点。
S140:通过相关算法计算出贡献度阈值,并筛选出贡献程度值高于贡献度阈值的基因位点作为与疾病表型数据显著关联的基因位点;
本步骤中,由于每个基因位点都有一个贡献度,为了筛选出对模型表达预测贡献度较高的基因位点,需要设定一个贡献度阈值对基因位点进行筛选。具体的,本申请实施例中的贡献度阈值计算方式为:使用统计学中的bootstrap抽样方法在所有基因片段上进行有放回抽样,针对设定数量的特定组织,在每个组织的所有基因片段中随机有放回地抽取一定数量的基因片段,得到N个基因片段,然后计算N个基因片段中所有基因位点的SHAP Value分布,取该SHAP Value分布的置信区间95%分位点的值作为贡献度阈值,根据贡献度阈值筛选出对应组织的重要基因片段中与疾病表型显著关联的基因位点。本申请实施例中,最终计算出的贡献度阈值为SHAPValue>0.05642。
基于上述,本申请第一实施例的基因与疾病关联分析方法通过构建基于深度学习的GDASA基因表达预测模型,利用GDASA基因表达预测模型进行基因表达预测,对基因表达与疾病表型进行关联分析,通过 单变量统计检验得到基因表达与疾病表型之间的显著关联信息,并根据设定阈值筛选出与疾病表型显著关联的重要基因片段,使用深度学习的可解释性方法计算出重要基因片段中各个基因位点对模型基因表达预测的贡献程度值,并筛选出贡献程度值高于设定贡献度阈值的基因位点作为与疾病显著关联的基因位点。本申请实施例的GDASA基因表达预测模型基于深度学习框架,使基因表达的预测精度更高,并大大提高了模型可解释分析的精确度,对疾病的靶点预测以及药物发现具有重要意义。
请参阅图3,是本申请第二实施例的基因与疾病关联分析方法的流程图。在本申请第二实施例中,将所述基因与疾病关联分析方法应用于抑郁症疾病表型的靶标分析中,在13个脑部组织中检测出了与抑郁症有显著关联的539个重要基因片段以及对应的基因位点,为抑郁症疾病的靶点发现与药物研发提供依据。具体的,本申请第二实施例的基因与疾病关联分析方法包括以下步骤:
S200:获取ROSMAP的原始数据,并对原始数据进行过滤和质量控制处理,得到用于进行基因表达预测的基因型以及表型数据;
本步骤中,所述疾病表型为与抑郁症相关的疾病表型,本申请实施例使用GTEx.v8的基因型与表达量数据进行GDASA基因表达预测模型的表达预测,针对抑郁症疾病在GTEx.v8数据集中挑选了包括大脑杏仁核组织、大脑下丘脑组织、大脑海马体组织等13个脑部组织的全转录组基因表达量数据。
S210:将基因型以及表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
S220:对基因表达与疾病表型进行关联分析,通过单变量统计检验得到基因表达与疾病表型之间的显著性关联信息,并根据设定阈值筛选出与疾病表型显著关联的重要基因片段;
本步骤中,针对基因表达与疾病表型的关联分析结果进行了TWAS分析并绘制曼哈顿图,具体如图4所示。显著性关联信息是指基因表达量与疾病表型单变量检验所得到的P值、AIC值、BIC值等各种统计量。本申请实施例选择P值作为评价指标,使用P值<0.05作为设定阈值筛选出与疾病表型显著关联的重要基因片段,同时对单变量统计检验得到的P值与设定阈值进行函数变换,变换公式如下:
PN=-Log(PO)     (4)
公式(4)中,PO代表变换前的P值,PN代表变换后的P值。
通过阈值筛选,最终在13个脑部组织中获得了总共539个与抑郁症疾病显著关联的重要基因片段,每个组织中与疾病显著关联的重要基因片段个数如表1所示:
表1每个脑部组织与疾病表型显著关联的重要基因片段个数
如表1所示,大部分组织中与疾病表型显著关联的重要基因片段个数都超过30个,只有大脑黑质组织、大脑脊髓颈组织、大脑杏仁核组织以及大脑下丘脑组织由于对应组织样本的数量限制等原因导致iGene个数较少,并导致能够被检测出来的基因更少,其中大脑黑质组织显著的基因个数最少仅仅只有12个基因片段,而大脑黑质组织以及大脑杏仁核组织在GETx数据集上由GDASA表达预测模型检测出的iGene个数也是比较少的只分别有2575个和2581个iGene。
S230:使用深度学习的可解释性方法计算出重要基因片段中各个基因位点的贡献程度值,并筛选出贡献程度值高于设定贡献度阈值的基因位点作为与疾病表型显著关联的基因位点;
本步骤中,在进行了基因片段与疾病表型的关联关系分析后,需要研究与疾病表型有显著关联的重要基因片段中有多少致病的基因位点。本申请实施例使用SHAP Value对训练好的GDASA基因表达预测模型进行可解释性分析,得到对模型基因表达预测贡献度较高的基因位点。具体的,针对TWAS分析中得到的每个组织中与疾病表型有显著关联的重要基因片段,在GTEx.v8数据集上做GDASA基因表达预测模型的训练,并且用SHAP方法计算其SHAP Value来表示其对模型基因表达预测的贡献程度,由于SHAP算法对于每个样本都会计算出其SHAP Value值,SHAP Value值有大于0或小于0的情况,SHAP Value大于0时表示该基因位点在对应样本下对基因表达预测的贡献为正面贡献,SHAP Value小于0则表示该基因位点在对应样本下对基因表达预测的贡献为负面贡献,无论是正面贡献还是负面贡献,都表示该基因位点对基因表达预测具有调控作用,因此对所有计算出来的SHAP Value值取绝对值来表示该基因位点的贡献值大小,同时针对每一个重要基因片段, 对计算出的SHAP Value矩阵的样本维度做平均化处理,以得到该重要基因片段上每个基因位点对应的贡献程度值,这样就可以提取出与疾病显著关联的重要基因片段上具有高贡献度的基因位点。具体的,使用贡献度阈值筛选出的13个脑部组织中所有与抑郁症疾病表型显著关联的基因位点个数如表2所示:
表2 13个脑部组织检测出的与抑郁症显著关联的基因位点个数
如表2所示,在13个脑部组织中总共提取出了15400个可能与抑郁症有显著关联的基因位点,其中大脑前扣带回皮层组织和大脑伏隔基底神经节脑核组织所检测出的显著基因位点个数最少分别只有845个和803个,其余组织所检测出的基因位点个数均超过1000个,最多的两个组织大脑脊髓颈组织和大脑脑黑质组织分别检测出了1738个和1639个显著的基因位点。
请参阅图5,为本申请实施例的基因与疾病关联分析装置结构示意图。本申请实施例的基因与疾病关联分析装置40包括:
模型构建模块41:用于构建基于深度学习的GDASA基因表达预测模型;
基因表达预测模块42:用于获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
关联分析模块43:用于对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;
基因位点筛选模块44:用于使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
请参阅图6,为本申请实施例的计算机设备结构示意图。该计算机设备50包括:
存储有可执行程序指令的存储器51;
与存储器51连接的处理器52;
处理器52用于调用存储器51中存储的可执行程序指令并执行以下步骤:构建基于深度学习的GDASA基因表达预测模型;获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
其中,处理器52还可以称为CPU(Central Processing Unit,中央处理单元)。处理器52可能是一种集成电路芯片,具有信号的处理能力。处理器52还可以是通用处理器、数字信号处理器(DSP)、专用集成电路(ASIC)、现成可编程门阵列(FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
请参阅图7,图7为本申请实施例的存储介质的结构示意图。本申请实施例的存储介质存储有能够实现以下步骤的程序指令61:构建基于深度学习的GDASA基因表达预测模型;获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。其中,该程序指令61可以以软件产品的形式存储在上述存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者 网络计算机设备等)或处理器(processor)执行本申请各个实施方式方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序指令的介质,或者是计算机、服务器、手机、平板等终端计算机设备。其中,服务器可以是独立的服务器,也可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(Content Delivery Network,CDN)、以及大数据和人工智能平台等基础云计算服务的云服务器。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统,装置和方法,可以通过其它的方式实现。例如,以上所描述的系统实施例仅仅是示意性的,例如,单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。以上仅为本申请的实施方式,并非因此限制本申请的专利范围,凡是利用本申请说明书及附图内容所作的等效结构或等效流程变换,或直接或间接运用在其他相关的技术领域,均同理包括在本申请的专利保护范围内。

Claims (10)

  1. 一种基因与疾病关联分析方法,其特征在于,包括:
    构建基于深度学习的GDASA基因表达预测模型;
    获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
    对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;
    使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
  2. 根据权利要求1所述的基因与疾病关联分析方法,其特征在于,所述基于深度学习的GDASA基因表达预测模型使用GTEx数据库第8版中的基因型及匹配基因表达数据进行训练得到。
  3. 根据权利要求2所述的基因与疾病关联分析方法,其特征在于,所述GDASA基因表达预测模型包括特征控制门、线性和非线性模块、先验知识融合模块和多任务模块,所述特征控制门用于过滤噪声和不显著的基因位点,所述线性和非线性模块用于集成基因位点之间的线性关联和非线性交互,所述先验知识融合模块用于融合先验知识,所述多任务模块用于利用主成分分析算法构建联合组织二级任务标签,对基因表达预测任务施加影响。
  4. 根据权利要求3所述的基因与疾病关联分析方法,其特征在于,所述获取基因型以及疾病表型数据具体为:
    根据疾病类型在GTEx.v8数据库中挑选设定数量的特定组织,并获取每个特定组织的全转录组的基因型以及疾病表型数据。
  5. 根据权利要求1至4任一项所述的基因与疾病关联分析方法,其特征在于,所述对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段具体为:
    使用单变量最小二乘回归算法对特定组织中基因表达和疾病表型数据的相关性进行建模为:
    上述公式中,D是疾病表型数据,δi是截距,是特定组织的第i个基因表达,αi是OLS回归模型的系数,εi是残差;构造用于表示表达性状效应大小的t统计值:
    上述公式中,ti表示第i个表达性状效应大小的t统计值,SE表示样本标准差,根据t值计算显著性关联的P值为:
    pi≡P(|T|>|ti|)
    上述公式中,pi表示表达性状效应大小的第i个P值,将P值作为评价指标筛选出与疾病表型显著关联的重要基因片段。
  6. 根据权利要求5所述的基因与疾病关联分析方法,其特征在于,所述使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值具体为:
    针对每个重要基因片段,使用SHAP算法计算该重要基因片段的SHAP Value;
    对所有计算出来的SHAP Value值,取其绝对值表示对应重要基因片段的贡献程度值;
    针对每一个重要基因片段的SHAP Value矩阵的样本维度做平均化处理,得到该重要基因片段上每个基因位点对应的贡献程度值。
  7. 根据权利要求6所述的基因与疾病关联分析方法,其特征在于,所述筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点具体为:
    针对设定数量的特定组织,使用统计学中的bootstrap抽样方法在每个特定组织的所有基因片段中随机有放回地抽取设定数量的基因片段,得到N个基因片段;
    计算N个基因片段中所有基因位点的SHAP Value分布,取所述SHAP Value分布的置信区间95%分位点的值作为贡献度阈值;
    筛选出所述贡献程度值高于贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
  8. 一种基因与疾病关联分析装置,其特征在于,包括:
    模型构建模块:用于构建基于深度学习的GDASA基因表达预测模型;
    基因表达预测模块:用于获取基因型以及疾病表型数据,将所述基因型以及疾病表型数据输入训练好的GDASA基因表达预测模型进行基因表达预测;
    关联分析模块:用于对所述基因表达预测结果与疾病表型数据进行关联分析,通过单变量统计检验得到基因表达预测结果与疾病表型数据之间的显著关联信息,并根据设定阈值筛选出与所述疾病表型数据显著关联的重要基因片段;
    基因位点筛选模块:用于使用深度学习的可解释性方法计算所述重要基因片段中各个基因位点对所述GDASA基因表达预测模型进行基因表达预测的贡献程度值,并筛选出所述贡献程度值高于设定贡献度阈值的基因位点作为与所述疾病表型数据显著关联的基因位点。
  9. 一种计算机设备,其特征在于,所述计算机设备包括处理器、与所述处理器耦接的存储器,其中,
    所述存储器存储有用于实现权利要求1-7任一项所述的基因与疾病关联分析方法的程序指令;
    所述处理器用于执行所述存储器存储的所述程序指令以控制基因与疾病关联分析方法。
  10. 一种存储介质,其特征在于,存储有处理器可运行的程序指令,所述程序指令用于执行权利要求1至7任一项所述基因与疾病关联分析方法。
PCT/CN2023/137633 2023-04-18 2023-12-08 基因与疾病关联分析方法、装置、计算机设备及存储介质 Ceased WO2024217010A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310461017.6 2023-04-18
CN202310461017.6A CN118824377A (zh) 2023-04-18 2023-04-18 基因与疾病关联分析方法、装置、计算机设备及存储介质

Publications (1)

Publication Number Publication Date
WO2024217010A1 true WO2024217010A1 (zh) 2024-10-24

Family

ID=93066471

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/137633 Ceased WO2024217010A1 (zh) 2023-04-18 2023-12-08 基因与疾病关联分析方法、装置、计算机设备及存储介质

Country Status (2)

Country Link
CN (1) CN118824377A (zh)
WO (1) WO2024217010A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119361018A (zh) * 2024-12-23 2025-01-24 微观纪元(合肥)量子科技有限公司 基于人工智能的分子虚拟筛选方法及相关设备
CN120183738A (zh) * 2025-05-23 2025-06-20 浙江医院 基于人工智能的脊柱-骨盆矢状位形态参数与心肺运动耐力参数相关性分析方法及系统
CN120526857A (zh) * 2025-07-23 2025-08-22 北京华大生物与信息融合技术研究有限公司 基因数据的处理方法、装置、存储介质及电子设备
CN120564828A (zh) * 2025-04-21 2025-08-29 中国农业大学 一种可解释性机器学习基因组预测方法及装置

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119580830B (zh) * 2024-12-03 2026-03-20 华中农业大学 一种基于特征约简和广义逆技术的自适应融合线性和非线性效应的表型预测方法及系统
CN121331235B (zh) * 2025-12-17 2026-03-03 内蒙古自治区农牧业科学院 一种使用高光谱鉴定马铃薯抗旱基因的方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210158967A1 (en) * 2019-11-26 2021-05-27 National Central University Method of prediction of potential health risk
CN113096816A (zh) * 2021-03-18 2021-07-09 西安交通大学 脑疾病发病风险预测模型建立方法、系统、设备及存储介质
CN114927163A (zh) * 2022-06-22 2022-08-19 浙江大学 一种基于单细胞图谱预测遗传模型的方法和存储介质
CN115698335A (zh) * 2020-05-22 2023-02-03 因斯特罗公司 使用机器学习模型预测疾病结果
WO2023033329A1 (ko) * 2021-09-06 2023-03-09 주식회사 바스젠바이오 질환 연관 유전자 변이 분석을 통한 질환별 위험 유전자 변이 정보 생성 장치 및 그 방법

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20210158967A1 (en) * 2019-11-26 2021-05-27 National Central University Method of prediction of potential health risk
CN115698335A (zh) * 2020-05-22 2023-02-03 因斯特罗公司 使用机器学习模型预测疾病结果
CN113096816A (zh) * 2021-03-18 2021-07-09 西安交通大学 脑疾病发病风险预测模型建立方法、系统、设备及存储介质
WO2023033329A1 (ko) * 2021-09-06 2023-03-09 주식회사 바스젠바이오 질환 연관 유전자 변이 분석을 통한 질환별 위험 유전자 변이 정보 생성 장치 및 그 방법
CN114927163A (zh) * 2022-06-22 2022-08-19 浙江大学 一种基于单细胞图谱预测遗传模型的方法和存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
MURRAY ALEX, CONGDON THOMAS R., TOMÁS RUBEN M. F., KILBRIDE PETER, GIBSON MATTHEW I.: "Red Blood Cell Cryopreservation with Minimal Post-Thaw Lysis Enabled by a Synergistic Combination of a Cryoprotecting Polyampholyte with DMSO/Trehalose", BIOMACROMOLECULES, AMERICAN CHEMICAL SOCIETY, US, vol. 23, no. 2, 14 February 2022 (2022-02-14), US , pages 467 - 477, XP093223104, ISSN: 1525-7797, DOI: 10.1021/acs.biomac.1c00599 *

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119361018A (zh) * 2024-12-23 2025-01-24 微观纪元(合肥)量子科技有限公司 基于人工智能的分子虚拟筛选方法及相关设备
CN120564828A (zh) * 2025-04-21 2025-08-29 中国农业大学 一种可解释性机器学习基因组预测方法及装置
CN120183738A (zh) * 2025-05-23 2025-06-20 浙江医院 基于人工智能的脊柱-骨盆矢状位形态参数与心肺运动耐力参数相关性分析方法及系统
CN120526857A (zh) * 2025-07-23 2025-08-22 北京华大生物与信息融合技术研究有限公司 基因数据的处理方法、装置、存储介质及电子设备

Also Published As

Publication number Publication date
CN118824377A (zh) 2024-10-22

Similar Documents

Publication Publication Date Title
WO2024217010A1 (zh) 基因与疾病关联分析方法、装置、计算机设备及存储介质
Fu et al. A gene prioritization method based on a swine multi-omics knowledgebase and a deep learning model
Wong et al. Decoding disease: from genomes to networks to phenotypes
Pasaniuc et al. Dissecting the genetics of complex traits using summary association statistics
JP4437050B2 (ja) 診断支援システム、診断支援方法および診断支援サービスの提供方法
Tiezzi et al. Accounting for trait architecture in genomic predictions of US Holstein cattle using a weighted realized relationship matrix
JP7350002B2 (ja) クロマチン相互作用データの分析用の方法および装置
Cuomo et al. CellRegMap: a statistical framework for mapping context‐specific regulatory variants using scRNA‐seq
Jia et al. Mapping quantitative trait loci for expression abundance
Nevado et al. Resequencing studies of nonmodel organisms using closely related reference genomes: optimal experimental designs and bioinformatics approaches for population genomics
Ning et al. Performance gains in genome-wide association studies for longitudinal traits via modeling time-varied effects
Huang et al. Evaluation of variant detection software for pooled next-generation sequence data
CN117457065A (zh) 一种基于单细胞多组学数据识别表型相关细胞类型的方法和系统
Monti et al. Identifying interpretable gene-biomarker associations with functionally informed kernel-based tests in 190,000 exomes
Vavoulis et al. DGEclust: differential expression analysis of clustered count data
WO2024187890A1 (zh) 基于snp数据的预测方法、装置、设备及存储介质
Cao et al. Gene selection by incorporating genetic networks into case-control association studies
WO2015105771A1 (en) Systems and methods for genomic variant analysis
JP2026507386A (ja) 定量的なバリアント病原性推定のための母集団内頻度モデル化
Gong et al. MethCP: differentially methylated region detection with change point models
US20190042697A1 (en) Computer-implemented methods for automated analysis and prioritization of variants in datasets
Che et al. Generalized linear mixed models for mapping multiple quantitative trait loci
Soller et al. Designing a QTL mapping study for implementation in the realized collaborative cross genetic reference population
CN119091963A (zh) 一种功能性染色质调控区域及靶基因鉴定方法及系统
CN109637582B (zh) 骨密度性状遗传力分析方法及装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23933875

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23933875

Country of ref document: EP

Kind code of ref document: A1