WO2023181572A1 - Covid-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム - Google Patents

Covid-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム Download PDF

Info

Publication number
WO2023181572A1
WO2023181572A1 PCT/JP2022/048200 JP2022048200W WO2023181572A1 WO 2023181572 A1 WO2023181572 A1 WO 2023181572A1 JP 2022048200 W JP2022048200 W JP 2022048200W WO 2023181572 A1 WO2023181572 A1 WO 2023181572A1
Authority
WO
WIPO (PCT)
Prior art keywords
motu
ref
data
risk
acid
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2022/048200
Other languages
English (en)
French (fr)
Inventor
尚義 永田
正裕 石金
典子 岩元
直志 竹内
博司 大野
亙 須田
裕美子 中西
弘晃 増岡
亮 青木
智彦 西嶋
博 井ノ岡
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ezaki Glico Co Ltd
National Center for Global Health and Medicine
Tokyo Medical University
RIKEN
Original Assignee
Ezaki Glico Co Ltd
National Center for Global Health and Medicine
Tokyo Medical University
RIKEN
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ezaki Glico Co Ltd, National Center for Global Health and Medicine, Tokyo Medical University, RIKEN filed Critical Ezaki Glico Co Ltd
Priority to JP2024509773A priority Critical patent/JPWO2023181572A1/ja
Publication of WO2023181572A1 publication Critical patent/WO2023181572A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12MAPPARATUS FOR ENZYMOLOGY OR MICROBIOLOGY; APPARATUS FOR CULTURING MICROORGANISMS FOR PRODUCING BIOMASS, FOR GROWING CELLS OR FOR OBTAINING FERMENTATION OR METABOLIC PRODUCTS, i.e. BIOREACTORS OR FERMENTERS
    • C12M1/00Apparatus for enzymology or microbiology
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/02Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving viable microorganisms
    • C12Q1/04Determining presence or kind of microorganism; Use of selective media for testing antibiotics or bacteriocides; Compositions containing a chemical indicator therefor
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6844Nucleic acid amplification reactions
    • C12Q1/6851Quantitative amplification
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/68Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
    • C12Q1/6869Methods for sequencing
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology

Definitions

  • the present invention relates to a method of testing at least one selected from the group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization when contracting COVID-19.
  • Non-Patent Document 1 Aggravation can be predicted by blood cytokine (CCL-17) (Non-Patent Document 2).
  • the risk of aggravation of COVID-19 can be tested using liver-type fatty acid binding protein in urine (Patent Document 1). Intestinal bacterial flora changes depending on the onset and severity of COVID-19 (Non-Patent Document 3).
  • Yamada, Gen, et al. Predicting respiratory failure for COVID-19 patients in Japan: a simple clinical score for evaluating the need for hospitalisation.”
  • Sugiyama, Masaya, et al. "Serum CCL17 level becomes a predictive marker to distinguish between mild/moderate and severe/critical disease in patients with COVID-19.”
  • the inventors of the present invention have found that conventional methods for predicting the risk of severe illness when contracting COVID-19 have a problem in that the prediction accuracy may not be sufficient.
  • One aspect of the present invention has been made in view of these circumstances, and its purpose is to reduce the risk of developing symptoms, worsening the severity, developing complications, and long-term hospitalization when infected with COVID-19. It is an object of the present invention to provide a method of testing at least one selected from a group consisting of risks with high precision based on an unprecedented new method.
  • the present invention adopts the following configuration in order to solve the above-mentioned problems.
  • the testing method evaluates at least one of the following: the risk of developing symptoms when contracting COVID-19, the risk of aggravation, the risk of developing complications, and the risk of long-term hospitalization.
  • the bacterial composition data of the intestinal flora is used to test at least one of the following: the risk of developing symptoms when contracting COVID-19, the risk of developing complications, and the risk of long-term hospitalization. It is used to test at least one selected from.
  • the inventors of the present invention have determined the risk of symptom onset and severity of symptoms when contracting COVID-19. It was confirmed that at least one selected from the group consisting of the risk of complications, the risk of developing complications, and the risk of long-term hospitalization can be predicted with higher accuracy than before.
  • the inventors of the present invention were able to determine the risk of developing symptoms, the risk of developing complications, and the long-term risk of developing symptoms when contracting COVID-19. It was confirmed that at least one selected from the group consisting of hospitalization risks could be predicted with high accuracy.
  • this configuration includes a measurement step of measuring at least one of the genetic composition data of the intestinal flora, the blood metabolite data, and the bacterial composition data of the intestinal flora in the biological sample collected from the subject.
  • a measurement step of measuring at least one of the genetic composition data of the intestinal flora, the blood metabolite data, and the bacterial composition data of the intestinal flora in the biological sample collected from the subject include.
  • at least one of the genetic composition data and blood metabolite data of the intestinal flora measured in the measuring step is the risk of symptom onset when contracting COVID-19. It is used to examine at least one selected from the group consisting of , risk of worsening, risk of developing complications, and risk of long-term hospitalization.
  • the bacterial composition data of the intestinal flora measured in the measuring step can be used to determine the risk of developing symptoms, the risk of developing complications, and It is used to test at least one selected from the group consisting of long-term hospitalization risk.
  • at least one selected from the group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization when contracting COVID-19 can be predicted with higher accuracy than before. can do.
  • one aspect of the present invention may be an inspection apparatus that implements all or part of each of the above configurations.
  • one aspect of the present invention may be an information processing method that realizes all or part of each of the above configurations, a program, or a computer or other device that stores such a program. It may be a storage medium readable by a device, machine, etc.
  • a computer-readable storage medium is a medium that stores information such as programs through electrical, magnetic, optical, mechanical, or chemical action.
  • the testing device can detect at least one of genetic composition data of intestinal flora, blood metabolite data, and bacterial composition data of intestinal flora in a biological sample collected from a subject. and inputting at least one of the acquired genetic composition data of the intestinal flora and the blood metabolite data into a pre-prepared prediction model. predict at least one selected from the group consisting of the risk of developing symptoms, the risk of aggravation, the risk of developing complications, and the risk of long-term hospitalization, or the obtained bacterial composition data of the intestinal flora. , predict at least one selected from the group consisting of the risk of developing symptoms when contracting COVID-19, the risk of developing complications, and the risk of long-term hospitalization by inputting it into a prediction model prepared in advance.
  • a prediction unit is a prediction unit.
  • the test program may store genetic composition data of intestinal flora, blood metabolite data, and bacterial composition data of intestinal flora in a biological sample collected from a subject on a computer. and inputting at least one of the acquired genetic composition data of the intestinal flora and the blood metabolite data into a prediction model prepared in advance. Bacteria of the intestinal flora obtained by predicting at least one selected from the group consisting of the risk of developing symptoms, the risk of worsening, the risk of developing complications, and the risk of long-term hospitalization when contracted with 19.
  • composition data By inputting the composition data into a prediction model prepared in advance, at least one selected from the group consisting of the risk of developing symptoms when contracting COVID-19, the risk of developing complications, and the risk of long-term hospitalization.
  • This is a prediction step for predicting , and a program for executing .
  • a method of testing with high accuracy at least one selected from the group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization when contracting COVID-19 etc. can be provided.
  • FIG. 1 is a flow diagram illustrating an overview of the inspection method according to this embodiment.
  • FIG. 2 is a bar graph showing the feature quantities extracted from the gene composition data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature quantities.
  • FIG. 3 is a bar graph showing the feature amounts extracted from the gene composition data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature amounts.
  • FIG. 4 is a bar graph showing the feature amounts extracted from blood metabolite data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature amounts.
  • FIG. 2 is a bar graph showing the feature quantities extracted from the gene composition data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature quantities.
  • FIG. 3 is a bar graph showing the feature amounts extracted from the gene composition data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature amounts.
  • FIG. 4 is a bar graph showing
  • FIG. 5 is a bar graph showing the feature amounts extracted from blood metabolite data for each outcome, and the predicted ROC-AUC score of a prediction model constructed using the feature amounts.
  • FIG. 6 is a bar graph showing the feature amounts extracted from the bacterial composition data for each outcome, and the predicted ROC-AUC score of the prediction model constructed using the feature amounts.
  • FIG. 7 is a bar graph showing the feature quantities extracted from the bacterial composition data for each outcome, and the predicted ROC-AUC score of the prediction model constructed using the feature quantities.
  • FIG. 8 is a bar graph showing the feature amounts extracted from the fecal metabolite concentration data for each outcome, and the predicted ROC-AUC score of the prediction model constructed using the feature amounts.
  • FIG. 9 is a bar graph showing the feature amounts extracted from the fecal metabolite concentration data for each outcome, and the predicted ROC-AUC score of the prediction model constructed using the feature amounts.
  • FIG. 10 is a diagram illustrating an overall outline of an analysis system including a model generation device and an inspection device according to this embodiment.
  • FIG. 11 schematically illustrates an example of the hardware configuration of the model generation device according to the embodiment.
  • FIG. 12 schematically illustrates an example of the hardware configuration of the inspection device according to the embodiment.
  • FIG. 13 schematically illustrates an example of the software configuration of the model generation device according to the embodiment.
  • FIG. 14 schematically illustrates an example of the software configuration of the inspection device according to the embodiment.
  • FIG. 15 is a flowchart illustrating an example of the processing procedure of the model generation device according to the embodiment.
  • FIG. 16 is a flowchart illustrating an example of a processing procedure of the inspection apparatus according to the embodiment.
  • this embodiment according to one aspect of the present invention will be described based on the drawings.
  • this embodiment described below is merely an illustration of the present invention in all respects. It goes without saying that various improvements and modifications can be made without departing from the scope of the invention. That is, in implementing the present invention, specific configurations depending on the embodiments may be adopted as appropriate.
  • FIG. 1 is a flow diagram illustrating an overview of the inspection method according to the present embodiment.
  • the present inventors have determined that COVID-19 can be detected by using at least one of genetic composition data of intestinal flora, blood metabolite data, and bacterial composition data of intestinal flora in biological samples collected from subjects. It has been confirmed that the risk of becoming seriously ill when infected with -19 can be predicted with higher accuracy than before.
  • the Boruta method as a machine learning algorithm to extract features from the gene composition data of the intestinal flora
  • several genes were extracted as feature genes, and a prediction model using random forest was created using these genes. It was confirmed that the risk of severe disease can be predicted with high accuracy.
  • the testing method includes gene composition data of intestinal flora (target gene composition data 221A), target blood metabolite data, and target blood metabolite data in a biological sample collected from a subject. 221B, and a measuring step (S11) of measuring at least one of bacterial composition data of intestinal flora (target bacterial composition data 221C).
  • target gene composition data 221A the target gene composition data 221A
  • target blood metabolite data 221B, and the target bacterial composition data 221C measured in the measurement step is infected with COVID-19.
  • the target gene composition data 221A and the target blood metabolite data 221B a group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization when contracting COVID-19 is determined.
  • using the target bacterial composition data 221C at least one selected from the group consisting of the risk of developing symptoms when contracting COVID-19, the risk of developing complications, and the risk of long-term hospitalization is tested.
  • measurements can be made from feces that can be collected by non-medical personnel using a non-invasive method.
  • the risk of aggravation, etc. can be examined.
  • the subject can be tested for the risk of aggravation, etc. using the subject gene composition data 221A or the subject bacterial composition data 221C in feces collected by the subject himself/herself.
  • target gene composition data 221A, target blood metabolite data 221B, and target bacterial composition data 221C in a biological sample collected from a subject can be read. At least one of these can be used to test for the risk of serious illness.
  • the background information includes, for example, height, weight, age, gender, presence or absence of a past illness, presence or absence of drinking habits, presence or absence of smoking habits, and the like. Details of COVID-19 non-infected individuals and COVID-19 infected individuals are explained below.
  • COVID-19 infected persons were defined as those who met all of the following conditions. That is, first, the condition that written consent for secondary use of specimens has been obtained regarding research participation. Second, the person must be a male or female who is 16 years old or older and 120 years old or younger at the time of obtaining the above consent.
  • the condition is that the person is Japanese (here refers to a person who was born in Japan and has Japanese nationality).Fourth, the condition is that the person is diagnosed with COVID-19 infection, and the National Research and Development Agency National Institute for Global Health Research The condition is that the patient was hospitalized and treated at the Center (Japan). Those who meet the fourth condition are particularly those who were diagnosed with COVID-19 during the period from January 31, 2020 to September 30, 2020. , who was hospitalized and treated at the National Center for Global Health and Medicine, and submitted fecal and blood samples. -003472-03) Principal Investigator: Atsuhito Sugiyama” handles clinical information, feces, and blood samples from those who have given consent for sample analysis between January 31, 2020 and September 30, 2020. Ta.
  • the ⁇ (alpha) strain which is a mutant strain of SARS coronavirus 2 (SARS-CoV-2), which is known as the virus that causes the new coronavirus infection (COVID-19), was introduced in the UK in November 2020. It was detected in a sample taken in September, when the virus was spreading. Therefore, it is believed that those diagnosed with COVID-19 in Japan during the period from January 31, 2020 to September 30, 2020 were not infected with specific mutant strains such as the alpha strain.
  • COVID-19 non-infected persons was defined as a person who met all of the first, second, and third conditions similar to those of a COVID-19 infected person, and also met the following condition as the fourth condition. In other words, those who meet the condition of having submitted fecal and plasma/serum samples between September 1, 2014, and June 31, 2019, clearly before the spread of COVID-19 infection, will be eligible for the fourth test. Those who meet the following conditions were selected.
  • Omics information (1) Microbial community data (intestinal bacterial composition data (intestinal flora bacterial composition data) and bacterial genetic composition data) from fecal metagenome analysis. Note that the identification method will be described later. (2) Metabolite community data of feces and plasma/serum. Note that the identification method will be described later.
  • fecal pellets were suspended in TE20 buffer (10 mmol/l Tris-HCl (pH 8.0) and 20 mmol/l EDTA (pH 8.0)), and lysozyme (Wako, 10mg per sample), akmopeptidase (Wako, 1,333 units per sample), sodium dodecyl sulfate (1%), and proteinase K (Merck, 0.67 mg per sample) were added for lysis. DNA was then recovered using phenol/chloroform/isoamyl alcohol (Wako) and 2-propanol (Wako). Furthermore, the recovered DNA was purified by treatment with RNase (Nippon Gene, 20 ⁇ g/ml) and used as fecal DNA.
  • RNase Nippon Gene, 20 ⁇ g/ml
  • Reads with low quality values (meanQV ⁇ 20) and duplicate reads were removed from the obtained metagenomic data. Furthermore, we trimmed bases with low accuracy where QV ⁇ 20 at the 5' end of each read, removed reads where more than half of the reads had QV ⁇ 20, and removed reads with a read length of less than 50 bp. Furthermore, metagenomic reads were mapped to the human genome (HG38) and the phiX bacteriophage genome using minimap2 (version 2.13-r850), and human genome sequences and PhiX sequences were removed by excluding reads with hits with 95% or more homology.
  • the created protein catalog was annotated in the KEGG database (as of October 7, 2019) using diamond (version 2.0.11).
  • the gene catalog was mapped to short reads using Bowtie2 (version 2.3.2). Functional composition data was calculated by combining the mapped data and the annotation results to KEGG.
  • a 10 ⁇ l aliquot of plasma was extracted using a mixture containing methanol:water:chloroform in a volume ratio of 5:5:2, and a 5 kDa cutoff filter column (Ultrafree MC-
  • the extract obtained by filtration with PLHCC Human Metabolome Technologies
  • a 25- ⁇ l aliquot of fecal methanol extract was extracted using a mixture containing methanol:water:chloroform at a volume ratio of 5:5:2, and the resulting extract was unfiltered. used for.
  • the extract prepared in this manner was divided into two for GC/MS/MS analysis and one for LC-MS/MS analysis.
  • GC/MS/MS analysis 200 ⁇ l of the extract was dried using an evaporator and freeze-dried using a freeze dryer. Lyophilized extracts were derivatized with methoxyamine hydrochloride (Sigma-Aldrich) and N-methyl-N-trimethylsilyl-trifluoroacetamide (MSTFA, GL Science) before analysis. The GC/MS/MS analysis program described in Reference 1 was used. For LC-MS/MS analysis, 100 ⁇ l of the extract was dried using an evaporator and freeze-dried using a freeze dryer. Prior to analysis, the lyophilized extract was reconstituted with 500 ⁇ l of water.
  • LC-MS/MS analysis was performed using a Shimadzu LCMS-8040 triple quadrupole mass spectrometer (Shimadzu).
  • a Discovery HS F5-3 2.1 mm.D. Solvent A (water with 0.1% formic acid) and Solvent B (acetonitrile with 0.1% formic acid) were used for analysis with the following gradient: 0.0 min (0% B); 0.0-3.5 min (0% B ⁇ 25%B); 3.5-7.5 minutes (25%B ⁇ 35%B); 7.5-10.3 minutes (35%B ⁇ 95%B); 10.3-13.7 minutes (95%B); 13.7-14.8 minutes (100% B); 14.8-17 minutes (0%B).
  • the injection volume was set to 1 ⁇ L.
  • Hydrophilic metabolites were monitored using multiple reaction monitoring (MRM). The parameters were set as follows: atomizing gas flow rate 3 l/min, heating gas flow rate 10 l/min, interface temperature 300 °C, DL temperature 250 °C, heating block temperature 400 °C, drying gas flow rate 15 l. /min.
  • MRM multiple reaction monitoring
  • the GC/MS/MS data and LC-MS/MS data were then processed using LabSolutions Insight (Shimadzu) to calculate the concentration.
  • SCFA single-chain fatty acids
  • Fecal samples were prepared by adding a 25 ⁇ l aliquot of fecal methanol extract to 10 ⁇ l of ultrapure water containing the internal standard, centrifuging at 40° C., and reconstituting by adding 100 ⁇ l of ultrapure water. 50 ⁇ l of hydrochloric acid and 200 ⁇ l of diethyl ether were added to each sample and stirred well. Then, after phase separation by centrifugation (3,000 ⁇ g, 10 min), 80 ⁇ l of the organic phase was transferred to a glass vial, and further added with N-tert-butyldimethylsilyl-N-trifluoroacetamide (MTBSTFA, Sigma-Aldrich). was added to derivatize.
  • MTBSTFA N-tert-butyldimethylsilyl-N-trifluoroacetamide
  • the training omics data 31 includes genetic composition data of intestinal flora (training gene composition data 31A), blood metabolite data (training blood metabolite data 31B), and bacterial composition data of intestinal flora (training bacterial composition data). 31C).
  • the training omics data 31 may further include fecal metabolite concentration data (training fecal metabolite concentration data 31D).
  • the training gene composition data of the intestinal flora may be abbreviated as "gene composition data” and the bacterial composition data of the intestinal flora may be abbreviated as "bacterial composition data.”
  • outcomes include "presence or absence of symptoms when contracted with COVID-19", “presence of worsening of symptoms”, “presence of complications”, and “presence of long-term hospitalization”.
  • the "symptoms” when contracting COVID-19 are at least one of respiratory symptoms, pneumonia, and diarrhea.
  • severe worsening (of corona infection) when contracting COVID-19 refers to the above-mentioned “severe worsening” of the respiratory tract.
  • a “complication” when contracting COVID-19 is at least one of liver damage, kidney damage, and thrombosis.
  • Thrombosis is defined as at least one of the following: a decrease in platelets to less than 20 x 10 4 / ⁇ l, an increase in D-dimer to 1.0 mg/ml or more, and an increase in fibrinogen to more than 400 mg/dl. be.
  • "Long-term hospitalization” indicates that the hospitalization period is 11 days or more. Hospitalization period refers to the period from ⁇ hospitalization'' to ⁇ discharge''. ⁇ Hospitalization'' means that the condition may worsen in the future based on the current state of the illness, or that the condition may be monitored at home. This refers to a condition in which managed treatment is provided at a medical institution under a 24-hour management system when it is determined that it is difficult to do so.
  • explanatory variables important for determining each outcome such as "presence or absence of severe disease” were selected, that is, feature quantity selection was performed.
  • explanatory variables (features) important for discrimination were extracted from the training omics data 31 using the Boruta algorithm (version 0.3). That is, when the outcome including the above-mentioned "presence or absence of severe disease” is used as the objective variable, the features (explanatory variables) that are important are the training gene composition data 31A, the training blood metabolite data 31B, the training bacterial composition data 31C, and extracted from each of the training fecal metabolite concentration data 31D.
  • a predictive model M was constructed by performing machine learning using the extracted feature quantities as explanatory variables and outcomes such as "presence or absence of severe disease” as objective variables.
  • a prediction model M was constructed using a random forest (scikit-learn version 0.24.1). That is, the prediction models MA, MB, MC, and , MD was constructed.
  • the outcomes include "presence or absence of symptoms when contracted with COVID-19", “presence of worsening of symptoms”, “presence of complications”, and “presence of long-term hospitalization”.
  • symptoms includes “respiratory symptoms,””pneumonia,” and “diarrhea.”
  • “complications” include “liver damage,” “kidney damage,” and “thrombosis,” and “thrombosis” includes “a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l,” and “D- This includes “an increase in dimer to 1.0 mg/ml or more” and “an increase in fibrinogen to more than 400 mg/dl.” Therefore, each of the prediction models MA, MB, MC, and MD includes a prediction model that predicts the presence or absence of each outcome described above.
  • the prediction model MA constructed using the features extracted from the training gene composition data 31A includes a prediction model MA1 that predicts "presence or absence of severe illness” and a prediction model MA1 that predicts "presence or absence of respiratory symptoms”.
  • MA2 includes a prediction model MA3 that predicts "the presence or absence of pneumonia.”
  • the predictive models MA include a predictive model MA4 that predicts "the occurrence of diarrhea”, a predictive model MA5 that predicts "the presence or absence of the onset of liver damage", a predictive model MA6 that predicts "the presence or absence of the onset of renal damage", and " It includes a prediction model MA7 that predicts the presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l.
  • Predictive model MA is predictive model MA8, which predicts "the presence or absence of an increase in D-dimer to 1.0 mg/ml or more," predictive model MA9, which predicts "the presence or absence of fibrinogen greater than 400 mg/dl,” and predictive model MA9, which predicts "the presence or absence of fibrinogen level greater than 400 mg/dl.”
  • the prediction model MA11 includes a prediction model MA11 that predicts "the presence or absence of”. The same applies to each of the prediction models MB, MC, and MD.
  • the prediction model M (MA, MB, MC) constructed using each of the training gene composition data 31A, training blood metabolite data 31B, and training bacterial composition data 31C was used to predict COVID-19. It has been confirmed that the risk of serious illness in the event of infection can be predicted with higher accuracy than before.
  • the ROC-AUC score for predicting "severe disease” using "patient background (weight, sex, age)" described in Non-Patent Document 1 is “0.71".
  • the ROC-AUC score for predicting "severe disease” of the prediction model MD ("fecal metabolites” in Table 1) constructed using the training fecal metabolite concentration data 31D is "0.72".
  • the ROC-AUC score of prediction model MA intestinal bacteria gene composition
  • training gene composition data 31A for predicting "severe disease” is "0.87". be.
  • the ROC-AUC score for predicting "severe disease” of the prediction model MB blood metabolites” in Table 1 constructed using the training blood metabolite data 31B is "0.92". That is, the prediction model M (at least one of the prediction models MA and MB) constructed using at least one of the trained gene composition data 31A and the trained blood metabolite data 31B predicts "severe disease” with higher accuracy than before. It was confirmed that it can be done.
  • prediction model MA constructed using the training gene composition data 31A and the prediction model MB constructed using the training blood metabolite data 31B both have "presence or absence of complications” and “presence or absence of long-term hospitalization”. It has also been confirmed that predictions can be made with higher accuracy than before.
  • the ROC of prediction model MC (intestinal bacterial species composition” in Table 1) constructed using training bacterial composition data 31C for "liver damage", “kidney damage”, etc.
  • the -AUC score is higher than the ROC-AUC score for predictions such as "liver disorder” and "kidney disorder” of the prediction model MD constructed using the training fecal metabolite concentration data 31D. Therefore, it has been confirmed that the predictive model MC constructed using training bacterial composition data 31C can predict with high accuracy the risk of developing symptoms, developing complications, and risk of long-term hospitalization when contracting COVID-19. Ta.
  • FIGS. 2 to 9 The feature amounts used in constructing each of the prediction models MA, MB, MC, and MD and the ROC-AUC scores of each prediction of the prediction models MA, MB, MC, and MD are shown in FIGS. 2 to 9. Shown below. Note that in the bar graphs shown in each of FIGS. 2 to 9, the length of the bar indicates the degree of contribution of each feature to the prediction of each outcome. In addition, features with dark gray bars indicate positive contributions to the prediction of each outcome, and feature quantities with light gray bars indicate negative contributions to the prediction of each outcome. It shows something.
  • Figures 2 and 3 show the feature quantities extracted from the training gene composition data 31A for each outcome and the predicted ROC-AUC scores of the prediction model MA constructed using the extracted feature quantities. ing.
  • ROC-AUC score of prediction model MA1 which was constructed using the genes (features) shown in (1) to (20) above, for predicting "presence or absence of severe disease” is "0.87". Met.
  • ROC-AUC score of prediction model MA2 which was constructed using the genes (features) shown in (21) to (40) above, for predicting “the presence or absence of respiratory symptoms” is “0. .78”.
  • ROC-AUC score of prediction model MA5 which was constructed using the genes (features) shown in (81) to (100) above, for predicting "presence or absence of onset of liver damage” is "0. 84".
  • ROC-AUC score of prediction model MA6 which was constructed using the genes (features) shown in (101) to (120) above, for predicting "presence or absence of onset of renal disorder” is "0. 88".
  • the prediction model MA8 which was constructed using the genes (features) shown in (141) to (160) above, relates to the prediction of "the presence or absence of an increase in D-dimer to 1.0 mg/ml or more".
  • the ROC-AUC score was "0.85".
  • ROC-AUC score related to prediction of "presence or absence of fibrinogen greater than 400 mg/dl" of prediction model MA9 constructed using the genes (features) shown in (161) to (180) above is: It was "0.89".
  • the genes shown in (201) to (220) below were extracted as feature quantities from the training gene composition data 31A. That is, (201) K01463 (bshB1), (202) K14188 (dltC), (203) K07457 (K07457), (204) K22958 (lpdC), (205) K15372 (toa), (206) K0 7069 (K07069), (207) K08153 (blt), (208) K14654 (RIB7, arfC), (209) K06156 (gntU), (210) K03734 (apbE), (211) K07505 (repA), (212) K0148 2 (DDAH, ddaH ), (213) K03328 (TC.PST), (214) K01635 (lacD), (215) K07217 (K07217), (216) K19973 (mntA), (217) K03753 (mobB),
  • ROC-AUC score of prediction model MA11 which was constructed using the genes (features) shown in (201) to (220) above, for predicting "presence or absence of long-term hospitalization” is "0.85". Met.
  • Figures 4 and 5 show the feature quantities extracted from the training blood metabolite data 31B for each outcome, and the predicted ROC-AUC score etc. of the prediction model MB constructed using the extracted feature quantities. It shows.
  • the blood metabolites shown below (301) to (320) are extracted as feature quantities from the training blood metabolite data 31B. It was done. That is, (301) phenylalanine, (302) mannose, (303) 2-hydroxybutyric acid, (304) inositol, (305) 4-hydroxyphenyllactic acid, (306) 3-phenyllactic acid, (307) tyrosine, (308) Glycine, (309) oxalic acid, (310) 2-aminoethanol, (311) threonic acid, (312) glycerol, (313) urea, (314) 2-hydroxyisovaleric acid, (315) L-carnitine, ( 316) Valine, (317) Glucuronic acid, (318) Dimethylglycine, (319) Sarcosine, and (320) Glucose were extracted as features.
  • ROC-AUC score of prediction model MB1 which was constructed using the blood metabolites (feature values) shown in (301) to (320) above, for predicting "presence or absence of severe disease” is "0. 92".
  • the blood metabolites shown in (321) to (335) below were extracted as feature quantities from the training blood metabolite data 31B. That is, (321) 3-hydroxyisobutyric acid, (322) propionic acid, (323) 2-hydroxyisobutyric acid, (324) inositol, (325) 2-ketoglutaric acid, (326) glyoxylic acid, (327) leucine, (328) 5-oxoproline, (329) phenylalanine, (330) valine, (331) glutamic acid, (332) 3-aminoisobutyric acid, (333) 2-aminoisobutyric acid, (334) nonanoic acid, (335) glucuronic acid was extracted as a feature.
  • ROC-AUC score related to prediction of "presence or absence of respiratory symptoms" of prediction model MB2 constructed using the blood metabolites (feature values) shown in (321) to (335) above is: It was "0.77".
  • the blood metabolites shown in (336) to (355) below were extracted as feature quantities from the training blood metabolite data 31B. That is, (336) erythrulose, (337) inositol, (338) mannitol, (339) formic acid, (340) serine, (341) glycine, (342) mannose, (343) fructose, (344) histidine, (345) 2-ketoglutaric acid, (346) indole-3-acetic acid, (347) trans-aconitic acid, (348) tyrosine, (349) ribulose, (350) glucose, (351) 2-aminoadipic acid, (352) ) Galacturonic acid, (353) uracil, (354) 3-hydroxyglutaric acid, and (355) serotonin were extracted as features.
  • ROC-AUC score of prediction model MB3 which was constructed using the blood metabolites (features) shown in (336) to (355) above, for predicting “presence or absence of pneumonia” is “0. .86”.
  • the blood metabolites shown in (356) to (361) below were extracted as feature quantities from the training blood metabolite data 31B. That is, (356) proline, (357) asparagine, (358) creatine, (359) indole-3-acetic acid, (360) phenylalanine, and (361) glyceric acid were extracted as feature quantities.
  • the blood metabolites shown below (362) to (380) were extracted as feature quantities from the training blood metabolite data 31B. That is, (362) 2-ketoglutaric acid, (363) glutamic acid, (364) histidine, (365) valine, (366) asparagine, (367) erythrulose, (368) citric acid, (369) isobutyric acid, (370) Glutamine, (371) Mannitol, (371) 2-aminopimelic acid, (372) Lysine, (373) Indole-3-acetic acid, (374) Succinic acid, (375) Leucine, (376) 2-ketoisocaproic acid, (377) ) Glycine, (378) 2-hydroxyisovaleric acid, (379) xylulose, and (380) lactic acid were extracted as feature quantities.
  • ROC-AUC score related to prediction of "presence or absence of onset of liver damage" of prediction model MB5 which was constructed using the blood metabolites (feature values) shown in (362) to (380) above, is " 0.79''.
  • the blood metabolites shown in (381) to (400) below were extracted as feature quantities from the training blood metabolite data 31B. Namely, (381) inositol, (382) urea, (383) creatinine, (384) arabinose, (385) 4-hydroxyphenyllactic acid, (386) cystine, (387) glutamic acid, (388) erythritol, (389) tryptophan.
  • ROC-AUC score related to prediction of "presence or absence of onset of renal disorder” of prediction model MB6 constructed using the blood metabolites (feature values) shown in (381) to (400) above is " 0.83''.
  • the objective variable is “presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l”
  • the blood metabolites shown below (401) to (417) are extracted as feature quantities from the training blood metabolite data 31B. It was done.
  • the predictive model MB7 which was constructed using the blood metabolites (features) shown in (401) to (417) above, is used to predict “the presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l”.
  • the ROC-AUC score was "0.81".
  • the objective variable is "the presence or absence of an increase in D-dimer to 1.0 mg/ml or more”
  • the blood metabolites shown in (418) to (433) below are selected as feature quantities from the training blood metabolite data 31B. Extracted.
  • the prediction model MB8 which was constructed using the blood metabolites (features) shown in (418) to (433) above, predicts "the presence or absence of an increase in D-dimer to 1.0 mg/ml or more".
  • the ROC-AUC score was "0.76".
  • the blood metabolites shown below (461) to (480) were extracted as feature quantities from the training blood metabolite data 31B. Namely, (461) erythrulose, (462) succinic acid, (463) mannose, (464) methionine sulfoxide, (465) maltose, (466) mannitol, (467) indole-3-acetic acid, (468) oxalic acid, ( 469) Asparagine, (470) Cysteine, (471) L-carnitine, (472) Urea, (473) Propionic acid, (474) Glutamine, (475) Inositol, (476) 2-ketoglutaric acid, (477) Lactic acid, (478) Glucose, (479) Glycerol, and (480) Decanoic acid were extracted as feature quantities.
  • ROC-AUC score of prediction model MB11 which was constructed using the blood metabolites (feature values) shown in (461) to (480) above, for predicting "presence or absence of long-term hospitalization” is "0. It was 8.
  • Figures 6 and 7 show the features extracted from the training bacterial composition data 31C for each outcome, and the predicted ROC-AUC scores of the prediction model MC constructed using the extracted features. ing.
  • the intestinal bacteria shown below (501) to (520) are extracted as feature quantities from the training bacterial composition data 31C. Ta. That is, (501) Synergistes sp (ref_mOTU_v25_11256), (502) Blautia producta (ref_mOTU_v25_02151), (503) Streptococcus oralis/pseudopneumoniae (ref_mOTU_v25_00296), (504) Streptococcus anginosus (ref_mOTU_v25_00569), (505) Blautia obeum/wexlerae (ref_mOTU_v25_02154) , (506) Flavonifactor plautii (ref_mOTU_v25_05238), (507) Ruthenibacterium lactatiformans (ref_mOTU_v25_0471
  • the ROC-AUC score of the prediction model MC1 which was constructed using the intestinal bacteria (feature values) shown in (501) to (520) above, for predicting "presence or absence of severe disease” is "0. 73”.
  • ROC-AUC score related to prediction of "presence or absence of respiratory symptoms" of prediction model MC2 constructed using the intestinal bacteria (features) shown in (521) to (540) above is: It was "0.72".
  • the intestinal bacteria shown in (541) to (560) below were extracted as features from the training bacterial composition data 31C. That is, (541) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367), (542) Firmicutes species incertae sedis (meta_mOTU_v25_12923), (543) Eggerthella lenta (ref_mOTU_v25_00719), (54 4) Streptococcus cristatus (ref_mOTU_v25_03967), (545) Bacteroides caccae (ref_mOTU_v25_03473) , (546) Bacteroides species incertae sedis (ext_mOTU_v26_17504), (547) Rothia mucilaginosa (ref_mOTU_v25_05265), (548) Streptococcus species incer
  • ROC-AUC score of prediction model MC3 which was constructed using the intestinal bacteria (features) shown in (541) to (560) above, for predicting "presence or absence of pneumonia” is "0. .57”.
  • ROC-AUC score of prediction model MC4 which was constructed using the intestinal bacteria (features) shown in (561) to (580) above, for predicting "the presence or absence of diarrhea” is "0. .73”.
  • the intestinal bacteria shown in (581) to (600) below were extracted as features from the training bacterial composition data 31C. That is, (581) Eggerthella lenta (ref_mOTU_v25_00719), (582) Eubacterium sp (meta_mOTU_v25_12688), (583) Clostridium sp (meta_mOTU_v25_12609), (584) Clostridiales spec ies incertae sedis (ext_mOTU_v26_26595), (585) Bacteroides sp.
  • ROC-AUC score of prediction model MC5 which was constructed using the intestinal bacteria (feature values) shown in (581) to (600) above, for predicting “presence or absence of onset of liver damage” is “ 0.74''.
  • the intestinal bacteria shown below (601) to (620) were extracted as features from the training bacterial composition data 31C. That is, (601) Clostridiales species incertae sedis (meta_mOTU_v25_13006), (602) Streptococcus anginosus (ref_mOTU_v25_00569), (603) Streptococcus intermedius/constellatus (ref_mOTU_v25_00572) , (604) Actinomyces marseillensis/pacaensis (ref_mOTU_v25_03846), (605) Firmicutes bacterium CAG :114 (ref_mOTU_v25_07728), (606) Methanobrevibacter smithii (ref_mOTU_v25_03695), (607) Clostridiales sp.
  • ROC-AUC score of prediction model MC6 which was constructed using the intestinal bacteria (feature values) shown in (601) to (620) above, for predicting "presence or absence of onset of renal disorder” is " 0.72''.
  • the intestinal bacteria shown in (621) to (640) below are determined from the training bacterial composition data 31C. , were extracted as features.
  • the prediction model MC7 which was constructed using the intestinal bacteria (feature values) shown in (621) to (640) above, was used to predict “the presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l”.
  • the ROC-AUC score was "0.71".
  • the intestinal bacteria shown in (641) to (660) below are determined from the training bacterial composition data 31C. was extracted as a feature.
  • the prediction model MC8 which was constructed using the intestinal bacteria (features) shown in (641) to (660) above, predicts “the presence or absence of an increase in D-dimer to 1.0 mg/ml or more”.
  • the ROC-AUC score was "0.73".
  • the intestinal bacteria shown below (661) to (680) are extracted as features from the training bacterial composition data 31C. It was done. Namely, (661) Faecalibacterium prausnitzii (ref_mOTU_v25_06108), (662) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716), (663) Collinsella species incertae sedis (ext_mOTU_v26_17347), (664) Blautia obeum/wexlerae (ref_mOTU_v25_02154), (665) Clostridiales sp.
  • the intestinal bacteria shown in (701) to (720) below were extracted as feature quantities from the training bacterial composition data 31C. That is, (701) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716), (702) Blautia masssiliensis (ref_mOTU_v25_03342), (703) Streptococcus anginosus/intermedius (ref_mOTU_v25_00567), (704) Massilioclostridium coli (ref_mOTU_v25_10237), (705) Eggerthella lenta (ref_mOTU_v25_00719), ( 706) Clostridiales Family XIII (ext_mOTU_v26_16402), (707) Clostridiales sp.
  • ROC-AUC score of prediction model MC11 which was constructed using the intestinal bacteria (feature values) shown in (701) to (720) above, for predicting "presence or absence of long-term hospitalization” is "0. 74”.
  • Figures 8 and 9 show the feature quantities extracted from the training fecal metabolite concentration data 31D for each outcome, and the predicted ROC-AUC score of the prediction model MD constructed using the extracted feature quantities. It shows.
  • the ROC-AUC score related to the prediction of "presence or absence of severe disease" of prediction model MD1 which was constructed using feature quantities extracted from training fecal metabolite concentration data 31D with "presence or absence of severe disease” as the objective variable, is "0.72''.
  • the ROC of prediction model MD2 which was constructed using features extracted from training fecal metabolite concentration data 31D with "presence or absence of respiratory symptoms” as the objective variable, for predicting "presence or absence of respiratory symptoms” -AUC score was "0.74".
  • the ROC-AUC score of prediction model MD3, which was constructed using features extracted from training fecal metabolite concentration data 31D with "presence or absence of pneumonia” as the objective variable, for predicting "presence or absence of pneumonia” is , "0.88".
  • the ROC-AUC score related to the prediction of "the presence or absence of diarrhea” of the prediction model MD4, which was constructed using the feature values extracted from the training fecal metabolite concentration data 31D with "the presence or absence of diarrhea” as the objective variable, is , "0.78".
  • ROC-AUC for prediction of "presence or absence of onset of liver damage" of prediction model MD5 which was constructed using features extracted from training fecal metabolite concentration data 31D with "presence or absence of onset of liver damage” as the objective variable.
  • the score was "0.7".
  • ROC-AUC for prediction of "presence or absence of onset of renal disorder” of prediction model MD6, which was constructed using features extracted from training fecal metabolite concentration data 31D with "presence or absence of onset of renal disorder” as the objective variable.
  • the score was "0.69”.
  • the prediction model MD7 which was constructed using features extracted from training fecal metabolite concentration data 31D, with "presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l" as an objective variable,
  • the ROC-AUC score for predicting "presence or absence of decrease to less than / ⁇ l" was "0.57".
  • the prediction model MD8 which was constructed using features extracted from training fecal metabolite concentration data 31D, with "presence or absence of an increase in D-dimer to 1.0 mg/ml or more" as an objective variable, The ROC-AUC score for predicting "the presence or absence of an increase to .0 mg/ml or higher” was "0.84".
  • the prediction model MD9 which was constructed using features extracted from training fecal metabolite concentration data 31D, with "presence or absence of fibrinogen greater than 400 mg/dl" as the objective variable, The ROC-AUC score related to prediction was "0.75".
  • At least the genetic composition data of the intestinal flora (target gene composition data 221A), the blood metabolite data (target blood metabolite data 221B), and the bacterial composition data of the intestinal flora (target bacterial composition data 221C).
  • a person to be tested for predicting each outcome such as "presence or absence of severe illness" using one of these methods may be a person infected with COVID-19 or a person not infected with COVID-19.
  • test subject is a COVID-19 infected person
  • each outcome such as "presence or absence of severe disease” after the time when the test subject was diagnosed as “infected” is recorded in the target gene composition data 221A
  • subject Prediction is made using at least one of blood metabolite data 221B and target bacterial composition data 221C.
  • each outcome such as "presence or absence of severe disease” if the test subject is infected with COVID-19 is recorded in the target gene composition data 221A and target blood metabolite data.
  • the prediction is made using at least one of the data 221B and the target bacterial composition data 221C.
  • FIG. 10 is a diagram illustrating an overall outline of an analysis system 100 including the model generation device 1 and the inspection device 2 according to the present embodiment.
  • an analysis system 100 according to this embodiment includes a model generation device 1 and an inspection device 2.
  • the model generation device 1 is a computer configured to generate a trained estimator 51 (that is, a prediction model M) by performing machine learning. Specifically, the model generation device 1 acquires a plurality of learning data sets 3. Each learning data set 3 is configured by a combination of training omics data 31 (or feature data extracted from the training omics data 31) and outcome data 32.
  • the training omics data 31 is omics data obtained from biological samples of COVID-19 infected patients (112 cases) at the time of hospitalization. That is, the training omics data 31 includes training gene composition data 31A, training blood metabolite data 31B, training bacterial composition data 31C, and training omics data 31A, training blood metabolite data 31B, and training bacterial composition data 31C obtained from biological samples of COVID-19 infected patients (112 cases) at the time of hospitalization. Contains at least one of training fecal metabolite concentration data 31D.
  • the method for constructing the prediction model MA from the training gene composition data 31A is to construct the prediction models MB, MC, and The method for constructing each of the MDs is similar.
  • the method of predicting the risk of severe disease, etc. from the target gene composition data 221A of the test subject (subject) using the prediction model MA is to This method is similar to the method of predicting the risk of aggravation from each of the target blood metabolite data 221B, target bacterial composition data 221C, and target fecal metabolite concentration data 221D. Therefore, an example in which the omics data is gene composition data, that is, an example in which the training omics data 31 is the training gene composition data 31A and the target omics data 221 is the target gene composition data 221A will be described below.
  • gene composition data of intestinal flora (for example, training gene composition data 31A and target gene composition data 221A) refers to genes (DNA ) and the number of each gene.
  • Genetic composition data of the intestinal flora can be obtained, for example, by extracting and analyzing the DNA of intestinal bacteria from feces.
  • blood metabolite data refers to information contained in a blood sample (e.g., whole blood, plasma, serum, etc.). This data consists of the type of metabolite and the concentration of each metabolite. Blood metabolite data can be obtained, for example, by metabolomic analysis of a blood sample by GC/MS/MS analysis, LC-MS/MS analysis, or the like.
  • bacterial composition data of intestinal flora (for example, training bacterial composition data 31C and target bacterial composition data 221C)" refers to the types of intestinal bacteria extracted from intestinal flora and each This data consists of the number (relative abundance ratio) of intestinal bacteria.
  • Bacterial composition data of the intestinal flora can be obtained, for example, by extracting intestinal bacterial DNA from feces and conducting metagenomic analysis.
  • outcome data 32 is data showing each outcome such as "presence or absence of severe disease” observed for each of the COVID-19 infected persons (112 cases).
  • outcome data 32 includes the "presence or absence of symptoms when contracting COVID-19", “presence of severity”, and “presence of complications” observed for each person infected with COVID-19 (112 cases). This data indicates the presence or absence of long-term hospitalization.
  • symptoms include “respiratory symptoms,””pneumonia,” and “diarrhea.”
  • complications include “liver damage,” “kidney damage,” and “thrombosis,” and “thrombosis” includes “a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l,” and “D- This includes “an increase in dimer to 1.0 mg/ml or more” and “an increase in fibrinogen to more than 400 mg/dl.”
  • each learning data set 3 the training omics data 31 and the outcome data 32 correspond.
  • one learning data set is created by combining training omics data 31 in a biological sample of a certain infected person with outcome data 32 indicating each outcome such as "presence or absence of severe disease" observed for the same infected person. 3 are made up.
  • the model generation device 1 performs machine learning of the estimator 51 (prediction model M) using the plurality of acquired learning data sets 3. As a result, a trained estimator 51 is generated. In particular, the model generation device 1 learns the relationship between the feature amounts extracted from the training omics data 31 included in each learning dataset 3 and each outcome indicated by the outcome data 32 included in each learning dataset 3. .
  • the testing device 2 is configured to use a trained estimator 51 (prediction model M) to predict each outcome such as "presence or absence of severe disease” from the omics data of the test subject (target omics data 221). It is a computer that has been Specifically, the inspection device 2 acquires the target omics data 221.
  • the target omics data 221 is omics data obtained from a biological sample of a test subject. As described above, here, an example will be described in which the target omics data 221 is the target gene composition data 221A obtained from a biological sample of a test subject.
  • the testing device 2 uses the estimator 51 trained by machine learning to predict (estimate) each outcome such as "presence or absence of severe disease" from the acquired target omics data 221. Then, the inspection device 2 outputs information regarding the predicted result (that is, the predicted result).
  • the estimator 51 receives the input of the feature amount of the target omics data 221 output from the extraction unit 112, and determines various information such as "presence or absence of severe disease” from the feature amount. Configured to predict (estimate) an outcome.
  • the estimator 51 uses a known algorithm such as random forest to extract each outcome, such as "presence or absence of severe disease,” from the explanatory variables (features) selected (extracted) for each outcome by the extraction unit 112. Predict.
  • the trained estimator 51 (prediction model M) used by the inspection device 2 performs machine learning on a plurality of learning data sets 3 each composed of a combination of training omics data 31 and outcome data 32. It is built.
  • the trained estimator 51 includes a plurality of learning methods each configured by a combination of outcome data 32 and at least one of training gene composition data 31A, training blood metabolite data 31B, and training bacterial composition data 31C, for example. It is constructed in advance by performing machine learning on Dataset 3.
  • the outcome data 32 indicates each outcome such as "presence or absence of severe disease" observed for COVID-19 patients (specifically, COVID-19 infected patients (112 cases)).
  • the trained estimator 51 uses machine learning that uses the feature quantities extracted from the target omics data 221 as explanatory variables and each outcome such as "presence or absence of severe disease" indicated by the outcome data 32 as an objective variable. It is pre-built by doing.
  • the trained estimator 51 uses, for example, a feature amount extracted from at least one of the training gene composition data 31A, the training blood metabolite data 31B, and the training bacterial composition data 31C as an explanatory variable, and uses the features indicated by the outcome data 32. It is constructed in advance by performing machine learning using each outcome as the objective variable.
  • the model generation device 1 and the inspection device 2 are connected to each other via a network.
  • the type of network may be appropriately selected from, for example, the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, and the like.
  • the method of exchanging data between the model generation device 1 and the inspection device 2 does not need to be limited to this example, and may be selected as appropriate depending on the embodiment.
  • data may be exchanged between the model generation device 1 and the inspection device 2 using a storage medium.
  • the model generation device 1 and the inspection device 2 are each configured by separate computers.
  • the configuration of the analysis system 100 according to the present embodiment does not need to be limited to such an example, and may be determined as appropriate depending on the embodiment.
  • the model generation device 1 and the inspection device 2 may be an integrated computer.
  • at least one of the model generation device 1 and the inspection device 2 may be configured by a plurality of computers.
  • the testing device 2 collects genetic composition data of the intestinal flora (target gene composition data 221A), blood metabolite data (target blood metabolite data 221B), and bacterial composition data of the intestinal flora (target bacterial composition data 221C).
  • a person to be tested for predicting each outcome such as "presence or absence of severe illness” using at least one of the following may be a person infected with COVID-19 or a person not infected with COVID-19. If the person to be tested is a person infected with COVID-19, the testing device 2 calculates each outcome such as ⁇ presence or absence of severe disease'' after the time when the person to be tested is diagnosed as having been infected, based on the target gene.
  • the prediction is made using at least one of composition data 221A, target blood metabolite data 221B, and target bacterial composition data 221C.
  • the testing device 2 records each outcome such as "presence or absence of severe disease” in the case that the test subject is infected with COVID-19, and the target gene composition data 221A, The prediction is made using at least one of the target blood metabolite data 221B and the target bacterial composition data 221C.
  • FIG. 11 schematically illustrates an example of the hardware configuration of the model generation device 1 according to this embodiment.
  • a control section 11 a storage section 12, a communication interface 13, an external interface 14, an input device 15, an output device 16, and a drive 17 are electrically connected.
  • the communication interface and external interface are described as "communication I/F" and "external I/F.”
  • the control unit 11 includes a CPU (Central Processing Unit) that is a hardware processor, a RAM (Random Access Memory), a ROM (Read Only Memory), etc., and is configured to execute information processing based on programs and various data. Ru.
  • the control unit 11 may further include a GPU (Graphics Processing Unit) not shown.
  • the storage unit 12 is an example of a memory, and includes, for example, a hard disk drive, a solid state drive, or the like. In this embodiment, the storage unit 12 stores various information such as the model generation program 81, the learning data set 3, and the learning result data 125.
  • the model generation program 81 is a program for causing the model generation device 1 to execute machine learning processing to generate the trained estimator 51 (prediction model M).
  • the model generation program 81 includes a series of instructions for the machine learning process.
  • the learning data set 3 is used for machine learning of the estimator 51.
  • the learning result data 125 indicates information regarding the trained estimator 51 generated by performing machine learning. In this embodiment, the learning result data 125 is generated as a result of executing the model generation program 81.
  • the communication interface 13 is, for example, a wired LAN (Local Area Network) module, a wireless LAN module, or the like, and is an interface for performing wired or wireless communication via a network.
  • the model generation device 1 can use the communication interface 13 to perform data communication with other information processing devices via a network.
  • the external interface 14 is, for example, a USB (Universal Serial Bus) port, a dedicated port, or the like, and is an interface for connecting to an external device. The type and number of external interfaces 14 may be selected arbitrarily.
  • the learning data set 3 (in particular, each of the training omics data 31 and outcome data 32) may be obtained from an external device.
  • the training omics data 31 may be obtained from an external measurement device, and the outcome data 32 may be obtained from an external outcome management server.
  • the model generation device 1 may be connected to the external measurement device or outcome management server via at least one of the communication interface 13 and the external interface 14.
  • the input device 15 is, for example, a device for performing input such as a mouse or a keyboard.
  • the output device 16 is, for example, a device for outputting, such as a display or a speaker.
  • An operator such as a user can operate the model generation device 1 by using the input device 15 and the output device 16.
  • the learning data set 3 may be obtained by inputting via the input device 15.
  • the drive 17 is, for example, a CD drive, a DVD drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91.
  • the storage medium 91 stores information such as programs through electrical, magnetic, optical, mechanical, or chemical action so that computers, other devices, machines, etc. can read various information such as stored programs. It is a medium that accumulates by At least one of the model generation program 81 and the learning data set 3 may be stored in the storage medium 91.
  • the model generation device 1 may acquire at least one of the model generation program 81 and the learning data set 3 from this storage medium 91.
  • a disk-type storage medium such as a CD or a DVD is illustrated as an example of the storage medium 91.
  • the type of storage medium 91 is not limited to the disk type, and may be other than the disk type.
  • An example of a storage medium other than a disk type is a semiconductor memory such as a flash memory.
  • the type of drive 17 may be arbitrarily selected depending on the type of storage medium 91.
  • the control unit 11 may include multiple hardware processors.
  • the hardware processor may include a microprocessor, a field-programmable gate array (FPGA), a digital signal processor (DSP), and the like.
  • the storage unit 12 may be configured by a RAM and a ROM included in the control unit 11. At least one of the communication interface 13, external interface 14, input device 15, output device 16, and drive 17 may be omitted.
  • the model generation device 1 may be composed of multiple computers. In this case, the hardware configurations of the computers may or may not match. Further, the model generation device 1 may be an information processing device designed exclusively for the provided service, or may be a general-purpose server device, a PC (Personal Computer), or the like.
  • FIG. 12 schematically illustrates an example of the hardware configuration of the inspection device 2 according to this embodiment.
  • a control section 21 a storage section 22, a communication interface 23, an external interface 24, an input device 25, an output device 26, and a drive 27 are electrically connected. It is a computer.
  • the control unit 21 to drive 27 and storage medium 92 of the inspection device 2 may be configured similarly to the control unit 11 to drive 17 and storage medium 91 of the model generation device 1, respectively.
  • the control unit 21 includes a CPU, which is a hardware processor, a RAM, a ROM, etc., and is configured to execute various information processing based on programs and data.
  • the storage unit 22 includes, for example, a hard disk drive, a solid state drive, or the like. In this embodiment, the storage unit 22 stores various information such as the inspection program 82 and learning result data 125.
  • the inspection program 82 is a program for causing the inspection device 2 to execute a prediction process to perform a prediction task (estimation task) using the trained estimator 51 (prediction model M).
  • the inspection program 82 includes a series of instructions for the prediction process.
  • At least one of the test program 82 and the learning result data 125 may be stored in the storage medium 92.
  • the inspection device 2 may acquire at least one of the inspection program 82 and the learning result data 125 from the storage medium 92.
  • the target omics data 221 may be obtained from an external device, for example, from an external measurement device.
  • the inspection device 2 may be connected to the external device via at least one of the communication interface 23 and the external interface 24.
  • the target omics data 221 may be obtained by input via the input device 25.
  • the control unit 21 may include multiple hardware processors.
  • the hardware processor may be comprised of a microprocessor, FPGA, DSP, etc.
  • the storage unit 22 may be configured by a RAM and a ROM included in the control unit 21. At least one of the communication interface 23, external interface 24, input device 25, output device 26, and drive 27 may be omitted.
  • the inspection device 2 may be composed of multiple computers. In this case, the hardware configurations of the computers may or may not match.
  • the inspection device 2 may be an information processing device designed exclusively for the service provided, as well as a general-purpose server device, a general-purpose PC, a tablet PC, a terminal device, or the like.
  • FIG. 13 schematically illustrates an example of the software configuration of the model generation device 1 according to this embodiment.
  • the control unit 11 of the model generation device 1 loads the model generation program 81 stored in the storage unit 12 into the RAM.
  • the control unit 11 uses the CPU to interpret and execute instructions included in the model generation program 81 developed in the RAM, thereby controlling each component.
  • the model generation device 1 operates as a computer including an acquisition section 111, an extraction section 112, a learning processing section 113, and a storage processing section 114 as software modules. That is, in this embodiment, each software module of the model generation device 1 is realized by the control unit 11 (CPU).
  • the acquisition unit 111 is configured to acquire a plurality of learning data sets 3 each configured by a combination of training omics data 31 and outcome data 32.
  • the extraction unit 112 extracts (selects) explanatory variables (features) used in the learning process of the estimator 51 (prediction model M) from the training omics data 31 included in each of the plurality of learning data sets 3.
  • the extraction unit 112 may extract explanatory variables (feature amounts) that explain each outcome such as "presence or absence of severe illness" from the training omics data 31 using a known algorithm such as the Boruta algorithm. Further, if the feature quantity to be extracted from the training omics data 31 can be specified in advance according to each outcome, the extraction unit 112 extracts the prespecified feature quantity from the training omics data 31 for each outcome. It may be extracted depending on the
  • the feature amount may include at least one of training gene composition data 31A, training blood metabolite data 31B, and training bacterial composition data 31C in a biological sample collected at the time of hospitalization of the person to be examined (subject). Then, the results are extracted by feature quantity selection using each outcome observed in the test subject, such as "presence or absence of severe illness" when infected with COVID-19, as the objective variable.
  • the extraction unit 112 that receives the input of the training omics data 31 uses the training omics as an explanatory variable (feature amount) that explains (predicts) "presence or absence of severe disease". From the data 31, data regarding the genes shown in (1) to (20) above are extracted.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts the genes shown in (21) to (40) from the training omics data 31 as feature quantities for predicting "the presence or absence of respiratory symptoms.” Extract the data.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts data on the genes shown in (41) to (60) above from the training omics data 31 as feature quantities for predicting "presence or absence of pneumonia”. Extract.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts data about the genes shown in (61) to (80) above from the training omics data 31 as feature quantities for predicting "the presence or absence of diarrhea”. Extract.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts data about the genes shown in (81) to (100) above from the training omics data 31 as feature quantities for predicting "presence or absence of onset of liver damage.” Extract.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts data about the genes shown in (101) to (120) above from the training omics data 31 as feature quantities for predicting "presence or absence of onset of renal disorder.” Extract.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts the above-mentioned (121) to ( 140) is extracted.
  • the extraction unit 112 that has received the input of the training omics data 31 uses the above-mentioned (141) to Data about the genes shown in (160) are extracted.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts the genes shown in (161) to (180) above from the training omics data 31 as a feature quantity for predicting "the presence or absence of fibrinogen greater than 400 mg/dl”. Extract data about.
  • the extraction unit 112 that has received the input of the training omics data 31 extracts data about the genes shown in (201) to (220) above from the training omics data 31 as feature quantities for predicting "presence or absence of long-term hospitalization.” do.
  • the learning processing unit 113 is configured to perform machine learning of the estimator 51 (prediction model M) using the acquired learning data set 3. For example, the learning processing unit 113 learns a machine learning model using the feature amount (explanatory variable) selected (extracted) from the training omics data 31 by the extraction unit 112 and the outcome data 32 (objective variable). In machine learning, for each learning data set 3, the outcome predicted by the estimator 51 is calculated based on the training omics data 31 by giving the training omics data 31 (in particular, the features extracted from the training omics data 31) as input. The estimator 51 is configured by training the estimator 51 to match the outcome indicated by the corresponding outcome data 32.
  • the estimator 51 performs a process such that the outcome predicted by the estimator 51 from the feature quantity extracted from the training omics data 31 matches the outcome indicated by the outcome data 32 corresponding to the training omics data 31. be trained.
  • the "presence or absence of severe disease” predicted from “the genes (features) shown in (1) to (20) above extracted from the training omics data 31” is the “severe disease” indicated by the outcome data 32.
  • the estimator 51 (A1) that is, the prediction model MA1 is trained to adapt to the "presence or absence of the The “presence or absence of respiratory symptoms” predicted from the “genes (features) shown in (21) to (40) above extracted from the training omics data 31” is the “respiratory symptoms” indicated by the outcome data 32.
  • the estimator 51 (A2) (that is, the prediction model MA2) is trained so as to adapt to "the presence or absence of expression.”
  • the “presence or absence of pneumonia” predicted from “the genes (features) shown in (41) to (60) above extracted from the training omics data 31” is the “presence or absence of pneumonia” indicated by the outcome data 32.
  • the estimator 51 (A3) i.e., the prediction model MA3) is trained to fit .
  • the “presence or absence of diarrhea” predicted from the “genes (features) shown in (61) to (80) above extracted from the training omics data 31” is the “presence or absence of diarrhea” indicated by the outcome data 32.
  • the estimator 51 (A4) (i.e., the prediction model MA4) is trained to fit ⁇ .
  • ⁇ Presence or absence of onset of liver damage'' predicted from ⁇ genes (feature amounts) shown in (81) to (100) above extracted from training omics data 31'' is predicted from ⁇ onset of liver damage'' shown by outcome data 32.
  • the estimator 51 (A5) (that is, the prediction model MA5) is trained to adapt to the "presence or absence of”.
  • the ⁇ presence or absence of the onset of renal disorder'' predicted from ⁇ the genes (feature amounts) shown in (101) to (120) above extracted from the training omics data 31'' is The estimator 51 (A6) (that is, the prediction model MA6) is trained to adapt to the "presence or absence of”.
  • the outcome data 32 predicts the presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l, which is predicted from the genes (features) shown in (121) to (140) above extracted from the training omics data 31.
  • the estimator 51 (A7) that is, the prediction model MA7) is trained to match "the presence or absence of a decrease in platelets to less than 20 ⁇ 10 4 / ⁇ l" indicated by .
  • the outcome data is ⁇ the presence or absence of an increase in D-dimer to 1.0 mg/ml or more'' predicted from ⁇ the genes (features) shown in (141) to (160) above extracted from the training omics data 31''.
  • the estimator 51 (A8) (that is, the prediction model MA8) is trained to match "the presence or absence of an increase in D-dimer to 1.0 mg/ml or more" indicated by 32.
  • the estimator 51 (A9) i.e., the prediction model MA9) is trained to adapt to "the presence or absence of fibrinogen greater than /dl".
  • the ⁇ presence or absence of long-term hospitalization'' predicted from ⁇ the genes (features) shown in (201) to (220) above extracted from the training omics data 31'' is the same as the ⁇ presence or absence of long-term hospitalization'' indicated by the outcome data 32.
  • Estimator 51 (A11) ie, predictive model MA11 is trained to adapt.
  • the storage processing unit 114 generates information regarding the trained estimator 51 (prediction model M) generated by machine learning as learning result data 125, and stores the generated learning result data 125 in a predetermined storage area. It is composed of The learning result data 125 may be configured as appropriate to include information for reproducing the trained estimator 51 (prediction model M).
  • each training omics data 31 (especially its feature amount) can be , it is possible to obtain prediction (estimation) results regarding outcomes.
  • the learning processing unit 113 uses prediction results (estimation results) obtained for each training omics data 31 (particularly, its feature amounts) and each outcome data 32 corresponding to each training omics data 31. It is configured to adjust the parameters of the estimator 51 so that the error between the estimator 51 and the outcome becomes smaller.
  • the storage processing unit 114 is configured to generate learning result data 125 for reproducing the trained estimator 51 (prediction model M) generated by the machine learning.
  • the configuration of the learning result data 125 is not particularly limited and may be determined as appropriate depending on the embodiment.
  • the learning result data 125 may include information indicating the values of each parameter obtained by adjusting the machine learning.
  • the learning result data 125 may include information indicating the structure of the estimator 51. The structure may be specified by, for example, the number of layers, the type of each layer, the number of nodes included in each layer, the connection relationship between nodes in adjacent layers, and the like.
  • FIG. 14 schematically illustrates an example of the software configuration of the inspection device 2 according to this embodiment.
  • the control unit 21 of the inspection device 2 loads the inspection program 82 stored in the storage unit 22 into the RAM. Then, the control unit 21 uses the CPU to interpret and execute instructions included in the inspection program 82 loaded in the RAM, thereby controlling each component.
  • the inspection device 2 according to the present embodiment operates as a computer including an acquisition section 211, an extraction section 212, a prediction section 213, and an output section 214 as software modules. That is, in this embodiment, each software module of the inspection device 2 is also realized by the control unit 21 (CPU) similarly to the model generation device 1.
  • the acquisition unit 211 is configured to acquire target omics data 221.
  • the acquisition unit 211 acquires, for example, at least one of target gene composition data 221A, target blood metabolite data 221B, and target bacterial composition data 221C in a biological sample collected from a test subject (subject).
  • the acquisition unit 211 may acquire the target omics data 221 from an external device, for example, from an external measurement device. That is, the testing device 2 does not need to measure the target omics data 221 from the biological sample collected from the test subject.
  • the extraction unit 212 extracts (selects) explanatory variables (features) used in the prediction process using the estimator 51 (prediction model M) from the target omics data 221.
  • the extraction unit 212 may extract explanatory variables (feature amounts) that explain each outcome such as "presence or absence of severe disease” from the target omics data 221 using a known algorithm such as the Boruta algorithm. For example, if the extraction unit 112 of the model generation device 1 is trained to extract feature quantities according to each outcome such as “presence or absence of severe disease” from the training omics data 31, the trained extraction unit 112 may be used as the extraction unit 212 of the inspection device 2.
  • the extraction unit 212 extracts the prespecified feature quantity from the target omics data 221 for each outcome. It may be extracted depending on the
  • the extraction unit 212 that has received the input of the target omics data 221 extracts the above-mentioned (1) from the target omics data 221 as an explanatory variable (feature amount) that explains (predicts) "presence or absence of severe disease”. Extract data regarding the genes shown in ⁇ (20).
  • the extraction unit 212 that has received the input of the target omics data 221 extracts information about the genes shown in (21) to (40) from the target omics data 221 as feature quantities for predicting "the presence or absence of respiratory symptoms.” Extract the data.
  • the extraction unit 212 that has received the input of the target omics data 221 extracts data about the genes shown in (41) to (60) above from the target omics data 221 as feature quantities for predicting "presence or absence of pneumonia”.
  • Extract The extraction unit 212 that has received the input of the target omics data 221 extracts data about the genes shown in (61) to (80) above from the target omics data 221 as feature quantities for predicting "the presence or absence of diarrhea”. Extract. The extraction unit 212 that has received the input of the target omics data 221 extracts data about the genes shown in (81) to (100) above from the target omics data 221 as feature quantities for predicting "presence or absence of onset of liver damage.” Extract. The extraction unit 212 that has received the input of the target omics data 221 extracts data about the genes shown in (101) to (120) above from the target omics data 221 as feature quantities for predicting "presence or absence of onset of renal disorder.” Extract.
  • the extraction unit 212 that has received the input of the target omics data 221 extracts the above-mentioned (121) to ( 140) is extracted.
  • the extraction unit 212 which has received the input of the target omics data 221, uses the above-mentioned (141) to Data about the genes shown in (160) are extracted.
  • the extraction unit 212 which has received the input of the target omics data 221, extracts the genes shown in (161) to (180) from the target omics data 221 as feature quantities for predicting "the presence or absence of fibrinogen greater than 400 mg/dl”. Extract data about.
  • the extraction unit 212 that has received the input of the target omics data 221 extracts data about the genes shown in (201) to (220) above from the target omics data 221 as feature quantities for predicting "presence or absence of long-term hospitalization". do.
  • the prediction unit 213 retains the learning result data 125 and is equipped with a trained estimator 51 (prediction model M).
  • the prediction unit 213 is configured to use the trained estimator 51 to predict (estimate) the outcome for the acquired target omics data 221.
  • the prediction unit 213 inputs at least one of the target gene composition data 221A, the target blood metabolite data 221B, and the target bacterial composition data 221C to the trained estimator 51 prepared in advance. , predict at least one selected from the group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization when contracting COVID-19.
  • the prediction unit 213 inputs the feature quantity extracted from the target omics data 221 by the extraction unit 212 to the trained estimator 51, and makes a prediction ( (estimation) result from the estimator 51.
  • the prediction unit 213 notifies the output unit 214 of the prediction results for each outcome obtained from the estimator 51.
  • the trained estimator 51 (A1) (that is, the prediction model MA1) uses the genes (features) shown in (1) to (20) above to estimate the test subject (target omics data 221). Predict whether the disease will worsen.
  • the trained estimator 51 (A2) (that is, the prediction model MA2) uses the genes (features) shown in (21) to (40) above to determine the “respiratory prediction” of the test subject (target omics data 221). Predict whether or not organic symptoms will occur.
  • the trained estimator 51 (A3) (that is, the prediction model MA3) uses the genes (features) shown in (41) to (60) above to determine whether the test subject (target omics data 221) has "pneumonia". The presence or absence of expression is predicted.
  • the trained estimator 51 (A4) (that is, the prediction model MA4) uses the genes (features) shown in (61) to (80) above to determine whether the "diarrhea" of the test subject (target omics data 221) The presence or absence of expression is predicted.
  • the trained estimator 51 (A5) (that is, the prediction model MA5) uses the genes (features) shown in (81) to (100) above to estimate the “hepatic function” of the test subject (target omics data 221). Predict whether or not a disorder will develop.
  • the trained estimator 51 (A6) (that is, the prediction model MA6) uses the genes (features) shown in (101) to (120) above to determine the “renal Predict whether or not a disorder will develop.
  • the trained estimator 51 (A7) (that is, the prediction model MA7) uses the genes (features) shown in (121) to (140) above to estimate the "platelet count” of the test subject (target omics data 221). Predict the presence or absence of a decrease to less than 20 ⁇ 10 4 / ⁇ l.
  • the trained estimator 51 (A8) (that is, the prediction model MA8) uses the genes (features) shown in (141) to (160) above to determine the “D” of the test subject (target omics data 221).
  • the trained estimator 51 (A9) (that is, the prediction model MA9) uses the genes (features) shown in (161) to (180) above to calculate the “400 mg /dl to predict the presence or absence of fibrinogen.
  • the trained estimator 51 (A11) (that is, the prediction model MA11) uses the genes (features) shown in (201) to (220) above to determine the “long-term prediction” of the test subject (target omics data 221). Predict whether or not you will be hospitalized.
  • the output unit 214 is configured to output information regarding the prediction results for each outcome notified from the prediction unit 213 (estimator 51). For example, the output unit 214 may cause the output device 16 to output information regarding the prediction results for each outcome.
  • each software module of the model generation device 1 and the inspection device 2 will be explained in detail in the operation example described later.
  • an example is described in which each of the software modules of the model generation device 1 and the inspection device 2 are both implemented by a general-purpose CPU.
  • some or all of the software modules may be implemented by one or more dedicated processors (eg, graphics processing units).
  • Each of the above modules may be implemented as a hardware module.
  • software modules may be omitted, replaced, or added as appropriate depending on the embodiment.
  • FIG. 15 is a flowchart illustrating an example of a processing procedure related to machine learning by the model generation device 1 according to the present embodiment.
  • the processing procedure of the model generation device 1 described below is an example of a model generation method.
  • the processing procedure of the model generation device 1 described below is only an example, and each step may be changed as much as possible. Further, steps may be omitted, replaced, or added as appropriate in the following processing procedure depending on the embodiment.
  • Step S101 In step S ⁇ b>101 , the control unit 11 operates as the acquisition unit 111 and acquires a plurality of learning data sets 3 each configured by a combination of training omics data 31 and outcome data 32 .
  • the number of learning data sets 3 to be acquired does not need to be particularly limited, and may be determined as appropriate depending on the embodiment so that machine learning can be performed.
  • the control unit 11 After acquiring the learning data set 3, the control unit 11 advances the process to the next step S102.
  • Step S102 the control unit 11 operates as the extraction unit 112 and extracts feature amounts from the training omics data 31 included in each acquired learning data set 3. That is, the control unit 11 extracts (selects) explanatory variables (features) used in the learning process of the estimator 51 (prediction model M). For example, the control unit 11 extracts, from the training omics data 31, explanatory variables (features) that explain each outcome such as "presence or absence of severe disease". After extracting the feature amount, the control unit 11 advances the process to the next step S103.
  • Step S103 the control unit 11 operates as the learning processing unit 113 and performs machine learning of the estimator 51 (prediction model M) using the acquired learning data set 3. Specifically, the control unit 11 performs machine learning of the estimator 51 using the feature amount extracted from the training omics data 31 in step S102 and the outcome data 32 corresponding to the training omics data 31. .
  • the control unit 11 initializes the estimator 51 (prediction model M) that is the subject of machine learning processing.
  • the structure and initial values of parameters of the estimator 51 (prediction model M) may be given by a template, or may be determined by input from an operator. Further, when performing additional learning or relearning, the control unit 11 may initialize the estimator 51 (prediction model M) based on learning result data obtained by past machine learning.
  • control unit 11 uses machine learning to change the outcome predicted (estimated) from each training omics data 31 (particularly its feature amount) to the outcome indicated by each outcome data 32 corresponding to each training omics data 31. Train the estimator 51 to fit. In other words, the control unit 11 uses machine learning to adjust the values of the parameters of the estimator 51 so that the outcome estimated from the feature amount of each training omics data 31 matches the outcome indicated by each corresponding outcome data 32. adjust.
  • This training process may use a stochastic gradient descent method, a mini-batch gradient descent method, or the like.
  • the control unit 11 inputs each training omics data 31 (particularly its feature amount) to the estimator 51, and performs forward calculation processing. Through this forward calculation process, the control unit 11 obtains the outcome prediction result (estimation result) for each training omics data 31 (particularly its feature amount) from the estimator 51.
  • control unit 11 calculates the error between the obtained estimation result and the outcome indicated by each outcome data 32 corresponding to each training omics data 31.
  • a loss function may be used to calculate the error.
  • the control unit 11 calculates the gradient of the calculated error.
  • the control unit 11 calculates the error in the value of the parameter of the estimator 51 using the calculated error gradient using the error backpropagation method.
  • the control unit 11 updates the values of the parameters of the estimator 51 based on the calculated error.
  • the degree to which parameter values are updated may be adjusted by the learning rate.
  • the learning rate may be specified by the operator or may be provided as a set value within the program.
  • the control unit 11 adjusts the parameter values for each training omics data 31 (particularly its feature amounts) so that the sum of the calculated errors becomes small. For example, the control unit 11 may repeat the adjustment of parameter values through the series of update processes described above until a predetermined condition is satisfied, such as executing the adjustment a specified number of times or the sum of calculated errors being equal to or less than a threshold value. As a result of this machine learning process, the control unit 11 uses a trained estimator 51 (prediction model M ) can be generated. When the machine learning process is completed, the control unit 11 advances the process to the next step S104.
  • Step S104 the control unit 11 operates as the storage processing unit 114 and generates information regarding the trained estimator 51 (prediction model M) generated by machine learning as learning result data 125. Then, the control unit 11 stores the generated learning result data 125 in a predetermined storage area.
  • the predetermined storage area may be, for example, the RAM in the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination thereof.
  • the storage medium may be, for example, a CD, a DVD, etc.
  • the control unit 11 may store the learning result data 125 in the storage medium via the drive 17.
  • the external storage device may be, for example, a data server such as a NAS (Network Attached Storage).
  • the control unit 11 may use the communication interface 13 to store the learning result data 125 in the data server via the network.
  • the external storage device may be, for example, an external storage device connected to the model generation device 1 via the external interface 14.
  • control unit 11 ends the processing procedure of the model generation device 1 according to this operation example.
  • the generated learning result data 125 may be provided to the inspection device 2 at any timing.
  • the inspection device 2 may obtain the learning result data 125 by accessing the model generation device 1 or the data server via the network using the communication interface 23. Furthermore, the inspection device 2 may acquire the learning result data 125 via the storage medium 92. Further, the learning result data 125 may be incorporated into the inspection device 2 in advance.
  • control unit 11 may update or newly generate the learning result data 125 by repeating the processes of steps S101 to S104 described above regularly or irregularly. During this repetition, at least a portion of the learning data set 3 used for machine learning may be changed, modified, added, deleted, etc. as appropriate. Then, the control unit 11 may update the learning result data 125 held by the testing device 2 by providing the updated or newly generated learning result data 125 to the testing device 2 using any method.
  • FIG. 16 is a flowchart illustrating an example of a processing procedure regarding execution of a prediction task by the inspection device 2 according to the present embodiment.
  • the processing procedure of the inspection device 2 described below is an example of an estimation method. However, the processing procedure of the inspection device 2 described below is only an example, and each step may be changed as much as possible. Further, steps may be omitted, replaced, or added as appropriate in the following processing procedure depending on the embodiment.
  • step S201 acquisition step
  • the control unit 21 operates as the acquisition unit 211 and acquires the target omics data 221 that is the target of the prediction task.
  • the control unit 21 acquires, for example, at least one of target gene composition data 221A, target blood metabolite data 221B, and target bacterial composition data 221C in a biological sample collected from a test subject (subject).
  • the target omics data 221 is obtained by measuring a biological sample collected from a test subject (subject).
  • the configuration of the target omics data 221 is similar to the training omics data 31 in the learning data set 3.
  • the data format of the target omics data 221 may be selected as appropriate depending on the embodiment.
  • the control unit 21 may directly acquire the target omics data 221.
  • control unit 21 may obtain the target omics data 221 indirectly, for example, via a network, a sensor, another computer, the storage medium 92, or the like. After acquiring the target omics data 221, the control unit 21 advances the process to the next step S202.
  • Step S202 the control unit 21 operates as the extraction unit 212, and extracts from the target omics data 221 feature amounts that explain each outcome such as “presence or absence of severe disease”. That is, the extraction unit 212 extracts (selects) explanatory variables (features) to be output to the estimator 51 (prediction model M) from the target omics data 221. After extracting the feature amount, the control unit 21 advances the process to the next step S203.
  • step S203 the control unit 21 operates as the prediction unit 213. That is, the control unit 21 inputs the target omics data 221 into a trained estimator 51 (prediction model M) prepared in advance to predict each outcome such as "presence or absence of severe disease". For example, the control unit 21 inputs at least one of the target gene composition data 221A, the target blood metabolite data 221B, and the target bacterial composition data 221C to the trained estimator 51, thereby determining whether the person is infected with COVID-19. At least one selected from the group consisting of the risk of developing symptoms, risk of worsening, risk of developing complications, and risk of long-term hospitalization is predicted.
  • step S203 the control unit 21 first refers to the learning result data 125 and sets the trained estimator 51 (prediction model M). Then, the control unit 21 uses the trained estimator 51 to estimate a solution to the prediction task for the acquired target omics data 221 (particularly its feature amount).
  • This estimation calculation process may be similar to the forward calculation process in the machine learning training process described above.
  • the control unit 21 inputs the feature amount extracted from the target omics data 221 to the trained estimator 51, and executes the forward calculation process of the trained estimator 51.
  • the control unit 21 can obtain from the estimator 51 the result (prediction result) of estimating the solution of the prediction task for the target omics data 221 (particularly its feature amount).
  • the control unit 21 advances the process to the next step S204.
  • step S204 output step
  • the control unit 21 operates as the output unit 214 and outputs information regarding the prediction result obtained in S203.
  • the output destination and the content of the information to be output may be determined as appropriate depending on the embodiment.
  • the control unit 21 may directly output the prediction result obtained in step S203 to the output device 26 or another computer output device. Further, the control unit 21 may perform some information processing based on the obtained prediction result. Then, the control unit 21 may output the result of performing the information processing as information regarding the prediction result.
  • the output of the result of performing this information processing may include controlling the operation of the controlled device according to the prediction result.
  • the output destination may be, for example, the output device 26, an output device of another computer, a controlled device, or the like.
  • control unit 21 ends the processing procedure of the inspection device 2 according to this operation example.
  • an example of constructing the predictive model M (estimator 51) using random forest has been described, but the learning method performed by the learning processing unit 113 to construct the predictive model M (estimator 51) is , but not limited to this.
  • an algorithm for constructing the predictive model M a known algorithm such as an algorithm used for machine learning can be used.
  • machine learning algorithms include support vector machines with linear kernels (SVM linear), support vector machines with RBF kernels (SVM rbf), neural networks, generalized linear models, Examples include regularized linear discriminant analysis and regularized logistic regression.
  • the extraction unit 112 (or extraction unit 212) and the estimator 51 have been described as two different functional units.
  • the extraction unit 112 (or the extraction unit 212) and the estimator 51 may be configured integrally, for example, as a neural network (neural network module).
  • a neural network module performs processing as the extraction unit 112 (or extraction unit 212), that is, an operation for extracting elements that satisfy a predetermined condition from a target set (hereinafter also referred to as "extraction operation").
  • extraction operation an operation for extracting elements that satisfy a predetermined condition from a target set
  • the contents of the extraction operation are not particularly limited as long as the extraction operation is a non-differentiable operation that narrows down the multiple elements calculated in the process of the regression operation to some elements (for example, desirable elements). It may be determined as appropriate depending on the embodiment.
  • Machine learning is performed by training the neural network module for each training data set 3 so that the values regressed from the training omics data 31 match each outcome indicated by the outcome data 32 using the neural network module. be done. Training the neural network module is performed by adjusting the value of each calculation parameter through regression trial processing using forward propagation and calculation parameter adjustment processing using back propagation. As a result of this machine learning, a trained neural network module that has acquired the ability to regress each outcome can be generated from the omics data (training omics data 31). By using such a trained neural network module, the testing device 2 can perform a prediction task of predicting each outcome of the test subject (subject) from the target omics data 221.
  • the testing device 2 uses a trained neural network module (prediction model M) to generate target gene composition data 221A and target blood metabolite data in a biological sample collected from a test subject (subject). From at least one of the data 221B and the target bacterial composition data 221C, it is possible to predict each outcome such as "presence or absence of severe illness" of the test subject.
  • prediction model M a trained neural network module
  • S11...Measurement step S201...Acquisition step, S203...Prediction step, S204...Output step, M...Prediction model, 2...Inspection device, 31A...Training gene composition data (genetic composition data of intestinal flora), 31B...Training blood metabolite data (blood metabolite data), 221A...Target gene composition data (genetic composition data of intestinal flora), 221B...Target blood metabolite data (blood metabolite data) 211... Acquisition unit, 213... Prediction unit, 214... Output unit, 51... Estimator (prediction model)

Landscapes

  • Chemical & Material Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Organic Chemistry (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biotechnology (AREA)
  • Microbiology (AREA)
  • Immunology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • Physics & Mathematics (AREA)
  • Analytical Chemistry (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Medicinal Chemistry (AREA)
  • Hematology (AREA)
  • Urology & Nephrology (AREA)
  • Sustainable Development (AREA)
  • Toxicology (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • Cell Biology (AREA)
  • Food Science & Technology (AREA)
  • General Physics & Mathematics (AREA)
  • Pathology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

本実施形態に係る検査方法は、被験者から採取された生体試料中の、腸内菌 叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを測定する。腸内菌叢の遺伝子組成データおよび血液代謝産物データの少なくとも一方は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用され、腸内菌叢の細菌組成データは、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される。

Description

COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム
 本発明は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法等に関する。
 従来、COVID-19に罹患した際の重症化リスク等に関する因子について、以下の報告がある。すなわち、COVID-19に罹患した際の重症化の予測には、罹患者の体重、性別、年齢が重要である(非特許文献1)。血中サイトカイン(CCL-17)によって重症化を予測することができる(非特許文献2)。尿中の肝型脂肪酸結合タンパク質によって、COVID-19の重症化リスクを検査することができる(特許文献1)。COVID-19の発症、重症度によって、腸内細菌叢が変化する(非特許文献3)。
特許第6933834号
Yamada, Gen, et al. "Predicting respiratory failure for COVID-19 patients in Japan: a simple clinical score for evaluating the need for hospitalisation." Epidemiology & Infection 149 (2021). Sugiyama, Masaya, et al. "Serum CCL17 level becomes a predictive marker to distinguish between mild/moderate and severe/critical disease in patients with COVID-19." Gene 766 (2021): 145145. Gut. 2021 Apr;70(4):698-706. doi: 10.1136/gutjnl-2020-323020. Epub 2021 Jan 11.
 本件発明者らは、COVID-19に罹患した際の重症化リスク等を予測する従来の方法は、予測精度が十分ではないことがあるとの問題点があることを見出した。
 本発明は、一側面では、このような事情を鑑みてなされたものであり、その目的は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを、従来ない新しい手法に基づいて高精度で検査する方法等を提供することである。
 本発明は、上述した課題を解決するために、以下の構成を採用する。
 すなわち、本発明の一側面に係る検査方法は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法であって、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを測定する測定ステップを含み、腸内菌叢の遺伝子組成データおよび血液代謝産物データの少なくとも一方は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用され、腸内菌叢の細菌組成データは、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される。
 本件発明者らは、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データおよび血液代謝産物データの少なくとも一方を用いることにより、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを従来よりも高精度に予測できることを確認した。また、本件発明者らは、被験者から採取された生体試料中の腸内菌叢の細菌組成データを用いることにより、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを高精度に予測できることを確認した。
 そこで、当該構成では、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを測定する測定ステップを含む。そして、本発明の一側面に係る検査方法において、前記測定ステップにて測定された腸内菌叢の遺伝子組成データおよび血液代謝産物データの少なくとも一方は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される。また、本発明の一側面に係る検査方法において、前記測定ステップにて測定された腸内菌叢の細菌組成データは、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される。当該構成によれば、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを従来よりも高精度に予測することができる。
 本発明の形態は、上記検査方法の形態に限られなくてもよい。上記形態に係る検査方法の別の態様として、本発明の一側面は、以上の各構成の全部又はその一部を実現する検査装置であってもよい。また、本発明の一側面は、以上の各構成の全部又はその一部を実現する情報処理方法であってもよいし、プログラムであってもよいし、このようなプログラムを記憶した、コンピュータその他装置、機械等が読み取り可能な記憶媒体であってもよい。ここで、コンピュータ等が読み取り可能な記憶媒体とは、プログラム等の情報を、電気的、磁気的、光学的、機械的、又は、化学的作用によって蓄積する媒体である。
 例えば、本発明の一側面に係る検査装置は、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを取得する取得部と、取得された前記腸内菌叢の遺伝子組成データおよび前記血液代謝産物データの少なくとも一方を、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測し、または、取得された前記腸内菌叢の細菌組成データを、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する予測部と、を備える。
 例えば、本発明の一側面に係る検査プログラムは、コンピュータに、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを取得する取得ステップと、取得された前記腸内菌叢の遺伝子組成データおよび前記血液代謝産物データの少なくとも一方を、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測し、または、取得された前記腸内菌叢の細菌組成データを、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する予測ステップと、を実行させるための、プログラムである。
 本発明によれば、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを、高精度で検査する方法等を提供することができる。
図1は、本実施形態に係る検査方法の概要を説明するフロー図である。 図2は、各アウトカムに対して遺伝子組成データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図3は、各アウトカムに対して遺伝子組成データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図4は、各アウトカムに対して血液代謝産物データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図5は、各アウトカムに対して血液代謝産物データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図6は、各アウトカムに対して細菌組成データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図7は、各アウトカムに対して細菌組成データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図8は、各アウトカムに対して糞便代謝産物濃度データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図9は、各アウトカムに対して糞便代謝産物濃度データから抽出された特徴量、および、該特徴量を用いて構築された予測モデルの予測のROC-AUCスコア等を示す棒グラフである。 図10は、本実施形態に係るモデル生成装置および検査装置を含む分析システムの全体概要を説明する図である。 図11は、実施の形態に係るモデル生成装置のハードウェア構成の一例を模式的に例示する。 図12は、実施の形態に係る検査装置のハードウェア構成の一例を模式的に例示する。 図13は、実施の形態に係るモデル生成装置のソフトウェア構成の一例を模式的に例示する。 図14は、実施の形態に係る検査装置のソフトウェア構成の一例を模式的に例示する。 図15は、実施の形態に係るモデル生成装置の処理手順の一例を示すフローチャートである。 図16は、実施の形態に係る検査装置の処理手順の一例を示すフローチャートである。
 以下、本発明の一側面に係る実施の形態(以下、「本実施形態」とも表記する)を、図面に基づいて説明する。ただし、以下で説明する本実施形態は、あらゆる点において本発明の例示に過ぎない。本発明の範囲を逸脱することなく種々の改良や変形を行うことができることは言うまでもない。つまり、本発明の実施にあたって、実施形態に応じた具体的構成が適宜採用されてもよい。
 §1 全体概要
 図1は、本実施形態に係る検査方法の概要を説明するフロー図である。本件発明者らは、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを用いることにより、COVID-19に罹患した際の重症化リスク等を従来よりも高精度に予測できることを確認した。すなわち、機械学習アルゴリズムとしてBoruta法を用いて腸内菌叢の遺伝子組成データから特徴量の抽出を行ったところ、いくつかの遺伝子が特徴量遺伝子として抽出され、これを用いたランダムフォレストによる予測モデルで重症化リスク等を高精度に予測できることが確認された。同様に、Boruta法を用いて血液代謝産物データから抽出した血液代謝産物を用いた予測モデル、および、Boruta法を用いて腸内菌叢の細菌組成データから抽出した腸内細菌を用いた予測モデルでも、重症化リスク等を高精度に予測できることが確認された。
 そこで、本実施形態に係る検査方法は、図1に例示するように、被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ(対象遺伝子組成データ221A)、対象血液代謝産物データ221B、および、腸内菌叢の細菌組成データ(対象細菌組成データ221C)の少なくとも1つを測定する測定ステップ(S11)を含む。そして、本実施形態に係る検査方法において、前記測定ステップにて測定された対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つは、COVID-19に罹患した際の重症化リスク等を検査するために使用される。具体的には、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される。例えば、対象遺伝子組成データ221Aおよび対象血液代謝産物データ221Bの少なくとも一方を用いて、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する。または、対象細菌組成データ221Cを用いて、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する。
 特に、被験者から採取する生体試料を糞便とし、糞便中の対象遺伝子組成データ221Aまたは対象細菌組成データ221Cを用いる場合、非侵襲的な方法により、非医療従事者によっても採取可能な糞便から測定した対象遺伝子組成データ221Aまたは対象細菌組成データ221Cを用いて、重症化リスク等を検査できる。例えば、被験者自身が採取した糞便中の対象遺伝子組成データ221Aまたは対象細菌組成データ221Cを用いて、被験者について、重症化リスク等を検査することができる。また、一般的なコンピュータに、後述する検査プログラム82を読み込ませることにより、被験者から採取された生体試料中の、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つから、重症化リスク等を検査することができる。
 §2 データ準備の概要
 以下では先ず、COVID-19に罹患した際の重症化リスク等を高精度に予測できる予測モデルMを構築するために本件発明者らが行なった作業の概要を説明する。
 2.1. 対象
 先ず、本件発明者らは、COVID-19非感染者(N=112)およびCOVID-19感染者(N=112、PCR検査の結果が陽性であって、国立国際医療センターに入院した者)の糞便サンプル、および、これらの者の背景情報を取得した。背景情報は、例えば、身長・体重・年齢・性別・既往症の有無・飲酒習慣の有無・喫煙習慣の有無などである。以下に、COVID-19非感染者およびCOVID-19感染者の詳細を説明する。
 2.1.1. <COVID-19感染者>
 COVID-19感染者は、以下の条件を全て満たす者とした。すなわち、第1に、研究参加に関して文書による検体の二次利用の同意が得られたとの条件。第2に、上記同意の取得時の年齢が16歳以上、120歳以下の男性または女性であるとの条件。第3に、日本人(ここでは、「日本生まれ、かつ、日本国籍の者)を指す)であるとの条件。第4に、COVID-19感染と診断され、国立研究開発法人国立国際医療研究センター(日本国)で入院加療した患者であるとの条件。第4の条件を満たす者は、特に、2020年1月31日から2020年9月30日までの期間にCOVID-19と診断され、国立研究開発法人国立国際医療研究センターで入院加療し、糞便・血液検体を提出した者である。具体的には、「新型コロナウイルス感染症(COVID-19)の前向き観察研究(NCGM-G-003472-03)主任研究者:杉山温人」で2020年1月31日から2020年9月30日までの期間に検体解析の同意が得られた者の臨床情報・糞便・血液検体を取り扱った。上記期間内で、入院加療を行った患者の入院時の臨床情報と検体とを解析に使用した。なお、新型コロナウイルス感染症(COVID-19)の原因ウイルスとして知られるSARSコロナウイルス2(SARS-CoV-2)の変異株であるα(アルファ)株は、イギリスで2020年11月、新型コロナウイルスの流行拡大中の9月に採取されたサンプルから検出された。そのため、2020年1月31日から2020年9月30日までの期間において、日本でCOVID-19と診断された者は、α株等の特定の変異株に感染した者ではないと考えられる。ただし、本発明の一側面に係る検査方法、検査装置等は、α株、β(ベータ)株、γ(ガンマ)株、および、δ(デルタ)株等の特定の変異株に感染した場合の各種のリスクの分析にも適用可能である。
 2.1.2. <COVID-19非感染者>
 COVID-19非感染者は、COVID-19感染者と同様の第1、第2、および、第3の条件を全て満たし、かつ、第4の条件として、以下の条件を満たす者とした。すなわち、明らかにCOVID-19感染が流行する以前の2014年9月1日から2019年6月31日までに糞便かつ血漿・血清検体を提出した者であるとの条件を満たす者を、第4の条件を満たす者とした。具体的には、2014年9月1日から2019年6月31日までの間に国際医療研究センターで「日本人の大規模データベース構築から明らかにする腸内微生物叢と病気・薬剤使用との関連(NCGM-G-001690-05)実施期間の主任研究者:永田尚義」で二次利用の同意が得られた臨床情報・糞便検体をコントロールコホートして取り扱った。
 2.2. 研究デザイン
 以下の研究デザインに則り、COVID-19感染者(112例)、および、COVID-19非感染者(112例)を解析対象とした。
 2.2.1. マッチングケースコントロール研究とコホート研究
 ケースは、COVID-19感染者(112例)とした。コントロールは、COVID-19非感染者から、ケースであるCOVID-19感染者と患者背景を揃えるために、背景因子(年齢、性別、BMI(Body Mass Index)、疾患、薬剤内服歴)でマッチしたコントロール症例を抽出した。マッチングの比率は、「ケース:コントロール」を「1:1」とした。その結果、COVID-19非感染者(112例)を抽出し、解析対象とした。
 ケースとコントロールとの間で、明瞭に異なるオミックス情報(マーカー)を同定する。さらに、前述したCOVID-19感染者(112例)を追跡し(コホート研究)、アウトカムを調査した。アウトカムと関連するオミックス情報を明らかにし、そのオミックス情報からアウトカムの予測モデルを構築することを目指した。
 2.2.2. アウトカムの評価項目
 臨床情報は入院時と退院後まで継続的にカルテに記載された。そのうち、年齢、性別、既往歴、基礎疾患、薬剤内服歴、酸素飽和度、呼吸数、酸素マスク有無と酸素投与量、呼吸器装着の有無、LDH、CRP、血小板数、凝固・線溶検査(PT(プロトロンビン時間 国際標準比)、APTT(Activated Partial Thromboplastin Time:活性化部分トロンボプラスチン時間)、D-ダイマー等)、画像検査、血栓症の有無、抗凝固療法の有無と内容、出血性有害事象の有無、消化管症状、転帰を抽出した。
 アウトカムは、COVID-19発症後に認めた、下痢、肝障害、腎障害、凝固線溶系異常、肺炎、呼吸器重症化、予後とした。下痢は、便形状と便の回数により判断し、肝障害、腎障害、凝固異常は、採血のデータで判断した。肺炎は、CT所見で放射線科の読影による肺炎と診断したものとした。呼吸器「重症化」の定義は、厚生労働省「新型コロナウイルス感染症診療の手引き2021COVID-19第4.2版(https://www.mhlw.go.jp/content/000742297.pdf)」を参考に、以下のいずれかを満たす者とした。すなわち、(1)室内気でSpO2≦93%のため酸素投与が必要な肺炎患者、(2)人工呼吸器管理を必要とする呼吸不全患者、(3)ショック、多臓器不全のために集中治療管理を必要とした患者のいずれかを満たす者を、呼吸器「重症化」患者とした。
 2.2.3. オミックス情報
(1)糞便のメタゲノム解析による微生物群集データ(腸内細菌組成データ(腸内菌叢の細菌組成データ)と細菌の遺伝子組成データ)。なお、その同定方法は後述する。
(2)糞便および血漿・血清の代謝産物群集データ。なお、その同定方法は後述する。
 2.2.4. 微生物群集データの同定方法
 2.2.4.1. 糞便DNAの抽出
 糞便DNAは、参考文献2(Takeuchi T. et al., Acetate differentially regulates IgA reactivity to commensal bacteria. Nature. 2021 Jul;595(7868):560-564. doi: 10.1038/s41586-021-03727-5. Epub 2021 Jul 14. PMID: 34262176.)に記載された方法に僅かな変更を加えて調製した。具体的には、糞便ペレットをTE20緩衝液(10 mmol/l Tris-HCl (pH 8.0)および20 mmol/l EDTA (pH 8.0))に懸濁し、リゾチーム(Wako, 10mg per sample)、アクモペプチダーゼ(Wako, 1,333 units per sample)、ドデシル硫酸ナトリウム(1%)、およびプロテイナーゼK(Merck, 0.67 mg per sample)を添加して溶菌した。次いで、フェノール/クロロホルム/イソアミルアルコール(Wako)および2-プロパノール(Wako)を用いてDNAを回収した。更に、回収したDNAをRNase (Nippon Gene, 20 μg/ml)で処理することにより精製し、糞便DNAとして使用した。
 2.2.4.2. ライブラリ調製、シーケンス、データ処理
 抽出した糞便DNAは超音波(ultra sonicator ME220: Covaris)によって物理的に断片化後、AmpureXP(ベックマン・コールター)にて精製し、ThruPLEX DNA-Seq キット(タカラ)を用いてメタゲノムショットガンライブラリー(インサートサイズ350 bp)として調製した。調製済のライブラリはqPCRによって定量後、NovaSeq6000プラットフォーム(イルミナ)で150base paired endモードを使用し、ショットガンシークエンスを行った。
 得られたメタゲノムデータからクオリティ値(meanQV≦20)が低いリード、および、duplicateリードを除去した。さらに、各リードから5'末端でQV≦20となった精度の低い塩基のトリム、リードの半分以上がQV≦20となったリード、リード長が50bp未満のリードの除去を行った。さらに、メタゲノムリードをminimap2(version 2.13-r850)によってヒトゲノム(HG38)とphiXバクテリオファージゲノムにマッピングし、相同性が95%以上でヒットしたリードを除外することでヒトゲノム配列およびPhiX配列を除去した。
 2.2.4.3. 菌組成の算出手法
 mOTUs(v2.6.1)を用いて、得られた高品質のメタゲノムリードに含まれる菌種の相対存在比の算出を行った。その後、R software(version 4.1.1)上でMaAsLin2 package(version 1.6.0)を用いた解析を行い、算出した存在比データから、COVID-19患者サンプル、あるいは重症度の高い患者サンプルで存在比が有意に高いmOTU(菌種)を検出した。
 2.2.4.4. 遺伝子組成算出手法
 フィルタリングしたメタゲノムショートリードに対し、MEGAHIT(v1.2.4-beta)を用いてアセンブリし、コンティグを構築した。構築したコンティグからseqkit(version: 0.8.0)によって500 bp未満の配列を除き、prodigal(v2.6.3)によって遺伝子およびタンパク質配列をコールした。各サンプルから得られた遺伝子配列とタンパク質配列とを統合し、seqkitによって100 bp未満のリードを除去後に、CD-HIT(version 4.6)を用いたクラスタリング(identity 95%, coverage 90%)を行った。その後、Integrated Gene Catalogue(Li, Junhua, et al. (2014))とのクラスタリングも行い、遺伝子およびタンパク質のカタログを作成した。
 作成したタンパク質カタログをdiamond(version 2.0.11)によってKEGG database (2019/10/7時点)にアノテーションした。また、Bowtie2(version 2.3.2)によって遺伝子カタログをショートリードにマッピングした。マッピングしたデータとKEGGへのアノテーション結果とを組み合わせ、機能組成データを算出した。
 2.2.5. 代謝産物群集データの同定方法
 2.2.5.1. 糞便の前処理
 糞便のアリコート(200 mg)を1,200 μlのメタノールに添加して激しく撹拌した後に、100 μmフィルターで濾過し、食物残渣を除去した。得られたろ液を遠心分離(15,000 × g、10分間、4°C)して、上清(糞便メタノール抽出物)を代謝産物分析に使用した。
 2.2.5.2. 糞便および血漿サンプルの親水性代謝産物の抽出および測定 糞便および血漿サンプルは、参考文献1(Sato K. et al., Obesity-Related Gut Microbiota Aggravates Alveolar Bone Destruction in Experimental Periodontitis through Elevation of Uric Acid. mBio. 2021 Jun 29;12(3):e0077121. doi: 10.1128/mBio.00771-21. Epub 2021 Jun 1. PMID: 34061595; PMCID: PMC8262938.)に記載された方法に僅かな変更を加えて調製した。具体的には、血漿サンプルについては、血漿のアリコート10μlを、メタノール:水:クロロホルムを容量比5:5:2で含む混液を使用して抽出処理し、5 kDaカットオフフィルターカラム(Ultrafree MC-PLHCC, Human Metabolome Technologies)でろ過し、得られた抽出液を使用した。また、糞便サンプルについては、糞便メタノール抽出物のアリコート25μlを、メタノール:水:クロロホルムを容量比5:5:2で含む混液を使用して抽出処理することにより得られた抽出液をろ過せずに使用した。このように調製した抽出液をGC/MS/MS分析用とLC-MS/MS分析用に分けた。
 GC/MS/MS分析では、抽出液200μlをエバポレーターで乾燥させ、凍結乾燥機で凍結乾燥した。凍結乾燥した抽出物を、分析前に、メトキシアミン塩酸塩(Sigma-Aldrich)およびN-メチル-N-トリメチルシリル-トリフルオロアセトアミド(MSTFA, GL Science)で誘導体化した。GC/MS/MS分析のプログラムは、参考文献1に記載のものを使用した。
 LC-MS/MS分析では、抽出液100μlをエバポレーターで乾燥させ、凍結乾燥機で凍結乾燥した。分析前に、凍結乾燥した抽出物に水500μlを加えて再構成した。LC-MS/MS分析は、Shimadzu LCMS-8040 triple quadrupole mass spectrometer (Shimadzu)を用いて実施した。クロマトグラフィー分離では、Discovery HS F5-3 (2.1 mm.D. x 150 mmL., 3 μm) (Supelco)を使用し、流速0.35 ml/min、カラム温度40 °Cに設定した。溶媒A(ギ酸0.1%を含む水)および溶媒B(ギ酸0.1%を含むアセトニトリル)を使用して、次のグラジエントで分析した:0.0分 (0% B); 0.0-3.5分 (0%B→25%B); 3.5-7.5分(25%B→35%B); 7.5-10.3分(35%B→95%B); 10.3-13.7分(95%B); 13.7-14.8分(100%B); 14.8-17 分(0%B)。インジェクション量は1μLに設定した。多重反応モニタリング(MRM)を使用して、親水性代謝産物をモニタリングした。パラメータは、次の通りに設定した:噴霧ガス流量3 l/分、加熱ガス流量10 l/分、インターフェイス温度300 °C、DL温度250 °C、ヒートブロック温度400 °C、乾燥ガス流量15 l/分。そして、LabSolutions Insight (Shimadzu)で、GC/MS/MSデータ及びLC-MS/MSデータを処理して、濃度を算出した。
 単鎖脂肪酸(SCFA)の抽出および測定は、参考文献2(Takeuchi T. et al., Acetate differentially regulates IgA reactivity to commensal bacteria. Nature. 2021 Jul;595(7868):560-564. doi: 10.1038/s41586-021-03727-5. Epub 2021 Jul 14. PMID: 34262176.)に記載された方法に幾つか変更を加えて行った。具体的には、血漿サンプルについては、血漿のアリコート90μlを、内部標準(2 mM [1,2-13C2]酢酸), 2 mM [2H7]酪酸、および2 mM クロトン酸)を含む超純水10μlに添加することにより調製した。糞便サンプルについては、糞便メタノール抽出物のアリコート25μlを、前記内部標準を含む超純水10μlに添加し、40℃で遠心濃縮し、超純水100μlを加えて再構成することにより調製した。各サンプルに、塩酸50μlおよびジエチルエーテル200μlを加えてよく撹拌した。次いで、遠心分離(3,000 × g、10分)して相分離させた後に、有機相80μlをガラスバイアルに移し、更にN-tert-ブチルジメチルシリル-N-トリフルオロアセトアミド(MTBSTFA, Sigma-Aldrich)を添加して誘導体化した。インジェクション前に、ガラスバイアルを80℃で20分間インキュベートし、48時間静置した。GC/MS/MS分析のプログラムは、参考文献2に記載のものを使用した。LabSolutions Insight (Shimadzu)で、GC/MS/MSデータを処理して、濃度を算出した。
 §3 予測モデル構築の概要
 本件発明者らは、上述の方法でCOVID-19感染者(112例)の各々の入院時の生体試料から取得した訓練オミックスデータ31を説明変数とし、COVID-19感染者(112例)の各々について観察された「重症化の有無」等の各アウトカムを目的変数とする、予測モデルMを機械学習により構築した。
 訓練オミックスデータ31は、腸内菌叢の遺伝子組成データ(訓練遺伝子組成データ31A)、血液代謝産物データ(訓練血液代謝産物データ31B)、および、腸内菌叢の細菌組成データ(訓練細菌組成データ31C)の少なくとも1つを含む。本実施形態では、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、および、訓練細菌組成データ31Cの少なくとも1つを用いて構築した予測モデルMの予測精度の検証のために、訓練オミックスデータ31は、さらに、糞便代謝産物濃度データ(訓練糞便代謝産物濃度データ31D)を含んでもよい。なお以下では、腸内菌叢の訓練遺伝子組成データは、「遺伝子組成データ」と、腸内菌叢の細菌組成データは、「細菌組成データ」と、略記することがある。
 また、アウトカムは、「COVID-19に罹患した際の症状の発現の有無」、「重症化の有無」、「合併症の有無」、「長期入院の有無」を含む。本実施形態において、COVID-19に罹患した際の「症状」は、呼吸器症状、肺炎、下痢の少なくとも1つである。また、COVID-19に罹患した際の「(コロナ感染症の)重症化」とは、上述の呼吸器「重症化」を示す。COVID-19に罹患した際の「合併症」は、肝障害、腎障害、血栓症の少なくとも1つである。「血栓症」は、血小板が20×104/μl未満に減少すること、D-ダイマーが1.0mg/ml以上に上昇すること、フィブリノゲンが400mg/dlよりも大きくなることの少なくとも1つである。「長期入院」は、入院期間が11日以上であることを示す。入院期間とは、「入院」から「退院」までの期間を指し、「入院」とは、現在の病気の状態などから、今後病状が悪化する可能性がある、あるいは、自宅では病状を経過観察することが困難と判断される場合に、医療機関において24時間管理体制のもとで管理治療をおこなう状態を指す。「入院」は、通常は24時間の管理治療が始まるタイミングで用いる語である。また、「退院」とは、現在の病気の状態が改善した場合、あるいは自宅で療養が可能と判断される場合において、医療機関の24時間管理体制が必要なくなった状態を指す。「退院」は、通常は自宅療養できるタイミングで用いる語である。
 予測モデルMの構築に際しては、先ず、訓練オミックスデータ31から、「重症化の有無」等の各アウトカムの判別に重要な説明変数を選択し、つまり、特徴量選択を実施した。本実施形態においては、Borutaアルゴリズム(version 0.3)を用いて、訓練オミックスデータ31から、判別に重要な説明変数(特徴量)を抽出した。すなわち、上述の「重症化の有無」を含むアウトカムを目的変数とした場合に重要となる特徴量(説明変数)を、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、訓練細菌組成データ31C、および、訓練糞便代謝産物濃度データ31Dの各々から、抽出した。
 次に、抽出した特徴量を説明変数とし、「重症化の有無」等の各アウトカムを目的変数とする機械学習を行なうことによって、予測モデルMを構築した。本実施形態においては、ランダムフォレスト(scikit-learn version 0.24.1)により、予測モデルMを構築した。すなわち、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、訓練細菌組成データ31C、および、訓練糞便代謝産物濃度データ31Dの各々から抽出した特徴量を用いて、予測モデルMA、MB、MC、および、MDを構築した。
 なお、上述の通り、アウトカムは、「COVID-19に罹患した際の症状の発現の有無」、「重症化の有無」、「合併症の有無」、「長期入院の有無」を含む。また、「症状」は、「呼吸器症状」、「肺炎」、「下痢」を含む。さらに、「合併症」は、「肝障害」、「腎障害」、「血栓症」を含み、「血栓症」は、「血小板が20×104/μl未満に減少すること」、「D-ダイマーが1.0mg/ml以上に上昇すること」、「フィブリノゲンが400mg/dlよりも大きくなること」を含む。それゆえ、予測モデルMA、MB、MC、および、MDの各々は、上述の各アウトカムの有無を予測する予測モデルを含む。
 例えば、訓練遺伝子組成データ31Aから抽出した特徴量を用いて構築される予測モデルMAは、「重症化の有無」を予測する予測モデルMA1、「呼吸器症状の発現の有無」を予測する予測モデルMA2、「肺炎の発現の有無」を予測する予測モデルMA3を含む。予測モデルMAは、「下痢の発現の有無」を予測する予測モデルMA4、「肝障害の発症の有無」を予測する予測モデルMA5、「腎障害の発症の有無」を予測する予測モデルMA6、「血小板の20×104/μl未満への減少の有無」を予測する予測モデルMA7を含む。予測モデルMAは、「D-dimerの1.0mg/ml以上への上昇の有無」を予測する予測モデルMA8、「400mg/dlよりも大きなフィブリノゲンの有無」を予測する予測モデルMA9、「長期入院の有無」を予測する予測モデルMA11を含む。予測モデルMB、MC、および、MDの各々についても同様である。
 さらに、構築した予測モデルMの予測精度を、10分割交差検証を行って得られたROC曲線(receiver operating characteristic curve)の曲線下面積(ROC-AUC)により評価した。予測モデルMの精度(ROC-AUC)を以下の表1に示す。
(表1)
 表1に示される通り、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、および、訓練細菌組成データ31Cの各々を用いて構築した予測モデルM(MA、MB、MC)により、COVID-19に罹患した際の重症化リスク等を従来よりも高精度に予測できることが確認された。
 例えば、非特許文献1に記載の「患者背景(体重、性別、年齢)」を用いた「重症化」の予測のROC-AUCスコアは「0.71」である。また、訓練糞便代謝産物濃度データ31Dを用いて構築した予測モデルMD(表1の「糞便代謝産物」)の、「重症化」の予測のROC-AUCスコアは「0.72」である。
 これに対して、訓練遺伝子組成データ31Aを用いて構築した予測モデルMA(表1の「腸内細菌遺伝子組成」)の、「重症化」の予測のROC-AUCスコアは「0.87」である。また、訓練血液代謝産物データ31Bを用いて構築した予測モデルMB(表1の「血液代謝産物」)の、「重症化」の予測のROC-AUCスコアは「0.92」である。すなわち、訓練遺伝子組成データ31Aおよび訓練血液代謝産物データ31Bの少なくとも一方を用いて構築した予測モデルM(予測モデルMAおよびMBの少なくとも一方)により、「重症化」を、従来よりも高精度に予測できることが確認された。また、訓練遺伝子組成データ31Aを用いて構築した予測モデルMA、および、訓練血液代謝産物データ31Bを用いて構築した予測モデルMBは、いずれも、「合併症の有無」、「長期入院の有無」等についても、従来よりも高精度に予測できることが確認された。
 また、表1に示すように、訓練細菌組成データ31Cを用いて構築した予測モデルMC(表1の「腸内細菌種組成」)の、「肝障害」、「腎障害」等の予測のROC-AUCスコアは、訓練糞便代謝産物濃度データ31Dを用いて構築した予測モデルMDの、「肝障害」、「腎障害」等の予測のROC-AUCスコアに比べて高い。それゆえ、訓練細菌組成データ31Cを用いて構築した予測モデルMCによって、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクなどを高精度に予測できることが確認された。
 予測モデルMA、MB、MC、および、MDの各々の構築に際して用いられた特徴量、および、予測モデルMA、MB、MC、および、MDの各々の予測のROC-AUCスコアを図2~図9に示す。なお、図2~図9の各々に示す棒グラフにおいて、棒の長さは、各特徴量の各アウトカムの予測への寄与度の大きさを示している。また、棒を濃い灰色で示した特徴量は、各アウトカムの予測への寄与度がポジティブであることを、棒を薄い灰色で示した特徴量は、各アウトカムの予測への寄与度がネガティブであることを、示している。
<腸内菌叢の遺伝子組成データを用いる場合について>
 図2および図3は、各アウトカムに対して訓練遺伝子組成データ31Aから抽出された特徴量、および、抽出された特徴量を用いて構築された予測モデルMAの予測のROC-AUCスコア等を示している。
 図2および図3に示すように、「重症化の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(1)~(20)に示す遺伝子が、特徴量として抽出された。すなわち、(1) K11636 (yxdM)、(2) K03932 (lpqC)、(3) K19064 (lysDH)、(4) K01455 (E3.5.1.49)、(5) K03753 (mobB)、(6) K03976 (ybaK, ebsC)、(7) K18123 (HOGA1)、(8) K03572 (mutL)、(9) K14654 (RIB7, arfC)、(10) K07991 (flaK-A, flaK)、(11) K23509 (ytfT, yjfF)、(12) K03191 (ureI)、(13) K00303 (soxB)、(14) K03685 (rnc, DROSHA, RNT1)、(15) K13891 (gsiD)、(16) K03394 (cobI-cbiL)、(17) K01581 (E4.1.1.17, ODC1, speC, speF)、(18) K03187 (ureE)、(19) K09817 (znuC)、(20) K04078 (groES, HSPE1)が、特徴量として抽出された。
 また、上述の(1)~(20)に示す遺伝子(特徴量)を用いて構築された予測モデルMA1の、「重症化の有無」の予測に係るROC-AUCスコアは、「0.87」であった。
 「呼吸器症状の発現の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(21)~(40)に示す遺伝子が、特徴量として抽出された。すなわち、(21) K01200 (pulA)、(22) K07650 (cssS)、(23) K02860 (rimM)、(24) K03769 (ppiC)、(25) K07770 (cssR)、(26) K01222 (E3.2.1.86A, celF)、(27) K07025 (K07025)、(28) K12573 (rnr, vacB)、(29) K19116 (cas5h)、(30) K02028 (ABC.PA.A)、(31) K01114 (plc)、(32) K01924 (murC)、(33) K01857 (pcaB)、(34) K01958 (PC, pyc)、(35) K03719 (lrp)、(36) K03402 (argR, ahrC)、(37) K02834 (rbfA)、(38) K01621 (xfp, xpk)、(39) K03650 (mnmE, trmE, MSS1)、(40) K07099 (K07099)が、特徴量として抽出された。
 また、上述の(21)~(40)に示す遺伝子(特徴量)を用いて構築された予測モデルMA2の、「呼吸器症状の発現の有無」の予測に係るROC-AUCスコアは、「0.78」であった。
 「肺炎の発現の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(41)~(60)に示す遺伝子が、特徴量として抽出された。すなわち、(41) K07171 (mazF, ndoA, chpA)、(42) K21469 (pbp4b)、(43) K06013 (STE24)、(44) K07459 (ybjD)、(45) K02107 (ATPVG, ahaH, atpH)、(46) K01463 (bshB1)、(47) K04652 (hypB)、(48) K07775 (resD)、(49) K10231 (kojP)、(50) K10540 (mglB)、(51) K19334 (tabA)、(52) K00172 (porG)、(53) K23265 (purQ)、(54) K01153 (hsdR)、(55) K01455 (E3.5.1.49)、(56) K17835 (griH)、(57) K15899 (pseF)、(58) K14623 (dinD)、(59) K09767 (yajQ)、(60) K13620 (wcaD)が、特徴量として抽出された。
 また、上述の(41)~(60)に示す遺伝子(特徴量)を用いて構築された予測モデルMA3の、「肺炎の発現の有無」の予測に係るROC-AUCスコアは、「0.87」であった。
 「下痢の発現の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(61)~(80)に示す遺伝子が、特徴量として抽出された。すなわち、(61) K07309 (ynfE)、(62) K03092 (rpoN)、(63) K03395 (aac3-I)、(64) K19883 (aacA-aphD)、(65) K03489 (yydK)、(66) K01004 (pcs)、(67) K13075 (ahlD, aiiA, attM, blcC)、(68) K11928 (putP)、(69) K21469 (pbp4b)、(70) K06311 (yndE)、(71) K07192 (FLOT)、(72) K07668 (vicR)、(73) K22579 (rph)、(74) K01506 (davA)、(75) K15524 (mngB)、(76) K11050 (ABC-2.CYL.A, cylA)、(77) K03587 (ftsI)、(78) K02745 (PTS-Aga-EIIB, agaV)、(79) K11178 (yagS)、(80) K03325 (ACR3, arsB)が、特徴量として抽出された。
 また、上述の(61)~(80)に示す遺伝子(特徴量)を用いて構築された予測モデルMA4の、「下痢の発現の有無」の予測に係るROC-AUCスコアは、「0.81」であった。
 「肝障害の発症の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(81)~(100)に示す遺伝子が、特徴量として抽出された。すなわち、(81) K01546 (kdpA)、(82) K09858 (K09858)、(83) K21636 (nrdD)、(84) K06895 (lysE, argO)、(85) K15868 (baiB)、(86) K05810 (yfiH)、(87) K00613 (GATM)、(88) K17363 (urdA)、(89) K03299 (TC.GNTP)、(90) K01547 (kdpB)、(91) K13821 (putA)、(92) K18889 (mdlA, smdA)、(93) K11194 (PTS-Fru1-EIIA, levD)、(94) K18350 (vanSC, vanSE, vanSG)、(95) K02558 (mpl)、(96) K03503 (umuD)、(97) K19169 (dndB)、(98) K00183 (K00183)、(99) K02445 (glpT)、(100) K05966 (citG)が、特徴量として抽出された。
 また、上述の(81)~(100)に示す遺伝子(特徴量)を用いて構築された予測モデルMA5の、「肝障害の発症の有無」の予測に係るROC-AUCスコアは、「0.84」であった。
 「腎障害の発症の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(101)~(120)に示す遺伝子が、特徴量として抽出された。すなわち、(101) K02919 (RP-L36, MRPL36, rpmJ)、(102) K13668 (pimB)、(103) K19221 (cobA, btuR)、(104) K05873 (cyaB)、(105) K18214 (tetP_A, tet40)、(106) K18906 (mgrA)、(107) K01227 (E3.2.1.96)、(108) K02057 (ABC.SS.P)、(109) K20453 (dmdB)、(110) K13730 (inlA)、(111) K06941 (rlmN)、(112) K05814 (ugpA)、(113) K12276 (mshE)、(114) K08258 (sspB2, sspP, scpA)、(115) K00973 (E2.7.7.24, rfbA, rffH)、(116) K02110 (ATPF0C, atpE)、(117) K07038 (K07038)、(118) K07010 (K07010)、(119) K06904 (K06904)、(120) K11690 (dctM)が、特徴量として抽出された。
 また、上述の(101)~(120)に示す遺伝子(特徴量)を用いて構築された予測モデルMA6の、「腎障害の発症の有無」の予測に係るROC-AUCスコアは、「0.88」であった。
 「血小板の20×104/μl未満への減少の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(121)~(140)に示す遺伝子が、特徴量として抽出された。すなわち、(121) K02124 (ATPVK, ntpK, atpK)、(122) K02118 (ATPVB, ntpB, atpB)、(123) K00891 (E2.7.1.71, aroK, aroL)、(124) K07069 (K07069)、(125) K03319 (TC.DASS)、(126) K02123 (ATPVI, ntpI, atpI)、(127) K00784 (rnz)、(128) K18349 (vanRC, vanRE, vanRG)、(129) K22960 (lpdD)、(130) K01785 (galM, GALM)、(131) K01712 (hutU, UROC1)、(132) K02073 (metQ)、(133) K03800 (lplA, lplJ, lipL1)、(134) K07469 (mop)、(135) K21613 (yxeP)、(136) K00782 (lldG)、(137) K00640 (cysE)、(138) K17234 (araN)、(139) K00823 (puuE)、(140) K09136 (ycaO)が、特徴量として抽出された。
 また、上述の(121)~(140)に示す遺伝子(特徴量)を用いて構築された予測モデルMA7の、「血小板の20×104/μl未満への減少の有無」の予測に係るROC-AUCスコアは、「0.79」であった。
 「D-dimerの1.0mg/ml以上への上昇の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(141)~(160)に示す遺伝子が、特徴量として抽出された。すなわち、(141) K03424 (tatD)、(142) K21465 (pbpA)、(143) K02744 (PTS-Aga-EIIA, agaF)、(144) K03781 (katE, CAT, catB, srpA)、(145) K06218 (relE, stbE)、(146) K01590 (hdc, HDC)、(147) K05305 (FUK)、(148) K09811 (ftsX)、(149) K02193 (ccmA)、(150) K09711 (K09711)、(151) K23375 (sbnC)、(152) K16389 (amphN, nysN, fscP, pimG)、(153) K01733 (thrC)、(154) K15599 (thiX)、(155) K03498 (trkH, trkG, ktrB)、(156) K10542 (mglA)、(157) K01216 (E3.2.1.73)、(158) K08303 (K08303)、(159) K08641 (vanX)、(160) K03274 (gmhD, rfaD)が、特徴量として抽出された。
 また、上述の(141)~(160)に示す遺伝子(特徴量)を用いて構築された予測モデルMA8の、「D-dimerの1.0mg/ml以上への上昇の有無」の予測に係るROC-AUCスコアは、「0.85」であった。
 「400mg/dlよりも大きなフィブリノゲンの有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(161)~(180)に示す遺伝子が、特徴量として抽出された。すなわち、(161) K02076 (zurR, zur)、(162) K19115 (csh2)、(163) K12552 (pbpA)、(164) K19334 (tabA)、(165) K17329 (dasA)、(166) K21744 (tipA)、(167) K03722 (dinG)、(168) K23265 (purQ)、(169) K19116 (cas5h)、(170) K00145 (argC)、(171) K09819 (ABC.MN.P)、(172) K03189 (ureG)、(173) K01579 (panD)、(174) K16302 (CNNM)、(175) K01996 (livF)、(176) K06156 (gntU)、(177) K22902 (K22902)、(178) K03524 (birA)、(179) K00803 (AGPS, agpS)、(180) K00209 (fabV, ter)が、特徴量として抽出された。
 また、上述の(161)~(180)に示す遺伝子(特徴量)を用いて構築された予測モデルMA9の、「400mg/dlよりも大きなフィブリノゲンの有無」の予測に係るROC-AUCスコアは、「0.89」であった。
 「長期入院の有無」を目的変数とした場合、訓練遺伝子組成データ31Aから、下記(201)~(220)に示す遺伝子が、特徴量として抽出された。すなわち、(201) K01463 (bshB1)、(202) K14188 (dltC)、(203) K07457 (K07457)、(204) K22958 (lpdC)、(205) K15372 (toa)、(206) K07069 (K07069)、(207) K08153 (blt)、(208) K14654 (RIB7, arfC)、(209) K06156 (gntU)、(210) K03734 (apbE)、(211) K07505 (repA)、(212) K01482 (DDAH, ddaH)、(213) K03328 (TC.PST)、(214) K01635 (lacD)、(215) K07217 (K07217)、(216) K19973 (mntA)、(217) K03753 (mobB)、(218) K23245 (apnO)、(219) K16203 (dppA1)、(220) K05770 (TSPO, BZRP)が、特徴量として抽出された。
 また、上述の(201)~(220)に示す遺伝子(特徴量)を用いて構築された予測モデルMA11の、「長期入院の有無」の予測に係るROC-AUCスコアは、「0.85」であった。
<血液代謝産物データを用いる場合について>
 図4および図5は、各アウトカムに対して訓練血液代謝産物データ31Bから抽出された特徴量、および、抽出された特徴量を用いて構築された予測モデルMBの予測のROC-AUCスコア等を示している。
 図4および図5に示すように、「重症化の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(301)~(320)に示す血液代謝産物が、特徴量として抽出された。すなわち、(301) フェニルアラニン、(302) マンノース、(303) 2-ヒドロキシ酪酸、(304) イノシトール、(305) 4-ヒドロキシフェニル乳酸、(306) 3-フェニル乳酸、(307) チロシン、(308) グリシン、(309) シュウ酸、(310) 2-アミノエタノール、(311) トレオン酸、(312) グリセロール、(313) 尿素、(314) 2-ヒドロキシイソ吉草酸、(315) L-カルニチン、(316) バリン、(317) グルクロン酸、(318) ジメチルグリシン、(319) サルコシン、(320) グルコースが、特徴量として抽出された。
 また、上述の(301)~(320)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB1の、「重症化の有無」の予測に係るROC-AUCスコアは、「0.92」であった。
 「呼吸器症状の発現の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(321)~(335)に示す血液代謝産物が、特徴量として抽出された。すなわち、(321) 3-ヒドロキシイソ酪酸、(322) プロピオン酸、(323) 2-ヒドロキシイソ酪酸、(324) イノシトール、(325) 2-ケトグルタル酸、(326) グリオキシル酸、(327) ロイシン、(328) 5-オキソプロリン、(329) フェニルアラニン、(330) バリン、(331) グルタミン酸、(332) 3-アミノイソ酪酸、(333) 2-アミノイソ酪酸、(334) ノナン酸、(335) グルクロン酸が、特徴量として抽出された。
 また、上述の(321)~(335)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB2の、「呼吸器症状の発現の有無」の予測に係るROC-AUCスコアは、「0.77」であった。
 「肺炎の発現の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(336)~(355)に示す血液代謝産物が、特徴量として抽出された。すなわち、(336) エリトルロース、(337) イノシトール、(338) マンニトール、(339) ギ酸、(340) セリン、(341) グリシン、(342) マンノース、(343) フルクトース、(344) ヒスチジン、(345) 2-ケトグルタル酸、(346) インドール-3-酢酸、(347) trans-アコニット酸、(348) チロシン、(349) リブロ―ス、(350) グルコース、(351) 2-アミノアジピン酸、(352) ガラクツロン酸、(353) ウラシル、(354) 3-ヒドロキシグルタル酸、(355) セロトニンが、特徴量として抽出された。
 また、上述の(336)~(355)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB3の、「肺炎の発現の有無」の予測に係るROC-AUCスコアは、「0.86」であった。
 「下痢の発現の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(356)~(361)に示す血液代謝産物が、特徴量として抽出された。すなわち、(356) プロリン、(357) アスパラギン、(358) クレアチン、(359) インドール-3-酢酸、(360) フェニルアラニン、(361) グリセリン酸が、特徴量として抽出された。
 また、上述の(356)~(361)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB4の、「下痢の発現の有無」の予測に係るROC-AUCスコアは、「0.65」であった。
 「肝障害の発症の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(362)~(380)に示す血液代謝産物が、特徴量として抽出された。すなわち、(362) 2-ケトグルタル酸、(363) グルタミン酸、(364) ヒスチジン、(365) バリン、(366) アスパラギン、(367) エリトルロース、(368) クエン酸、(369) イソ酪酸、(370) グルタミン、(371) マンニトール、(371) 2-アミノピメリン酸、(372) リジン、(373) インドール-3-酢酸、(374) コハク酸、(375) ロイシン、(376) 2-ケトイソカプロン酸、(377) グリシン、(378) 2-ヒドロキシイソ吉草酸、(379) キシルロース、(380) 乳酸が、特徴量として抽出された。
 また、上述の(362)~(380)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB5の、「肝障害の発症の有無」の予測に係るROC-AUCスコアは、「0.79」であった。
 「腎障害の発症の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(381)~(400)に示す血液代謝産物が、特徴量として抽出された。すなわち、(381) イノシトール、(382) 尿素、(383) クレアチニン、(384) アラビノース、(385) 4-ヒドロキシフェニル乳酸、(386) シスチン、(387) グルタミン酸、(388) エリスリトール、(389) トリプトファン、(390) マンニトール、(391) グルクロン酸、(392) アラビトール、(393) 2-ヒドロキシイソ酪酸、(394) 2-ヒドロキシ酪酸、(395) セリン、(396) コリン、(397) プロピオン酸、(398) トリメチルアミンオキシド、(399) 2-ケトグルタル酸、(400) グリセリン酸が、特徴量として抽出された。
 また、上述の(381)~(400)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB6の、「腎障害の発症の有無」の予測に係るROC-AUCスコアは、「0.83」であった。
 「血小板の20×104/μl未満への減少の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(401)~(417)に示す血液代謝産物が、特徴量として抽出された。すなわち、(401) グルタミン、(402) タウリン、(403) アセチルグリシン、(404) 乳酸、(405) ノナン酸、(406) トレオニン、(407) ヒスチジン、(408) イソ酪酸、(409) マンニトール、(410) グルタミン酸、(411) マンノース、(412) 3-ヒドロキシ酪酸、(413) N6-アセチルリジン、(414) 2-アミノピメリン酸、(415) カテコール、(416) イノシトール、(417) トリプトファンが、特徴量として抽出された。
 また、上述の(401)~(417)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB7の、「血小板の20×104/μl未満への減少の有無」の予測に係るROC-AUCスコアは、「0.81」であった。
 「D-dimerの1.0mg/ml以上への上昇の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(418)~(433)に示す血液代謝産物が、特徴量として抽出された。すなわち、(418) マルトース、(419) トリプトファン、(420) チロシン、(421) クエン酸、(422) シュウ酸、(423) クレアチン、(424) 2-ヒドロキシイソ吉草酸、(425) トレオン酸、(426) サルコシン、(427) 尿酸、(428) ラムノース、(429) グルカル酸、(430) 4-ヒドロキシプロリン、(431) イソ吉草酸、(432) ヒスチジン、(433) グルクロン酸が、特徴量として抽出された。
 また、上述の(418)~(433)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB8の、「D-dimerの1.0mg/ml以上への上昇の有無」の予測に係るROC-AUCスコアは、「0.76」であった。
 「400mg/dlよりも大きなフィブリノゲンの有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(434)~(453)に示す血液代謝産物が、特徴量として抽出された。すなわち、(434) マンノース、(435) セロトニン、(436) システイン、(437) 4-ヒドロキシプロリン、(438) フェニルアラニン、(439) バリン、(440) 3-ヒドロキシイソ酪酸、(441) ギ酸、(442) 2-アミノピメリン酸、(443) ロイシン、(444) クエン酸、(445) イノシトール、(446) グリシン、(447) トリプトファン、(448) マルトース、(449) インドール-3-酢酸、(450) 3-フェニル乳酸、(451) 尿酸、(452) チロシン、(453) アラニンが、特徴量として抽出された。
 また、上述の(434)~(453)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB9の、「400mg/dlよりも大きなフィブリノゲンの有無」の予測に係るROC-AUCスコアは、「0.88」であった。
 「長期入院の有無」を目的変数とした場合、訓練血液代謝産物データ31Bから、下記(461)~(480)に示す血液代謝産物が、特徴量として抽出された。すなわち、(461) エリトルロース、(462) コハク酸、(463) マンノース、(464) メチオニンスルホキシド、(465) マルトース、(466) マンニトール、(467) インドール-3-酢酸、(468) シュウ酸、(469) アスパラギン、(470) システイン、(471) L-カルニチン、(472) 尿素、(473) プロピオン酸、(474) グルタミン、(475) イノシトール、(476) 2-ケトグルタル酸、(477) 乳酸、(478) グルコース、(479) グリセロール、(480) デカン酸が、特徴量として抽出された。
 また、上述の(461)~(480)に示す血液代謝産物(特徴量)を用いて構築された予測モデルMB11の、「長期入院の有無」の予測に係るROC-AUCスコアは、「0.8」であった。
<腸内菌叢の細菌組成データを用いる場合について>
 図6および図7は、各アウトカムに対して訓練細菌組成データ31Cから抽出された特徴量、および、抽出された特徴量を用いて構築された予測モデルMCの予測のROC-AUCスコア等を示している。
 図6および図7に示すように、「重症化の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(501)~(520)に示す腸内細菌が、特徴量として抽出された。すなわち、(501) Synergistes sp (ref_mOTU_v25_11256)、(502) Blautia producta (ref_mOTU_v25_02151)、(503) Streptococcus oralis/pseudopneumoniae (ref_mOTU_v25_00296)、(504) Streptococcus anginosus (ref_mOTU_v25_00569)、(505) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、(506) Flavonifractor plautii (ref_mOTU_v25_05238)、(507) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、(508) Clostridiales species incertae sedis (meta_mOTU_v25_12301)、(509) Flavonifractor plautii (ref_mOTU_v25_02971)、(510) Intestinimonas butyriciproducens (ref_mOTU_v25_03928)、(511) Eubacterium ramulus (ref_mOTU_v25_03570)、(512) Bacteroides caccae (ref_mOTU_v25_03473)、(513) Clostridiales species incertae sedis (meta_mOTU_v25_12635)、(514) Dorea longicatena (ref_mOTU_v25_03692)、(515) Faecalicatena contorta (ref_mOTU_v25_03445)、(516) Clostridiales species incertae sedis (ext_mOTU_v26_26664)、(517) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、(518) Ruminococcaceae species incertae sedis (meta_mOTU_v25_12286)、(519) [Clostridium] clostridioforme/bolteae (ref_mOTU_v25_03442)、(520) Blautia species incertae sedis (ext_mOTU_v26_15352)が、特徴量として抽出された。
 また、上述の(501)~(520)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC1の、「重症化の有無」の予測に係るROC-AUCスコアは、「0.73」であった。
 「呼吸器症状の発現の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(521)~(540)に示す腸内細菌が、特徴量として抽出された。すなわち、(521) Bacteroides species incertae sedis (ext_mOTU_v26_26380)、(522) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、(523) Bifidobacterium breve (ref_mOTU_v25_01098)、(524) Eggerthella lenta (ref_mOTU_v25_00719)、(525) Sutterella wadsworthensis (ref_mOTU_v25_03066)、(526) Actinomyces graevenitzii (ref_mOTU_v25_04054)、(527) Collinsella species incertae sedis (ext_mOTU_v26_17345)、(528) Acidaminococcus intestini (ref_mOTU_v25_01949)、(529) Bifidobacterium adolescentis (ref_mOTU_v25_02703)、(530) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、(531) Veillonella atypica (ref_mOTU_v25_01941)、(532) Bifidobacterium longum (ref_mOTU_v25_01099)、(533) Gemella sanguinis (ref_mOTU_v25_04303)、(534) Streptococcus parasanguinis (ref_mOTU_v25_00312)、(535) Veillonella species incertae sedis (meta_mOTU_v25_13135)、(536) Actinomyces sp (ref_mOTU_v25_01914)、(537) Bifidobacterium bifidum (ref_mOTU_v25_03116)、(538) Lactobacillus fermentum/oris (ref_mOTU_v25_01407)、(539) Bacteroides sp. (ref_mOTU_v25_03475)、(540) Ruminococcus bromii (ref_mOTU_v25_00853)が、特徴量として抽出された。
 また、上述の(521)~(540)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC2の、「呼吸器症状の発現の有無」の予測に係るROC-AUCスコアは、「0.72」であった。
 「肺炎の発現の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(541)~(560)に示す腸内細菌が、特徴量として抽出された。すなわち、(541) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、(542) Firmicutes species incertae sedis (meta_mOTU_v25_12923)、(543) Eggerthella lenta (ref_mOTU_v25_00719)、(544) Streptococcus cristatus (ref_mOTU_v25_03967)、(545) Bacteroides caccae (ref_mOTU_v25_03473)、(546) Bacteroides species incertae sedis (ext_mOTU_v26_17504)、(547) Rothia mucilaginosa (ref_mOTU_v25_05265)、(548) Streptococcus species incertae sedis (ext_mOTU_v26_28826)、(549) Actinomyces sp (ref_mOTU_v25_12049)、(550) Lachnospiraceae species incertae sedis (meta_mOTU_v25_12240)、(551) Escherichia coli (ref_mOTU_v25_00095)、(552) Acidaminococcus intestini (ref_mOTU_v25_01949)、(553) Clostridiales sp. (ref_mOTU_v25_03444)、(554) Tyzzerella nexilis (ref_mOTU_v25_03689)、(555) Streptococcus thermophilus (ref_mOTU_v25_01348)、(556) Bacteroides rodentium/uniformis (ref_mOTU_v25_00855)、(557) Firmicutes species incertae sedis (meta_mOTU_v25_12227)、(558) Clostridiales species incertae sedis (ext_mOTU_v26_18011)、(559) Lachnospiraceae species incertae sedis (ext_mOTU_v26_16388)、(560) Isoptericola variabilis (ref_mOTU_v25_01910)が、特徴量として抽出された。
 また、上述の(541)~(560)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC3の、「肺炎の発現の有無」の予測に係るROC-AUCスコアは、「0.57」であった。
 「下痢の発現の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(561)~(580)に示す腸内細菌が、特徴量として抽出された。すなわち、(561) Streptococcus species incertae sedis (ext_mOTU_v26_28879)、(562) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、(563) Streptococcus australis (ref_mOTU_v25_00311)、(564) Tyzzerella species incertae sedis (meta_mOTU_v25_12266)、(565) Holdemanella biformis (meta_mOTU_v25_12329)、(566) [Ruminococcus] torques (ref_mOTU_v25_03703)、(567) Anaerotignum lactatifermentans (ref_mOTU_v25_02190)、(568) Butyricicoccus sp. (ref_mOTU_v25_02967)、(569) Eubacterium species incertae sedis (ext_mOTU_v26_16240)、(570) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、(571) Actinobacteria sp. (ref_mOTU_v25_01911)、(572) Streptococcus parasanguinis (ref_mOTU_v25_00312)、(573) [Clostridium] clostridioforme/bolteae (ref_mOTU_v25_03442)、(574) Ruminococcaceae species incertae sedis (ext_mOTU_v26_16263)、(575) Erysipelatoclostridium ramosum (ref_mOTU_v25_03439)、(576) Clostridiales species incertae sedis (meta_mOTU_v25_13392)、(577) Provencibacterium massiliense (ref_mOTU_v25_10126)、(578) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、(579) Clostridium sp (ref_mOTU_v25_09167)、(580) Sutterella wadsworthensis (ref_mOTU_v25_03066)が、特徴量として抽出された。
 また、上述の(561)~(580)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC4の、「下痢の発現の有無」の予測に係るROC-AUCスコアは、「0.73」であった。
 「肝障害の発症の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(581)~(600)に示す腸内細菌が、特徴量として抽出された。すなわち、(581) Eggerthella lenta (ref_mOTU_v25_00719)、(582) Eubacterium sp (meta_mOTU_v25_12688)、(583) Clostridium sp (meta_mOTU_v25_12609)、(584) Clostridiales species incertae sedis (ext_mOTU_v26_26595)、(585) Bacteroides sp. (ref_mOTU_v25_03475)、(586) Clostridiales species incertae sedis (meta_mOTU_v25_13012)、(587) Megasphaera sp (ref_mOTU_v25_03433)、(588) Bacteroides caecimuris (ref_mOTU_v25_03476)、(589) Bacteroides faecis/thetaiotaomicron (ref_mOTU_v25_01657)、(590) Lactococcus lactis (ref_mOTU_v25_01300)、(591) Clostridiales Family XIII (ext_mOTU_v26_16402)、(592) Lachnospiraceae species incertae sedis (ext_mOTU_v26_26654)、(593) Mogibacterium timidum (ref_mOTU_v25_04269)、(594) [Eubacterium] hallii (ref_mOTU_v25_03632)、(595) Coprococcus sp. (ref_mOTU_v25_01683)、(596) Bacteroides species incertae sedis (ext_mOTU_v26_26291)、(597) Eggerthellaceae species incertae sedis (ext_mOTU_v26_15442)、(598) uncultured Flavonifractor sp. (ref_mOTU_v25_07315)、(599) Clostridium species incertae sedis (ext_mOTU_v26_17379)、(600) Collinsella aerofaciens (ref_mOTU_v25_03626)が、特徴量として抽出された。
 また、上述の(581)~(600)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC5の、「肝障害の発症の有無」の予測に係るROC-AUCスコアは、「0.74」であった。
 「腎障害の発症の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(601)~(620)に示す腸内細菌が、特徴量として抽出された。すなわち、(601) Clostridiales species incertae sedis (meta_mOTU_v25_13006)、(602) Streptococcus anginosus (ref_mOTU_v25_00569)、(603) Streptococcus intermedius/constellatus (ref_mOTU_v25_00572)、(604) Actinomyces marseillensis/pacaensis (ref_mOTU_v25_03846)、(605) Firmicutes bacterium CAG:114 (ref_mOTU_v25_07728)、(606) Methanobrevibacter smithii (ref_mOTU_v25_03695)、(607) Clostridiales sp. (ref_mOTU_v25_03661)、(608) Megamonas funiformis/rupellensis (ref_mOTU_v25_02318)、(609) Streptococcus anginosus (ref_mOTU_v25_00570)、(610) Bacteroides species incertae sedis (ext_mOTU_v26_18132)、(611) Blautia massiliensis (ref_mOTU_v25_03342)、(612) Collinsella species incertae sedis (ext_mOTU_v26_17328)、(613) Bacteroides plebeius (ref_mOTU_v25_05069)、(614) Rothia dentocariosa (ref_mOTU_v25_04800)、(615) Clostridiales species incertae sedis (meta_mOTU_v25_12635)、(615) Anaerotruncus colihominis (ref_mOTU_v25_03438)、(617) Staphylococcaceae sp. (ref_mOTU_v25_04195)、(618) Streptococcus oralis (ref_mOTU_v25_00289)、(619) Alistipes finegoldii (ref_mOTU_v25_03682)、(620) Firmicutes species incertae sedis (meta_mOTU_v25_14410)が、特徴量として抽出された。
 また、上述の(601)~(620)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC6の、「腎障害の発症の有無」の予測に係るROC-AUCスコアは、「0.72」であった。
 「血小板の20×104/μl未満への減少の有無」(血栓症リスク1)を目的変数とした場合、訓練細菌組成データ31Cから、下記(621)~(640)に示す腸内細菌が、特徴量として抽出された。すなわち、(621) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、(622) Parabacteroides distasonis (ref_mOTU_v25_03640)、(623) [Clostridium] leptum (ref_mOTU_v25_03688)、(624) Erysipelatoclostridium ramosum (ref_mOTU_v25_03439)、(625) Coprococcus catus (ref_mOTU_v25_11851)、(626) Veillonella parvula (ref_mOTU_v25_01938)、(627) Rothia aeria/dentocariosa (ref_mOTU_v25_02524)、(628) Bacteroides species incertae sedis (ext_mOTU_v26_17684)、(629) Adlercreutzia equolifaciens (ref_mOTU_v25_08094)、(630) Clostridiales sp. (ref_mOTU_v25_05138)、(631) Eggerthella lenta (ref_mOTU_v25_00719)、(632) Blautia producta (ref_mOTU_v25_02152)、(633) Actinomyces sp. (ref_mOTU_v25_01913)、(634) Parabacteroides species incertae sedis (ext_mOTU_v26_26659)、(635) Clostridiales species incertae sedis (meta_mOTU_v25_13012)、(636) Clostridiales species incertae sedis (meta_mOTU_v25_12654)、(637) Clostridium sp (meta_mOTU_v25_12609)、(638) Eggerthellaceae species incertae sedis (ext_mOTU_v26_17308)、(639) Hungatella hathewayi (ref_mOTU_v25_03435)、(640) Firmicutes sp. (ref_mOTU_v25_02743)が、特徴量として抽出された。
 また、上述の(621)~(640)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC7の、「血小板の20×104/μl未満への減少の有無」の予測に係るROC-AUCスコアは、「0.71」であった。
 「D-dimerの1.0mg/ml以上への上昇の有無」(血栓症リスク2)を目的変数とした場合、訓練細菌組成データ31Cから、下記(641)~(660)に示す腸内細菌が、特徴量として抽出された。すなわち、(641) Bifidobacterium pseudocatenulatum (ref_mOTU_v25_02700)、(642) Bacteroides species incertae sedis (ext_mOTU_v26_17504)、(643) Akkermansia muciniphila (ref_mOTU_v25_03591)、(644) uncultured Eubacterium sp (meta_mOTU_v25_13063)、(645) Catabacter hongkongensis (ref_mOTU_v25_06126)、(646) Lactobacillus plantarum (ref_mOTU_v25_00930)、(647) Rothia dentocariosa (ref_mOTU_v25_04800)、(648) Slackia exigua (ref_mOTU_v25_01958)、(649) Enterococcus faecium/durans (ref_mOTU_v25_00323)、(650) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、(651) Streptococcus oralis/pseudopneumoniae (ref_mOTU_v25_00296)、(652) Bacteroides cellulosilyticus/fragilis (ref_mOTU_v25_01597)、(653) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、(654) Bacteroides sp. (ref_mOTU_v25_03475)、(655) Hungatella hathewayi (ref_mOTU_v25_03436)、(656) Staphylococcus aureus (ref_mOTU_v25_00340)、(657) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、(658) Lactobacillus casei/paracasei (ref_mOTU_v25_01406)、(659) Butyricicoccus sp. (ref_mOTU_v25_02967)、(660) Alistipes species incertae sedis (meta_mOTU_v25_12829)が、特徴量として抽出された。
 また、上述の(641)~(660)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC8の、「D-dimerの1.0mg/ml以上への上昇の有無」の予測に係るROC-AUCスコアは、「0.73」であった。
 「400mg/dlよりも大きなフィブリノゲンの有無」(血栓症リスク3)を目的変数とした場合、訓練細菌組成データ31Cから、下記(661)~(680)に示す腸内細菌が、特徴量として抽出された。すなわち、(661) Faecalibacterium prausnitzii (ref_mOTU_v25_06108)、(662) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、(663) Collinsella species incertae sedis (ext_mOTU_v26_17347)、(664) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、(665) Clostridiales sp. (ref_mOTU_v25_03444)、(666) Lachnoclostridium species incertae sedis (ext_mOTU_v26_16258)、(667) Eggerthella lenta (ref_mOTU_v25_00719)、(668) [Clostridium] asparagiforme/lavalense (ref_mOTU_v25_06317)、(669) Bacteroides sp. (ref_mOTU_v25_03475)、(670) Bacteroides rodentium/uniformis (ref_mOTU_v25_00855)、(671) Collinsella aerofaciens (ref_mOTU_v25_03626)、(672) [Ruminococcus] torques (ref_mOTU_v25_03703)、(673) Firmicutes sp. (ref_mOTU_v25_02743)、(674) Streptococcus sp. (ref_mOTU_v25_00283)、(675) Streptococcus oralis/pseudopneumoniae (ref_mOTU_v25_00296)、(676) Bifidobacterium longum (ref_mOTU_v25_01099)、(677) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、(678) Odoribacter splanchnicus (ref_mOTU_v25_03697)、(679) [Clostridium] clostridioforme/bolteae (ref_mOTU_v25_03442)、(680) [Ruminococcus] gnavus (ref_mOTU_v25_01594)が、特徴量として抽出された。
 また、上述の(661)~(680)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC9の、「400mg/dlよりも大きなフィブリノゲンの有無」の予測に係るROC-AUCスコアは、「0.69」であった。
 「長期入院の有無」を目的変数とした場合、訓練細菌組成データ31Cから、下記(701)~(720)に示す腸内細菌が、特徴量として抽出された。すなわち、(701) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、(702) Blautia massiliensis (ref_mOTU_v25_03342)、(703) Streptococcus anginosus/intermedius (ref_mOTU_v25_00567)、(704) Massilioclostridium coli (ref_mOTU_v25_10237)、(705) Eggerthella lenta (ref_mOTU_v25_00719)、(706) Clostridiales Family XIII (ext_mOTU_v26_16402)、(707) Clostridiales sp. (ref_mOTU_v25_04568)、(708) Firmicutes sp. (ref_mOTU_v25_02743)、(709) Bacteroides plebeius (ref_mOTU_v25_05069)、(710) Eisenbergiella tayi (ref_mOTU_v25_03446)、(711) Phascolarctobacterium succinatutens (ref_mOTU_v25_03700)、(712) Granulicatella species incertae sedis (ext_mOTU_v26_19463)、(713) Faecalibacterium prausnitzii (ref_mOTU_v25_06112)、(714) Veillonella dispar (ref_mOTU_v25_01940)、(715) Actinomyces graevenitzii (ref_mOTU_v25_04054)、(716) Pseudoflavonifractor species incertae sedis (meta_mOTU_v25_14283)、(717) Bacteroides faecis/thetaiotaomicron (ref_mOTU_v25_01657)、(718) Dorea longicatena (ref_mOTU_v25_03692)、(719) [Clostridium] leptum (ref_mOTU_v25_03688)、(720) Blautia obeum/wexlerae (ref_mOTU_v25_02154)が、特徴量として抽出された。
 また、上述の(701)~(720)に示す腸内細菌(特徴量)を用いて構築された予測モデルMC11の、「長期入院の有無」の予測に係るROC-AUCスコアは、「0.74」であった。
<糞便代謝産物濃度データを用いる場合について>
 図8および図9は、各アウトカムに対して訓練糞便代謝産物濃度データ31Dから抽出された特徴量、および、抽出された特徴量を用いて構築された予測モデルMDの予測のROC-AUCスコア等を示している。「重症化の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD1の、「重症化の有無」の予測に係るROC-AUCスコアは、「0.72」であった。「呼吸器症状の発現の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD2の、「呼吸器症状の発現の有無」の予測に係るROC-AUCスコアは、「0.74」であった。「肺炎の発現の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD3の、「肺炎の発現の有無」の予測に係るROC-AUCスコアは、「0.88」であった。「下痢の発現の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD4の、「下痢の発現の有無」の予測に係るROC-AUCスコアは、「0.78」であった。「肝障害の発症の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD5の、「肝障害の発症の有無」の予測に係るROC-AUCスコアは、「0.7」であった。「腎障害の発症の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD6の、「腎障害の発症の有無」の予測に係るROC-AUCスコアは、「0.69」であった。「血小板の20×104/μl未満への減少の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD7の、「血小板の20×104/μl未満への減少の有無」の予測に係るROC-AUCスコアは、「0.57」であった。「D-dimerの1.0mg/ml以上への上昇の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD8の、「D-dimerの1.0mg/ml以上への上昇の有無」の予測に係るROC-AUCスコアは、「0.84」であった。「400mg/dlよりも大きなフィブリノゲンの有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD9の、「400mg/dlよりも大きなフィブリノゲンの有無」の予測に係るROC-AUCスコアは、「0.75」であった。「長期入院の有無」を目的変数として訓練糞便代謝産物濃度データ31Dから抽出した特徴量を用いて構築された予測モデルMD11の、「長期入院の有無」の予測に係るROC-AUCスコアは、「0.75」であった。
 なお、腸内菌叢の遺伝子組成データ(対象遺伝子組成データ221A)、血液代謝産物データ(対象血液代謝産物データ221B)、および、腸内菌叢の細菌組成データ(対象細菌組成データ221C)の少なくとも1つを用いて「重症化の有無」等の各アウトカムを予測する検査対象者は、COVID-19感染者であってもよいし、COVID-19非感染者であってもよい。検査対象者がCOVID-19感染者である場合、その検査対象者が「感染した」と診断された時点以降の時点における「重症化の有無」等の各アウトカムを、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを用いて予測する。検査対象者がCOVID-19非感染者である場合、その検査対象者がCOVID-19に感染した場合の「重症化の有無」等の各アウトカムを、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを用いて予測する。
 これまで、予測モデルMについて、その概要を説明してきた。以下では、予測モデルMを構築するモデル生成装置1、および、モデル生成装置1により構築された予測モデルMを用いて重症化リスク等を予測する検査装置2について、図10等を参照して詳細に説明する。
 §4 システムへの適用例
 図10は、本実施形態に係るモデル生成装置1および検査装置2を含む分析システム100の全体概要を説明する図である。図10に示される通り、本実施形態に係る分析システム100は、モデル生成装置1および検査装置2を備える。
 (モデル生成装置)
 モデル生成装置1は、機械学習を実施することで訓練済みの推定器51(つまり、予測モデルM)を生成するように構成されたコンピュータである。具体的に、モデル生成装置1は、複数の学習データセット3を取得する。各学習データセット3は、訓練オミックスデータ31(または、訓練オミックスデータ31から抽出された特徴量データ)と、アウトカムデータ32との組み合わせにより構成される。
 訓練オミックスデータ31は、COVID-19感染者(112例)の各々の入院時の生体試料から取得したオミックスデータである。すなわち、訓練オミックスデータ31は、COVID-19感染者(112例)の各々の入院時の生体試料から取得した、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、訓練細菌組成データ31C、および、訓練糞便代謝産物濃度データ31Dの少なくとも1つを含む。
 なお、訓練遺伝子組成データ31Aから予測モデルMAを構築する方法は、訓練血液代謝産物データ31B、訓練細菌組成データ31C、および、訓練糞便代謝産物濃度データ31Dの各々から予測モデルMB、MC、および、MDの各々を構築する方法と同様である。また、予測モデルMAを用いて、検査対象者(被験者)の対象遺伝子組成データ221Aから重症化リスク等を予測する方法は、予測モデルMB、MC、および、MDの各々を用いて、検査対象者の対象血液代謝産物データ221B、対象細菌組成データ221C、および、対象糞便代謝産物濃度データ221Dの各々から重症化リスク等を予測する方法と同様である。そこで、以下では、オミックスデータが遺伝子組成データである例について、つまり、訓練オミックスデータ31が訓練遺伝子組成データ31Aであり、対象オミックスデータ221が対象遺伝子組成データ221Aである例について説明する。
 また、上述の通り、本実施形態において、「腸内菌叢の遺伝子組成データ(例えば、訓練遺伝子組成データ31Aおよび対象遺伝子組成データ221A)」とは、腸内菌叢から抽出される遺伝子(DNA)の種類及び各遺伝子の数からなるデータである。腸内菌叢の遺伝子組成データは、例えば、糞便から腸内細菌のDNAを抽出して解析することにより得ることができる。
 同様に、本実施形態において、「血液代謝産物データ(例えば、訓練血液代謝産物データ31Bおよび対象血液代謝産物データ221B)」とは、血液サンプル(例えば、全血、血漿、血清等)に含まれる代謝産物の種類及び各代謝産物の濃度からなるデータである。血液代謝産物データは、例えば、血液サンプルをGC/MS/MS分析、LC-MS/MS分析等によりメタボローム解析することにより得ることができる。また、本実施形態において、「腸内菌叢の細菌組成データ(例えば、訓練細菌組成データ31Cおよび対象細菌組成データ221C)」とは、腸内菌叢から抽出される腸内細菌の種類及び各腸内細菌の数(相対存在比)からなるデータである。腸内菌叢の細菌組成データは、例えば、糞便から腸内細菌のDNAを抽出してメタゲノム解析することにより得ることができる。
 また、アウトカムデータ32は、COVID-19感染者(112例)の各々について観察された「重症化の有無」等の各アウトカムを示すデータである。すなわち、アウトカムデータ32は、COVID-19感染者(112例)の各々について観察された、「COVID-19に罹患した際の症状の発現の有無」、「重症化の有無」、「合併症の有無」、「長期入院の有無」を示すデータである。前述の通り、「症状」は、「呼吸器症状」、「肺炎」、「下痢」を含む。さらに、「合併症」は、「肝障害」、「腎障害」、「血栓症」を含み、「血栓症」は、「血小板が20×104/μl未満に減少すること」、「D-ダイマーが1.0mg/ml以上に上昇すること」、「フィブリノゲンが400mg/dlよりも大きくなること」を含む。
 各学習データセット3において、訓練オミックスデータ31と、アウトカムデータ32とは対応している。例えば、或る感染者の生体試料中の訓練オミックスデータ31と、その同じ感染者について観察された「重症化の有無」等の各アウトカムを示すアウトカムデータ32との組合せによって、1つの学習データセット3が構成されている。
 モデル生成装置1は、取得された複数の学習データセット3を使用して、推定器51(予測モデルM)の機械学習を実施する。これにより、訓練済みの推定器51が生成される。特に、モデル生成装置1は、各学習データセット3に含まれる訓練オミックスデータ31から抽出された特徴量と、各学習データセット3に含まれるアウトカムデータ32により示される各アウトカムとの関係を学習する。
 (検査装置)
 検査装置2は、訓練済みの推定器51(予測モデルM)を使用して、検査対象者のオミックスデータ(対象オミックスデータ221)から「重症化の有無」等の各アウトカムを予測するように構成されたコンピュータである。具体的に、検査装置2は、対象オミックスデータ221を取得する。
 対象オミックスデータ221は、検査対象者の生体試料から取得したオミックスデータである。上述の通り、ここでは、対象オミックスデータ221が検査対象者の生体試料から取得した対象遺伝子組成データ221Aである例について説明する。検査装置2は、機械学習により訓練済みの推定器51を使用して、取得された対象オミックスデータ221から、「重症化の有無」等の各アウトカムを予測(推定)する。そして、検査装置2は、予測した結果(つまり、予測結果)に関する情報を出力する。
 (予測モデル)
 図10に例示するように、推定器51(予測モデルM)は、抽出部112より出力される対象オミックスデータ221の特徴量の入力を受け付けて、特徴量から「重症化の有無」などの各アウトカムを予測(推定)するように構成される。例えば、推定器51は、ランダムフォレストなどの公知のアルゴリズムを利用して、抽出部112よって各アウトカムについて選択(抽出)された説明変数(特徴量)から、「重症化の有無」などの各アウトカムを予測する。
 検査装置2が利用する訓練済みの推定器51(予測モデルM)は、訓練オミックスデータ31とアウトカムデータ32との組み合わせによりそれぞれ構成される複数の学習データセット3に対する機械学習を行なうことによって、予め構築されている。訓練済みの推定器51は、例えば、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、および、訓練細菌組成データ31Cの少なくとも1つと、アウトカムデータ32と、の組み合わせによりそれぞれ構成される複数の学習データセット3に対する機械学習を行なうことによって、予め構築されている。前述の通り、アウトカムデータ32は、COVID-19の罹患者(具体的には、COVID-19感染者(112例))について観察された、「重症化の有無」等の各アウトカムを示す。
 特に、訓練済みの推定器51は、対象オミックスデータ221から抽出された特徴量を説明変数とし、アウトカムデータ32により示される「重症化の有無」等の各アウトカムを目的変数とする、機械学習を行なうことによって、予め構築されている。訓練済みの推定器51は、例えば、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、および、訓練細菌組成データ31Cの少なくとも1つから抽出された特徴量を説明変数とし、アウトカムデータ32により示される各アウトカムを目的変数とする、機械学習を行なうことによって、予め構築されている。
 なお、図10の例では、モデル生成装置1および検査装置2は、ネットワークを介して互いに接続されている。ネットワークの種類は、例えば、インターネット、無線通信網、移動通信網、電話網、専用網等から適宜選択されてよい。ただし、モデル生成装置1および検査装置2の間でデータをやり取りする方法は、このような例に限定されなくてもよく、実施の形態に応じて適宜選択されてよい。例えば、モデル生成装置1および検査装置2の間では、記憶媒体を利用して、データがやり取りされてよい。
 また、図10の例では、モデル生成装置1および検査装置2は、それぞれ別個のコンピュータにより構成されている。しかしながら、本実施形態に係る分析システム100の構成は、このような例に限定されなくてもよく、実施の形態に応じて適宜決定されてよい。たとえば、モデル生成装置1および検査装置2は一体のコンピュータであってもよい。また、例えば、モデル生成装置1および検査装置2のうちの少なくとも一方は、複数台のコンピュータにより構成されてもよい。
 検査装置2が腸内菌叢の遺伝子組成データ(対象遺伝子組成データ221A)、血液代謝産物データ(対象血液代謝産物データ221B)、および、腸内菌叢の細菌組成データ(対象細菌組成データ221C)の少なくとも1つを用いて「重症化の有無」等の各アウトカムを予測する検査対象者は、COVID-19感染者であってもよいし、COVID-19非感染者であってもよい。検査対象者がCOVID-19感染者である場合、検査装置2は、その検査対象者が「感染した」と診断された時点以降の時点における「重症化の有無」等の各アウトカムを、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを用いて予測する。検査対象者がCOVID-19非感染者である場合、検査装置2は、その検査対象者がCOVID-19に感染した場合の「重症化の有無」等の各アウトカムを、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを用いて予測する。
 §5 各装置の構成例
 [ハードウェア構成]
 <モデル生成装置>
 図11は、本実施形態に係るモデル生成装置1のハードウェア構成の一例を模式的に例示する。図11に示されるとおり、本実施形態に係るモデル生成装置1は、制御部11、記憶部12、通信インタフェース13、外部インタフェース14、入力装置15、出力装置16、およびドライブ17が電気的に接続されたコンピュータである。なお、図11では、通信インタフェースおよび外部インタフェースを「通信I/F」および「外部I/F」と記載している。
 制御部11は、ハードウェアプロセッサであるCPU(Central Processing Unit)、RAM(Random Access Memory)、ROM(Read Only Memory)等を含み、プログラムおよび各種データに基づいて情報処理を実行するように構成される。制御部11は、さらに、不図示のGPU(Graphics Processing Unit)を含んでもよい。記憶部12は、メモリの一例であり、例えば、ハードディスクドライブ、ソリッドステートドライブ等で構成される。本実施形態では、記憶部12は、モデル生成プログラム81、学習データセット3、学習結果データ125等の各種情報を記憶する。
 モデル生成プログラム81は、訓練済みの推定器51(予測モデルM)を生成する機械学習処理をモデル生成装置1に実行させるためのプログラムである。モデル生成プログラム81は、当該機械学習処理の一連の命令を含む。学習データセット3は、推定器51の機械学習に使用される。学習結果データ125は、機械学習の実施により生成された訓練済みの推定器51に関する情報を示す。本実施形態では、学習結果データ125は、モデル生成プログラム81を実行した結果として生成される。
 通信インタフェース13は、例えば、有線LAN(Local Area Network)モジュール、無線LANモジュール等であり、ネットワークを介した有線又は無線通信を行うためのインタフェースである。モデル生成装置1は、通信インタフェース13を利用して、他の情報処理装置との間で、ネットワークを介したデータ通信を実行することができる。外部インタフェース14は、例えば、USB(Universal Serial Bus)ポート、専用ポート等であり、外部装置と接続するためのインタフェースである。外部インタフェース14の種類および数は任意に選択されてよい。学習データセット3(特に、訓練オミックスデータ31およびアウトカムデータ32の各々)は、外部の装置から取得されてもよい。例えば、訓練オミックスデータ31は、外部の測定装置から取得されてもよく、また、アウトカムデータ32は、外部のアウトカム管理サーバから取得されてもよい。この場合に、モデル生成装置1は、通信インタフェース13および外部インタフェース14の少なくとも一方を介して、当該外部の測定装置またはアウトカム管理サーバに接続されてよい。
 入力装置15は、例えば、マウス、キーボード等の入力を行うための装置である。また、出力装置16は、例えば、ディスプレイ、スピーカ等の出力を行うための装置である。ユーザ等のオペレータは、入力装置15および出力装置16を利用することで、モデル生成装置1を操作することができる。学習データセット3は、入力装置15を介した入力により得られてもよい。
 ドライブ17は、例えば、CDドライブ、DVDドライブ等であり、記憶媒体91に記憶されたプログラム等の各種情報を読み込むためのドライブ装置である。記憶媒体91は、コンピュータその他装置、機械等が、記憶されたプログラム等の各種情報を読み取り可能なように、当該プログラム等の情報を、電気的、磁気的、光学的、機械的又は化学的作用によって蓄積する媒体である。上記モデル生成プログラム81および学習データセット3の少なくともいずれかは、記憶媒体91に記憶されていてもよい。モデル生成装置1は、この記憶媒体91から、上記モデル生成プログラム81および学習データセット3の少なくともいずれかを取得してもよい。なお、図11では、記憶媒体91の一例として、CD、DVD等のディスク型の記憶媒体を例示している。しかしながら、記憶媒体91の種類は、ディスク型に限られなくてもよく、ディスク型以外であってもよい。ディスク型以外の記憶媒体として、例えば、フラッシュメモリ等の半導体メモリを挙げることができる。ドライブ17の種類は、記憶媒体91の種類に応じて任意に選択されてよい。
 なお、モデル生成装置1の具体的なハードウェア構成に関して、実施形態に応じて、適宜、構成要素の省略、置換および追加が可能である。例えば、制御部11は、複数のハードウェアプロセッサを含んでもよい。ハードウェアプロセッサは、マイクロプロセッサ、FPGA(field-programmable gate array)、DSP(digital signal processor)等で構成されてよい。記憶部12は、制御部11に含まれるRAMおよびROMにより構成されてもよい。通信インタフェース13、外部インタフェース14、入力装置15、出力装置16およびドライブ17の少なくともいずれかは省略されてもよい。モデル生成装置1は、複数台のコンピュータで構成されてもよい。この場合、各コンピュータのハードウェア構成は、一致していてもよいし、一致していなくてもよい。また、モデル生成装置1は、提供されるサービス専用に設計された情報処理装置の他、汎用のサーバ装置、PC(Personal Computer)等であってもよい。
 <検査装置>
 図12は、本実施形態に係る検査装置2のハードウェア構成の一例を模式的に例示する。図12に示されるとおり、本実施形態に係る検査装置2は、制御部21、記憶部22、通信インタフェース23、外部インタフェース24、入力装置25、出力装置26、およびドライブ27が電気的に接続されたコンピュータである。
 検査装置2の制御部21~ドライブ27および記憶媒体92はそれぞれ、上記モデル生成装置1の制御部11~ドライブ17および記憶媒体91それぞれと同様に構成されてよい。制御部21は、ハードウェアプロセッサであるCPU、RAM、ROM等を含み、プログラムおよびデータに基づいて各種情報処理を実行するように構成される。記憶部22は、例えば、ハードディスクドライブ、ソリッドステートドライブ等で構成される。本実施形態では、記憶部22は、検査プログラム82、学習結果データ125等の各種情報を記憶する。
 検査プログラム82は、訓練済みの推定器51(予測モデルM)を使用して予測タスク(推定タスク)を遂行する予測処理を検査装置2に実行させるためのプログラムである。検査プログラム82は、当該予測処理の一連の命令を含む。検査プログラム82および学習結果データ125の少なくともいずれかは、記憶媒体92に記憶されていてもよい。また、検査装置2は、検査プログラム82および学習結果データ125の少なくともいずれかを記憶媒体92から取得してもよい。
 学習データセット3と同様に、対象オミックスデータ221は、外部の装置から取得されてもよく、例えば、外部の測定装置から取得されてもよい。この場合に、検査装置2は、通信インタフェース23および外部インタフェース24の少なくとも一方を介して、当該外部の装置に接続されてよい。或いは、対象オミックスデータ221は、入力装置25を介した入力により得られてもよい。
 なお、検査装置2の具体的なハードウェア構成に関して、実施形態に応じて、適宜、構成要素の省略、置換および追加が可能である。例えば、制御部21は、複数のハードウェアプロセッサを含んでもよい。ハードウェアプロセッサは、マイクロプロセッサ、FPGA、DSP等で構成されてよい。記憶部22は、制御部21に含まれるRAMおよびROMにより構成されてもよい。通信インタフェース23、外部インタフェース24、入力装置25、出力装置26、およびドライブ27の少なくともいずれかは省略されてもよい。検査装置2は、複数台のコンピュータで構成されてもよい。この場合、各コンピュータのハードウェア構成は、一致していてもよいし、一致していなくてもよい。また、検査装置2は、提供されるサービス専用に設計された情報処理装置の他、汎用のサーバ装置、汎用のPC、タブレットPC、端末装置等であってもよい。
 [ソフトウェア構成]
 <モデル生成装置>
 図13は、本実施形態に係るモデル生成装置1のソフトウェア構成の一例を模式的に例示する。モデル生成装置1の制御部11は、記憶部12に記憶されたモデル生成プログラム81をRAMに展開する。そして、制御部11は、RAMに展開されたモデル生成プログラム81に含まれる命令をCPUにより解釈および実行して、各構成要素を制御する。これにより、図13に示されるとおり、本実施形態に係るモデル生成装置1は、取得部111、抽出部112、学習処理部113、および保存処理部114をソフトウェアモジュールとして備えるコンピュータとして動作する。すなわち、本実施形態では、モデル生成装置1の各ソフトウェアモジュールは、制御部11(CPU)により実現される。
 取得部111は、訓練オミックスデータ31とアウトカムデータ32との組み合わせによりそれぞれ構成される複数の学習データセット3を取得するように構成される。
 抽出部112は、複数の学習データセット3の各々に含まれる訓練オミックスデータ31から、推定器51(予測モデルM)の学習処理で用いる説明変数(特徴量)を抽出(選択)する。抽出部112は、Borutaアルゴリズムなどの公知のアルゴリズムを利用して、訓練オミックスデータ31から、「重症化の有無」などの各アウトカムを説明する説明変数(特徴量)を抽出してもよい。また、各アウトカムに応じて訓練オミックスデータ31から抽出すべき特徴量を予め特定することができている場合、抽出部112は、訓練オミックスデータ31から、予め特定されている特徴量を、各アウトカムに応じて抽出してもよい。
 例えば、特徴量は、検査対象者(被験者)の入院時に採取された生体試料中の、訓練遺伝子組成データ31A、訓練血液代謝産物データ31B、および、訓練細菌組成データ31Cの少なくとも1つを説明変数とし、前記検査対象者について観察された、COVID-19に罹患した際の「重症化の有無」等の各アウトカムを目的変数とする、特徴量選択により抽出される。
 具体的には、訓練オミックスデータ31(例えば、訓練遺伝子組成データ31A)の入力を受け付けた抽出部112は、「重症化の有無」を説明(予測)する説明変数(特徴量)として、訓練オミックスデータ31から、上述の(1)~(20)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「呼吸器症状の発現の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(21)~(40)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「肺炎の発現の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(41)~(60)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「下痢の発現の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(61)~(80)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「肝障害の発症の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(81)~(100)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「腎障害の発症の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(101)~(120)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「血小板の20×104/μl未満への減少の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(121)~(140)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「D-dimerの1.0mg/ml以上への上昇の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(141)~(160)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「400mg/dlよりも大きなフィブリノゲンの有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(161)~(180)に示す遺伝子についてのデータを抽出する。訓練オミックスデータ31の入力を受け付けた抽出部112は、「長期入院の有無」を予測する特徴量として、訓練オミックスデータ31から、上述の(201)~(220)に示す遺伝子についてのデータを抽出する。
 学習処理部113は、取得された学習データセット3を使用して、推定器51(予測モデルM)の機械学習を実施するように構成される。例えば、学習処理部113は、抽出部112によって訓練オミックスデータ31から選択(抽出)された特徴量(説明変数)と、アウトカムデータ32(目的変数)と、を用いて機械学習モデルを学習する。機械学習は、各学習データセット3について、訓練オミックスデータ31(特に、訓練オミックスデータ31から抽出された特徴量)を入力として与えることで推定器51により予測されるアウトカムが、訓練オミックスデータ31に対応するアウトカムデータ32により示されるアウトカムに適合するように、推定器51を訓練することにより構成される。
 例えば、機械学習において推定器51は、訓練オミックスデータ31から抽出された特徴量から推定器51が予測するアウトカムが、訓練オミックスデータ31に対応するアウトカムデータ32により示されるアウトカムに適合するように、訓練される。
 具体的には、「訓練オミックスデータ31から抽出された上述の(1)~(20)に示す遺伝子(特徴量)」から予測する「重症化の有無」が、アウトカムデータ32により示される「重症化の有無」に適合するように、推定器51(A1)(つまり、予測モデルMA1)は訓練される。「訓練オミックスデータ31から抽出された上述の(21)~(40)に示す遺伝子(特徴量)」から予測する「呼吸器症状の発現の有無」が、アウトカムデータ32により示される「呼吸器症状の発現の有無」に適合するように、推定器51(A2)(つまり、予測モデルMA2)は訓練される。「訓練オミックスデータ31から抽出された上述の(41)~(60)に示す遺伝子(特徴量)」から予測する「肺炎の発現の有無」が、アウトカムデータ32により示される「肺炎の発現の有無」に適合するように、推定器51(A3)(つまり、予測モデルMA3)は訓練される。「訓練オミックスデータ31から抽出された上述の(61)~(80)に示す遺伝子(特徴量)」から予測する「下痢の発現の有無」が、アウトカムデータ32により示される「下痢の発現の有無」に適合するように、推定器51(A4)(つまり、予測モデルMA4)は訓練される。「訓練オミックスデータ31から抽出された上述の(81)~(100)に示す遺伝子(特徴量)」から予測する「肝障害の発症の有無」が、アウトカムデータ32により示される「肝障害の発症の有無」に適合するように、推定器51(A5)(つまり、予測モデルMA5)は訓練される。「訓練オミックスデータ31から抽出された上述の(101)~(120)に示す遺伝子(特徴量)」から予測する「腎障害の発症の有無」が、アウトカムデータ32により示される「腎障害の発症の有無」に適合するように、推定器51(A6)(つまり、予測モデルMA6)は訓練される。「訓練オミックスデータ31から抽出された上述の(121)~(140)に示す遺伝子(特徴量)」から予測する「血小板の20×104/μl未満への減少の有無」が、アウトカムデータ32により示される「血小板の20×104/μl未満への減少の有無」に適合するように、推定器51(A7)(つまり、予測モデルMA7)は訓練される。「訓練オミックスデータ31から抽出された上述の(141)~(160)に示す遺伝子(特徴量)」から予測する「D-dimerの1.0mg/ml以上への上昇の有無」が、アウトカムデータ32により示される「D-dimerの1.0mg/ml以上への上昇の有無」に適合するように、推定器51(A8)(つまり、予測モデルMA8)は訓練される。「訓練オミックスデータ31から抽出された上述の(161)~(180)に示す遺伝子(特徴量)」から予測する「400mg/dlよりも大きなフィブリノゲンの有無」が、アウトカムデータ32により示される「400mg/dlよりも大きなフィブリノゲンの有無」に適合するように、推定器51(A9)(つまり、予測モデルMA9)は訓練される。「訓練オミックスデータ31から抽出された上述の(201)~(220)に示す遺伝子(特徴量)」から予測する「長期入院の有無」が、アウトカムデータ32により示される「長期入院の有無」に適合するように、推定器51(A11)(つまり、予測モデルMA11)は訓練される。
 保存処理部114は、機械学習により生成された訓練済みの推定器51(予測モデルM)に関する情報を学習結果データ125として生成し、生成された学習結果データ125を所定の記憶領域に保存するように構成される。学習結果データ125は、訓練済みの推定器51(予測モデルM)を再生するための情報を含むように適宜構成されてよい。
 (機械学習)
 図13に示されるとおり、抽出部112によって各訓練オミックスデータ31から抽出された特徴量を推定器51(予測モデルM)に入力することで、各訓練オミックスデータ31(特に、その特徴量)に対する、アウトカムについての予測(推定)結果を得ることができる。学習処理部113は、機械学習において、各訓練オミックスデータ31(特に、その特徴量)に対して得られる予測結果(推定結果)と、各訓練オミックスデータ31に対応する各アウトカムデータ32により示されるアウトカムとの間の誤差が小さくなるように、推定器51のパラメータを調整するように構成される。
 保存処理部114は、上記機械学習により生成された訓練済みの推定器51(予測モデルM)を再生するための学習結果データ125を生成するように構成される。訓練済みの推定器51を再生可能であれば、学習結果データ125の構成は、特に限定されなくてよく、実施の形態に応じて適宜決定されてよい。一例として、学習結果データ125は、上記機械学習の調整により得られた各パラメータの値を示す情報を含んでよい。場合によって、学習結果データ125は、推定器51の構造を示す情報を含んでよい。構造は、例えば、層の数、各層の種類、各層に含まれるノードの数、隣接する層のノード同士の結合関係等により特定されてよい。
 <検査装置>
 図14は、本実施形態に係る検査装置2のソフトウェア構成の一例を模式的に例示する。検査装置2の制御部21は、記憶部22に記憶された検査プログラム82をRAMに展開する。そして、制御部21は、RAMに展開された検査プログラム82に含まれる命令をCPUにより解釈および実行して、各構成要素を制御する。これにより、図14に示されるとおり、本実施形態に係る検査装置2は、取得部211、抽出部212、予測部213、および出力部214をソフトウェアモジュールとして備えるコンピュータとして動作する。すなわち、本実施形態では、検査装置2の各ソフトウェアモジュールも、モデル生成装置1と同様に、制御部21(CPU)により実現される。
 取得部211は、対象オミックスデータ221を取得するように構成される。取得部211は、例えば、検査対象者(被験者)から採取された生体試料中の、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを取得する。取得部211は、対象オミックスデータ221を、外部の装置から取得してもよく、例えば、外部の測定装置から取得してもよい。すなわち、検査装置2にとって、検査装置2自体が、検査対象者から採取された生体試料から、対象オミックスデータ221を測定する必要はない。
 抽出部212は、対象オミックスデータ221から、推定器51(予測モデルM)を用いた予測処理で用いる説明変数(特徴量)を抽出(選択)する。抽出部212は、Borutaアルゴリズムなどの公知のアルゴリズムを利用して、対象オミックスデータ221から、「重症化の有無」などの各アウトカムを説明する説明変数(特徴量)を抽出してもよい。例えば、モデル生成装置1の抽出部112が、訓練オミックスデータ31から、「重症化の有無」などの各アウトカムに応じた特徴量を抽出するように訓練されている場合、訓練済の抽出部112を、検査装置2の抽出部212として用いてもよい。また、各アウトカムに応じて対象オミックスデータ221から抽出すべき特徴量を予め特定することができている場合、抽出部212は、対象オミックスデータ221から、予め特定されている特徴量を、各アウトカムに応じて抽出してもよい。
 具体的には、対象オミックスデータ221の入力を受け付けた抽出部212は、「重症化の有無」を説明(予測)する説明変数(特徴量)として、対象オミックスデータ221から、上述の(1)~(20)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「呼吸器症状の発現の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(21)~(40)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「肺炎の発現の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(41)~(60)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「下痢の発現の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(61)~(80)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「肝障害の発症の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(81)~(100)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「腎障害の発症の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(101)~(120)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「血小板の20×104/μl未満への減少の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(121)~(140)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「D-dimerの1.0mg/ml以上への上昇の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(141)~(160)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「400mg/dlよりも大きなフィブリノゲンの有無」を予測する特徴量として、対象オミックスデータ221から、上述の(161)~(180)に示す遺伝子についてのデータを抽出する。対象オミックスデータ221の入力を受け付けた抽出部212は、「長期入院の有無」を予測する特徴量として、対象オミックスデータ221から、上述の(201)~(220)に示す遺伝子についてのデータを抽出する。
 予測部213は、学習結果データ125を保持していることで、訓練済みの推定器51(予測モデルM)を備える。予測部213は、訓練済みの推定器51を使用して、取得された対象オミックスデータ221に対するアウトカムを予測(推定)するように構成される。予測部213は、例えば、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを、予め準備しておいた訓練済みの推定器51に入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する。
 具体的には、予測部213は、抽出部212により対象オミックスデータ221から抽出された特徴量を訓練済みの推定器51を入力して、「重症化の有無」等の各アウトカムについての予測(推定)の結果を推定器51から得るように構成される。予測部213は、推定器51から得た各アウトカムについての予測結果を、出力部214に通知する。
 例えば、訓練済みの推定器51(A1)(つまり、予測モデルMA1)は、上述の(1)~(20)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「重症化の有無」を予測する。訓練済みの推定器51(A2)(つまり、予測モデルMA2)は、上述の(21)~(40)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「呼吸器症状の発現の有無」を予測する。訓練済みの推定器51(A3)(つまり、予測モデルMA3)は、上述の(41)~(60)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「肺炎の発現の有無」を予測する。訓練済みの推定器51(A4)(つまり、予測モデルMA4)は、上述の(61)~(80)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「下痢の発現の有無」を予測する。訓練済みの推定器51(A5)(つまり、予測モデルMA5)は、上述の(81)~(100)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「肝障害の発症の有無」を予測する。訓練済みの推定器51(A6)(つまり、予測モデルMA6)は、上述の(101)~(120)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「腎障害の発症の有無」を予測する。訓練済みの推定器51(A7)(つまり、予測モデルMA7)は、上述の(121)~(140)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「血小板の20×104/μl未満への減少の有無」を予測する。訓練済みの推定器51(A8)(つまり、予測モデルMA8)は、上述の(141)~(160)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「D-dimerの1.0mg/ml以上への上昇の有無」を予測する。訓練済みの推定器51(A9)(つまり、予測モデルMA9)は、上述の(161)~(180)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「400mg/dlよりも大きなフィブリノゲンの有無」を予測する。訓練済みの推定器51(A11)(つまり、予測モデルMA11)は、上述の(201)~(220)に示す遺伝子(特徴量)を用いて、検査対象者(対象オミックスデータ221)の「長期入院の有無」を予測する。
 出力部214は、予測部213(推定器51)から通知された各アウトカムについての予測結果に関する情報を出力するように構成される。例えば、出力部214は、各アウトカムについての予測結果に関する情報を、出力装置16に出力させてもよい。
 <その他>
 モデル生成装置1および検査装置2の各ソフトウェアモジュールに関しては後述する動作例で詳細に説明する。なお、本実施形態では、モデル生成装置1および検査装置2の各ソフトウェアモジュールがいずれも汎用のCPUによって実現される例について説明している。しかしながら、上記ソフトウェアモジュールの一部又は全部が、1又は複数の専用のプロセッサ(例えば、グラフィックスプロセッシングユニット)により実現されてもよい。上記各モジュールは、ハードウェアモジュールとして実現されてもよい。また、モデル生成装置1および検査装置2それぞれのソフトウェア構成に関して、実施形態に応じて、適宜、ソフトウェアモジュールの省略、置換および追加が行われてもよい。
 §6 各装置の動作例
 [モデル生成装置]
 図15は、本実施形態に係るモデル生成装置1による機械学習に関する処理手順の一例を示すフローチャートである。以下で説明するモデル生成装置1の処理手順は、モデル生成方法の一例である。ただし、以下で説明するモデル生成装置1の処理手順は一例に過ぎず、各ステップは可能な限り変更されてよい。また、以下の処理手順について、実施の形態に応じて、適宜、ステップの省略、置換、および追加が行われてよい。
 (ステップS101)
 ステップS101では、制御部11は、取得部111として動作し、訓練オミックスデータ31とアウトカムデータ32との組み合わせによりそれぞれ構成される複数の学習データセット3を取得する。取得する学習データセット3の数は、特に限定されなくてよく、機械学習を実施可能なように実施の形態に応じて適宜決定されてよい。学習データセット3を取得すると、制御部11は、次のステップS102に処理を進める。
 (ステップS102)
 ステップS102では、制御部11は、抽出部112として動作し、取得された各学習データセット3に含まれる訓練オミックスデータ31から特徴量を抽出する。すなわち、制御部11は、推定器51(予測モデルM)の学習処理で用いる説明変数(特徴量)を抽出(選択)する。例えば、制御部11は、訓練オミックスデータ31から、「重症化の有無」などの各アウトカムを説明する説明変数(特徴量)を抽出する。特徴量を抽出すると、制御部11は、次のステップS103に処理を進める。
 (ステップS103)
 ステップS103では、制御部11は、学習処理部113として動作し、取得された学習データセット3を使用して、推定器51(予測モデルM)の機械学習を実施する。具体的には、制御部11は、ステップS102で訓練オミックスデータ31から抽出された特徴量と、訓練オミックスデータ31に対応するアウトカムデータ32とを使用して、推定器51の機械学習を実施する。
 機械学習の処理の一例として、まず、制御部11は、機械学習の処理対象となる推定器51(予測モデルM)の初期設定を行う。推定器51(予測モデルM)の構造およびパラメータの初期値は、テンプレートにより与えられてもよいし、或いはオペレータの入力により決定されてもよい。また、追加学習又は再学習を行う場合、制御部11は、過去の機械学習により得られた学習結果データに基づいて、推定器51(予測モデルM)の初期設定を行ってもよい。
 次に、制御部11は、機械学習により、各訓練オミックスデータ31(特に、その特徴量)から予測(推定)したアウトカムが、各訓練オミックスデータ31に対応する各アウトカムデータ32により示されるアウトカムに適合するように推定器51を訓練する。言い換えれば、制御部11は、機械学習により、各訓練オミックスデータ31の特徴量から推定したアウトカムが、対応する各アウトカムデータ32により示されるアウトカムに適合するように、推定器51のパラメータの値を調整する。この訓練処理には、確率的勾配降下法、ミニバッチ勾配降下法等が用いられてよい。
 訓練処理の一例として、先ず、制御部11は、各訓練オミックスデータ31(特に、その特徴量)を推定器51に入力し、順方向の演算処理を実行する。この順方向の演算処理により、制御部11は、各訓練オミックスデータ31(特に、その特徴量)に対するアウトカムの予測結果(推定結果)を、推定器51から取得する。
 次に、制御部11は、得られた推定結果と、各訓練オミックスデータ31に対応する各アウトカムデータ32により示されるアウトカムとの間の誤差を算出する。誤差の算出には、損失関数が用いられてよい。続いて、制御部11は、算出された誤差の勾配を算出する。制御部11は、誤差逆伝播法により、算出された誤差の勾配を用いて、推定器51のパラメータの値の誤差を算出する。そして、制御部11は、算出された誤差に基づいて、推定器51のパラメータの値を更新する。パラメータの値を更新する程度は、学習率により調節されてよい。学習率は、オペレータの指定により与えられてもよいし、プログラム内の設定値として与えられてもよい。
 制御部11は、上記一連の更新処理により、各訓練オミックスデータ31(特に、その特徴量)について、算出される誤差の和が小さくなるように、パラメータの値を調整する。例えば、規定回数実行する、算出される誤差の和が閾値以下になる等の所定の条件を満たすまで、制御部11は、上記一連の更新処理によるパラメータの値の調整を繰り返してもよい。この機械学習の処理の結果として、制御部11は、与えられたオミックスデータ(特に、その特徴量)に対してアウトカムを予測(推定)する能力を獲得した訓練済みの推定器51(予測モデルM)を生成することができる。機械学習の処理が完了すると、制御部11は、次のステップS104に処理を進める。
 (ステップS104)
 ステップS104では、制御部11は、保存処理部114として動作し、機械学習により生成された訓練済みの推定器51(予測モデルM)に関する情報を学習結果データ125として生成する。そして、制御部11は、生成された学習結果データ125を所定の記憶領域に保存する。
 所定の記憶領域は、例えば、制御部11内のRAM、記憶部12、外部記憶装置、記憶メディア又はこれらの組み合わせであってよい。記憶メディアは、例えば、CD、DVD等であってよく、制御部11は、ドライブ17を介して記憶メディアに学習結果データ125を格納してもよい。外部記憶装置は、例えば、NAS(Network Attached Storage)等のデータサーバであってよい。この場合、制御部11は、通信インタフェース13を利用して、ネットワークを介してデータサーバに学習結果データ125を格納してもよい。また、外部記憶装置は、例えば、外部インタフェース14を介してモデル生成装置1に接続された外付けの記憶装置であってもよい。
 学習結果データ125の保存が完了すると、制御部11は、本動作例に係るモデル生成装置1の処理手順を終了する。
 なお、生成された学習結果データ125は、任意のタイミングで検査装置2に提供されてよい。検査装置2は、通信インタフェース23を利用して、モデル生成装置1又はデータサーバにネットワークを介してアクセスすることで、学習結果データ125を取得してもよい。また、検査装置2は、記憶媒体92を介して、学習結果データ125を取得してもよい。また、学習結果データ125は、検査装置2に予め組み込まれてもよい。
 更に、制御部11は、上記ステップS101~ステップS104の処理を定期又は不定期に繰り返すことで、学習結果データ125を更新又は新たに生成してもよい。この繰り返しの際に、機械学習に使用する学習データセット3の少なくとも一部の変更、修正、追加、削除等が適宜実行されてよい。そして、制御部11は、更新した又は新たに生成した学習結果データ125を任意の方法で検査装置2に提供することで、検査装置2の保持する学習結果データ125を更新してもよい。
 [検査装置]
 図16は、本実施形態に係る検査装置2による予測タスクの遂行に関する処理手順の一例を示すフローチャートである。以下で説明する検査装置2の処理手順は、推定方法の一例である。ただし、以下で説明する検査装置2の処理手順は一例に過ぎず、各ステップは可能な限り変更されてよい。また、以下の処理手順について、実施の形態に応じて、適宜、ステップの省略、置換、および追加が行われてよい。
 (ステップS201)
 ステップS201(取得ステップ)では、制御部21は、取得部211として動作し、予測タスクの対象となる対象オミックスデータ221を取得する。制御部21は、例えば、検査対象者(被験者)から採取された生体試料中の、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを取得する。前述の通り、対象オミックスデータ221は、検査対象者(被験者)から採取された生体試料を測定することで得えられる。対象オミックスデータ221の構成は、学習データセット3における訓練オミックスデータ31と同様である。対象オミックスデータ221のデータ形式は、実施の形態に応じて適宜選択されてよい。制御部21は、対象オミックスデータ221を直接的に取得してもよい。或いは、制御部21は、例えば、ネットワーク、センサ、他のコンピュータ、記憶媒体92等を介して対象オミックスデータ221を間接的に取得してもよい。対象オミックスデータ221を取得すると、制御部21は、次のステップS202に処理を進める。
 (ステップS202)
 ステップS202では、制御部21は、抽出部212として動作し、対象オミックスデータ221から、「重症化の有無」などの各アウトカムを説明する特徴量を抽出する。すなわち、抽出部212は、対象オミックスデータ221から、推定器51(予測モデルM)へと出力する説明変数(特徴量)を抽出(選択)する。特徴量を抽出すると、制御部21は、次のステップS203に処理を進める。
 (ステップS203)
 ステップS203(予測ステップ)では、制御部21は、予測部213として動作する。すなわち、制御部21は、対象オミックスデータ221を、予め準備しておいた訓練済みの推定器51(予測モデルM)に入力することで、「重症化の有無」等の各アウトカムを予測する。制御部21は、例えば、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つを、訓練済みの推定器51に入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する。
 具体的には、ステップS203において先ず、制御部21は、学習結果データ125を参照して、訓練済みの推定器51(予測モデルM)の設定を行う。そして、制御部21は、訓練済みの推定器51を使用して、取得された対象オミックスデータ221(特に、その特徴量)に対する予測タスクの解を推定する。この推定の演算処理は、上記機械学習の訓練処理における順方向の演算処理と同様であってよい。例えば、制御部21は、対象オミックスデータ221から抽出された特徴量を訓練済みの推定器51に入力し、訓練済みの推定器51の順方向の演算処理を実行する。この演算処理を実行した結果として、制御部21は、対象オミックスデータ221(特に、その特徴量)に対して予測タスクの解を推定した結果(予測結果)を推定器51から取得することができる。予測結果を取得すると、制御部21は、次のステップS204に処理を進める。
 (ステップS204)
 ステップS204(出力ステップ)では、制御部21は、出力部214として動作し、S203にて取得した予測結果に関する情報を出力する。
 出力先および出力する情報の内容はそれぞれ、実施の形態に応じて適宜決定されてよい。例えば、制御部21は、ステップS203により得られた予測結果を出力装置26または他のコンピュータの出力装置にそのまま出力してもよい。また、制御部21は、得られた予測結果に基づいて、何らかの情報処理を実行してもよい。そして、制御部21は、その情報処理を実行した結果を、予測結果に関する情報として出力してもよい。この情報処理を実行した結果の出力には、予測結果に応じて制御対象装置の動作を制御すること等が含まれてよい。出力先は、例えば、出力装置26、他のコンピュータの出力装置、制御対象装置等であってよい。
 予測結果に関する情報の出力が完了すると、制御部21は、本動作例に係る検査装置2の処理手順を終了する。
 §7 変形例
 以上、本発明の実施の形態を詳細に説明してきたが、前述までの説明はあらゆる点において本発明の例示に過ぎない。本発明の範囲を逸脱することなく種々の改良又は変形を行うことができることは言うまでもない。例えば、以下のような変更が可能である。なお、以下では、上記実施形態と同様の構成要素に関しては同様の符号を用い、上記実施形態と同様の点については、適宜説明を省略した。以下の変形例は適宜組み合わせ可能である。
 本実施形態においては、ランダムフォレストにより、予測モデルM(推定器51)を構築する例を説明したが、学習処理部113が予測モデルM(推定器51)を構築するために行なう学習の手法は、これに限られるものではない。予測モデルMの構築におけるアルゴリズムとしては、機械学習に用いるアルゴリズムなどの公知のものを利用することができる。機械学習アルゴリズムの例としては、ランダムフォレスト以外に、線形カーネルのサポートベクターマシン(SVM linear)、rbfカーネルのサポートベクターマシン(SVM rbf)ニューラルネットワーク(Nerural net)、一般線形モデル(Generalized linear model)、正則化線形判別分析(Regularized linear discriminant analysis)、正則化ロジスティック回帰(Regularized logistic regression)などが挙げられる。
 また、本実施形態においては、抽出部112(または抽出部212)と推定器51とを2つの異なる機能部として説明した。しかしながら、抽出部112(または抽出部212)と推定器51とは一体に構成されてもよく、例えば、ニューラルネットワーク(ニューラルネットワークモジュール)として一体に構成されてもよい。このようなニューラルネットワークモジュールは、抽出部112(または抽出部212)としての処理を、つまり、所定の条件を満たす要素を対象の集合から抽出する演算(以下、「抽出演算」とも記載する)を含む。抽出演算は、回帰演算の過程で算出される複数の要素を一部の要素(例えば、望ましい要素)に絞る微分不可能な演算であれば、抽出演算の内容は、特に限定されなくてよく、実施の形態に応じて適宜決定されてよい。機械学習は、各学習データセット3について、ニューラルネットワークモジュールを使用して訓練オミックスデータ31から回帰される値がアウトカムデータ32により示される各アウトカムに適合するようにニューラルネットワークモジュールを訓練することにより構成される。ニューラルネットワークモジュールを訓練することは、順伝播による回帰の試行処理、および、逆伝播による演算パラメータの調整処理により、各演算パラメータの値を調整することにより構成される。この機械学習の結果、オミックスデータ(訓練オミックスデータ31)から、各アウトカムを回帰する能力を獲得した訓練済みのニューラルネットワークモジュールを生成することができる。そして、検査装置2は、このような訓練済みのニューラルネットワークモジュールを使用することにより、対象オミックスデータ221から、検査対象者(被験者)の各アウトカムを予測する予測タスクを遂行することができる。例えば、検査装置2は、訓練済みのニューラルネットワークモジュール(予測モデルM)を使用することにより、検査対象者(被験者)から採取された生体試料中の、対象遺伝子組成データ221A、対象血液代謝産物データ221B、および、対象細菌組成データ221Cの少なくとも1つから、検査対象者の「重症化の有無」等の各アウトカムを予測することができる。
 S11…測定ステップ、S201…取得ステップ、S203…予測ステップ、
S204…出力ステップ、M…予測モデル、2…検査装置、
31A…訓練遺伝子組成データ(腸内菌叢の遺伝子組成データ)、
31B…訓練血液代謝産物データ(血液代謝産物データ)、
221A…対象遺伝子組成データ(腸内菌叢の遺伝子組成データ)、
221B…対象血液代謝産物データ(血液代謝産物データ)
211…取得部、213…予測部、214…出力部、51…推定器(予測モデル)

Claims (33)

  1.  COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法であって、
     被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを測定する測定ステップを含み、
     腸内菌叢の遺伝子組成データおよび血液代謝産物データの少なくとも一方は、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用され、
     腸内菌叢の細菌組成データは、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査するために使用される、
    前記検査方法。
  2.  前記COVID-19に罹患した際の症状は、呼吸器症状、肺炎、下痢の少なくとも1つである、
    請求項1に記載の検査方法。
  3.  前記合併症は、肝障害、腎障害、血栓症の少なくとも1つである、
    請求項1または2に記載の検査方法。
  4.  前記血栓症は、血小板が20×104/μl未満に減少すること、D-ダイマーが1.0mg/ml以上に上昇すること、フィブリノゲンが400mg/dlよりも大きくなることの少なくとも1つである、
    請求項3に記載の検査方法。
  5.  コンピュータが、
     前記測定ステップで得られた腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを取得する取得ステップと、
     前記取得ステップにて取得された腸内細菌の遺伝子組成データおよび血液代謝産物データの少なくとも一方を、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測し、または、前記取得ステップにて取得された腸内菌叢の細菌組成データを、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する予測ステップと、
     前記予測ステップにて予測された予測結果を出力する出力ステップと、
    を実行する、
    請求項1から4のいずれか1項に記載の検査方法。
  6.  前記予測モデルは、
      COVID-19の罹患者の入院時に採取された生体試料中の、前記腸内菌叢の遺伝子組成データ、前記血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つと、
      前記罹患者について観察された、COVID-19に罹患した際の前記症状の発現、
    前記重症化、前記合併症の発症、及び前記長期入院の少なくとも1つの有無を示す情報と、
    の組み合わせによりそれぞれ構成される複数の学習データセットに対する機械学習を行なうことによって、予め構築されている、
    請求項5に記載の検査方法。
  7.  前記予測モデルは、
      COVID-19の罹患者の入院時に採取された生体試料中の、前記腸内菌叢の遺伝子組成データ、前記血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つから抽出された特徴量を説明変数とし、
      前記罹患者について観察された、COVID-19に罹患した際の前記症状の発現、前記重症化、前記合併症の発症、及び前記長期入院の少なくとも1つの有無を目的変数とする、
    機械学習を行なうことによって、予め構築されている、
    請求項5または6に記載の検査方法。
  8.  前記特徴量は、
      前記罹患者の入院時に採取された生体試料中の、前記腸内菌叢の遺伝子組成データ、前記血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを説明変数とし、
      前記罹患者について観察された、COVID-19に罹患した際の前記症状の発現、前記重症化、前記合併症の発症、及び前記長期入院の少なくとも1つの有無を目的変数とする、
    特徴量選択により抽出される、
    請求項7に記載の検査方法。
  9.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記重症化の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(1)~(20)に示す遺伝子である、
    (1) K11636、(2) K03932、(3) K19064、(4) K01455、
    (5) K03753、(6) K03976、(7) K18123、(8) K03572、
    (9) K14654、(10) K07991、(11) K23509、(12) K03191、
    (13) K00303、(14) K03685、(15) K13891、(16) K03394、
    (17) K01581、(18) K03187、(19) K09817、(20) K04078
    請求項7または8に記載の検査方法。
  10.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての呼吸器症状の発現の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(21)~(40)に示す遺伝子である、
    (21) K01200、(22) K07650、(23) K02860、(24) K03769、
    (25) K07770、(26) K01222、(27) K07025、(28) K12573、
    (29) K19116、(30) K02028、(31) K01114、(32) K01924、
    (33) K01857、(34) K01958、(35) K03719、(36) K03402、
    (37) K02834、(38) K01621、(39) K03650、(40) K07099
    請求項7または8に記載の検査方法。
  11.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての肺炎の発現の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(41)~(60)に示す遺伝子である、
    (41) K07171、(42) K21469、(43) K06013、(44) K07459、
    (45) K02107、(46) K01463、(47) K04652、(48) K07775、
    (49) K10231、(50) K10540、(51) K19334、(52) K00172、
    (53) K23265、(54) K01153、(55) K01455、(56) K17835、
    (57) K15899、(58) K14623、(59) K09767、(60) K13620
    請求項7または8に記載の検査方法。
  12.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての下痢の発現の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(61)~(80)に示す遺伝子である、
    (61) K07309、(62) K03092、(63) K03395、(64) K19883、
    (65) K03489、(66) K01004、(67) K13075、(68) K1192、
    (69) K21469、(70) K06311、(71) K07192、(72) K07668、
    (73) K22579、(74) K01506、(75) K15524、(76) K11050、
    (77) K03587、(78) K02745、(79) K11178、(80) K03325
    請求項7または8に記載の検査方法。
  13.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての肝障害の発症の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(81)~(100)に示す遺伝子である、
    (81) K01546、(82) K09858、(83) K21636、(84) K06895、
    (85) K15868、(86) K05810、(87) K00613、(88) K17363、
    (89) K03299、(90) K01547、(91) K13821、(92) K18889、
    (93) K11194、(94) K18350、(95) K02558、(96) K03503、
    (97) K19169、(98) K00183、(99) K02445、(100) K05966
    請求項7または8に記載の検査方法。
  14.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての腎障害の発症の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(101)~(120)に示す遺伝子である、
    (101) K02919、(102) K13668、(103) K19221、
    (104) K05873、(105) K18214、(106) K18906、
    (107) K01227、(108) K02057、(109) K20453、
    (110) K13730、(111) K06941、(112) K05814、
    (113) K12276、(114) K08258、(115) K00973、
    (116) K02110、(117) K07038、(118) K07010、
    (119) K06904、(120) K11690
    請求項7または8に記載の検査方法。
  15.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としてのD-dimerの1.0mg/ml以上への上昇の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(141)~(160)に示す遺伝子である、
    (141) K03424、(142) K21465、(143) K02744、
    (144) K03781、(145) K06218、(146) K01590、
    (147) K05305、(148) K09811、(149) K02193、
    (150) K09711、(151) K23375、(152) K16389、
    (153) K01733、(154) K15599、(155) K03498、
    (156) K10542、(157) K01216、(158) K08303、
    (159) K08641、(160) K03274
    請求項7または8に記載の検査方法。
  16.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記長期入院の有無であり、
     前記特徴量は、前記腸内菌叢の遺伝子組成データから抽出された下記(201)~(220)に示す遺伝子である、
    (201) K01463、(202) K14188、(203) K07457、
    (204) K22958、(205) K15372、(206) K07069、
    (207) K08153、(208) K14654、(209) K06156、
    (210) K03734、(211) K07505、(212) K01482、
    (213) K03328、(214) K01635、(215) K07217、
    (216) K19973、(217) K03753、(218) K23245、
    (219) K16203、(220) K05770
    請求項7または8に記載の検査方法。
  17.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記重症化の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(301)~(320)に示す血液代謝産物である、
    (301) フェニルアラニン、(302) マンノース、
    (303) 2-ヒドロキシ酪酸、(304) イノシトール、
    (305) 4-ヒドロキシフェニル乳酸、(306) 3-フェニル乳酸、
    (307) チロシン、(308) グリシン、(309) シュウ酸、
    (310) 2-アミノエタノール、(311) トレオン酸、(312) グリセロール、
    (313) 尿素、(314) 2-ヒドロキシイソ吉草酸、(315) L-カルニチン、
    (316) バリン、(317) グルクロン酸、(318) ジメチルグリシン、
    (319) サルコシン、(320) グルコース
    請求項7または8に記載の検査方法。
  18.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての呼吸器症状の発現の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(321)~(335)に示す血液代謝産物である、
    (321) 3-ヒドロキシイソ酪酸、(322) プロピオン酸、
    (323) 2-ヒドロキシイソ酪酸、(324) イノシトール、
    (325) 2-ケトグルタル酸、(326) グリオキシル酸、(327) ロイシン、
    (328) 5-オキソプロリン、(329) フェニルアラニン、(330) バリン、
    (331) グルタミン酸、(332) 3-アミノイソ酪酸、
    (333) 2-アミノイソ酪酸、(334) ノナン酸、(335) グルクロン酸
    請求項7または8に記載の検査方法。
  19.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての肺炎の発現の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(336)~(355)に示す血液代謝産物である、
    (336) エリトルロース、(337) イノシトール、(338) マンニトール、
    (339) ギ酸、(340) セリン、(341) グリシン、
    (342) マンノース、(343) フルクトース、(344) ヒスチジン、
    (345) 2-ケトグルタル酸、(346) インドール-3-酢酸、
    (347) trans-アコニット酸、(348) チロシン、(349) リブロ―ス、
    (350) グルコース、(351) 2-アミノアジピン酸、
    (352) ガラクツロン酸、(353) ウラシル、
    (354) 3-ヒドロキシグルタル酸、(355) セロトニン
    請求項7または8に記載の検査方法。
  20.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての下痢の発現の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(356)~(361)に示す血液代謝産物である、
    (356) プロリン、(357) アスパラギン、(358) クレアチン、
    (359) インドール-3-酢酸、(360) フェニルアラニン、
    (361) グリセリン酸
    請求項7または8に記載の検査方法。
  21.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての肝障害の発症の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(362)~(380)に示す血液代謝産物である、
    (362) 2-ケトグルタル酸、(363) グルタミン酸、(364) ヒスチジン、
    (365) バリン、(366) アスパラギン、(367) エリトルロース、
    (368) クエン酸、(369) イソ酪酸、(370) グルタミン、
    (371) マンニトール、(371) 2-アミノピメリン酸、(372) リジン、
    (373) インドール-3-酢酸、(374) コハク酸、(375) ロイシン、
    (376) 2-ケトイソカプロン酸、(377) グリシン、
    (378) 2-ヒドロキシイソ吉草酸、(379) キシルロース、(380) 乳酸
    請求項7または8に記載の検査方法。
  22.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての腎障害の発症の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(381)~(400)に示す血液代謝産物である、
    (381) イノシトール、(382) 尿素、(383) クレアチニン、
    (384) アラビノース、(385) 4-ヒドロキシフェニル乳酸、
    (386) シスチン、(387) グルタミン酸、(388) エリスリトール、
    (389) トリプトファン、(390) マンニトール、(391) グルクロン酸、
    (392) アラビトール、(393) 2-ヒドロキシイソ酪酸、
    (394) 2-ヒドロキシ酪酸、(395) セリン、(396) コリン、
    (397) プロピオン酸、(398) トリメチルアミンオキシド、
    (399) 2-ケトグルタル酸、(400) グリセリン酸
    請求項7または8に記載の検査方法。
  23.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としてのD-dimerの1.0mg/ml以上への上昇の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(418)~(433)に示す血液代謝産物である、
    (418) マルトース、(419) トリプトファン、(420) チロシン、
    (421) クエン酸、(422) シュウ酸、(423) クレアチン、
    (424) 2-ヒドロキシイソ吉草酸、(425) トレオン酸、
    (426) サルコシン、(427) 尿酸、(428) ラムノース、
    (429) グルカル酸、(430) 4-ヒドロキシプロリン、
    (431) イソ吉草酸、(432) ヒスチジン、(433) グルクロン酸
    請求項7または8に記載の検査方法。
  24.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記長期入院の有無であり、
     前記特徴量は、前記血液代謝産物データから抽出された下記(461)~(480)に示す血液代謝産物である、
    (461) エリトルロース、(462) コハク酸、(463) マンノース、
    (464) メチオニンスルホキシド、(465) マルトース、
    (466) マンニトール、(467) インドール-3-酢酸、(468) シュウ酸、
    (469) アスパラギン、(470) システイン、(471) L-カルニチン、
    (472) 尿素、(473) プロピオン酸、(474) グルタミン、
    (475) イノシトール、(476) 2-ケトグルタル酸、(477) 乳酸、
    (478) グルコース、(479) グリセロール、(480) デカン酸
    請求項7または8に記載の検査方法。
  25.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての呼吸器症状の発現の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(521)~(540)に示す腸内細菌である、
    (521) Bacteroides species incertae sedis (ext_mOTU_v26_26380)、
    (522) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、
    (523) Bifidobacterium breve (ref_mOTU_v25_01098)、
    (524) Eggerthella lenta (ref_mOTU_v25_00719)、
    (525) Sutterella wadsworthensis (ref_mOTU_v25_03066)、
    (526) Actinomyces graevenitzii (ref_mOTU_v25_04054)、
    (527) Collinsella species incertae sedis (ext_mOTU_v26_17345)、
    (528) Acidaminococcus intestini (ref_mOTU_v25_01949)、
    (529) Bifidobacterium adolescentis (ref_mOTU_v25_02703)、
    (530) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、
    (531) Veillonella atypica (ref_mOTU_v25_01941)、
    (532) Bifidobacterium longum (ref_mOTU_v25_01099)、
    (533) Gemella sanguinis (ref_mOTU_v25_04303)、
    (534) Streptococcus parasanguinis (ref_mOTU_v25_00312)、
    (535) Veillonella species incertae sedis (meta_mOTU_v25_13135)、
    (536) Actinomyces sp (ref_mOTU_v25_01914)、
    (537) Bifidobacterium bifidum (ref_mOTU_v25_03116)、
    (538) Lactobacillus fermentum/oris (ref_mOTU_v25_01407)、
    (539) Bacteroides sp. (ref_mOTU_v25_03475)、
    (540) Ruminococcus bromii (ref_mOTU_v25_00853)
    請求項7または8に記載の検査方法。
  26.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての肺炎の発現の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(541)~(560)に示す腸内細菌である、
    (541) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、
    (542) Firmicutes species incertae sedis (meta_mOTU_v25_12923)、
    (543) Eggerthella lenta (ref_mOTU_v25_00719)、
    (544) Streptococcus cristatus (ref_mOTU_v25_03967)、
    (545) Bacteroides caccae (ref_mOTU_v25_03473)、
    (546) Bacteroides species incertae sedis (ext_mOTU_v26_17504)、
    (547) Rothia mucilaginosa (ref_mOTU_v25_05265)、
    (548) Streptococcus species incertae sedis (ext_mOTU_v26_28826)、
    (549) Actinomyces sp (ref_mOTU_v25_12049)、
    (550) Lachnospiraceae species incertae sedis (meta_mOTU_v25_12240)、
    (551) Escherichia coli (ref_mOTU_v25_00095)、
    (552) Acidaminococcus intestini (ref_mOTU_v25_01949)、
    (553) Clostridiales sp. (ref_mOTU_v25_03444)、
    (554) Tyzzerella nexilis (ref_mOTU_v25_03689)、
    (555) Streptococcus thermophilus (ref_mOTU_v25_01348)、
    (556) Bacteroides rodentium/uniformis (ref_mOTU_v25_00855)、
    (557) Firmicutes species incertae sedis (meta_mOTU_v25_12227)、
    (558) Clostridiales species incertae sedis (ext_mOTU_v26_18011)、
    (559) Lachnospiraceae species incertae sedis (ext_mOTU_v26_16388)、
    (560) Isoptericola variabilis (ref_mOTU_v25_01910)
    請求項7または8に記載の検査方法。
  27.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記症状としての下痢の発現の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(561)~(580)に示す腸内細菌である、
    (561) Streptococcus species incertae sedis (ext_mOTU_v26_28879)、
    (562) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、
    (563) Streptococcus australis (ref_mOTU_v25_00311)、
    (564) Tyzzerella species incertae sedis (meta_mOTU_v25_12266)、
    (565) Holdemanella biformis (meta_mOTU_v25_12329)、
    (566) [Ruminococcus] torques (ref_mOTU_v25_03703)、
    (567) Anaerotignum lactatifermentans (ref_mOTU_v25_02190)、
    (568) Butyricicoccus sp. (ref_mOTU_v25_02967)、
    (569) Eubacterium species incertae sedis (ext_mOTU_v26_16240)、
    (570) Bacteroides dorei/vulgatus (ref_mOTU_v25_02367)、
    (571) Actinobacteria sp. (ref_mOTU_v25_01911)、
    (572) Streptococcus parasanguinis (ref_mOTU_v25_00312)、
    (573) [Clostridium] clostridioforme/bolteae (ref_mOTU_v25_03442)、
    (574) Ruminococcaceae species incertae sedis (ext_mOTU_v26_16263)、
    (575) Erysipelatoclostridium ramosum (ref_mOTU_v25_03439)、
    (576) Clostridiales species incertae sedis (meta_mOTU_v25_13392)、
    (577) Provencibacterium massiliense (ref_mOTU_v25_10126)、
    (578) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、
    (579) Clostridium sp (ref_mOTU_v25_09167)、
    (580) Sutterella wadsworthensis (ref_mOTU_v25_03066)
    請求項7または8に記載の検査方法。
  28.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての肝障害の発症の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(581)~(600)に示す腸内細菌である、
    (581) Eggerthella lenta (ref_mOTU_v25_00719)、
    (582) Eubacterium sp (meta_mOTU_v25_12688)、
    (583) Clostridium sp (meta_mOTU_v25_12609)、
    (584) Clostridiales species incertae sedis (ext_mOTU_v26_26595)、
    (585) Bacteroides sp. (ref_mOTU_v25_03475)、
    (586) Clostridiales species incertae sedis (meta_mOTU_v25_13012)、
    (587) Megasphaera sp (ref_mOTU_v25_03433)、
    (588) Bacteroides caecimuris (ref_mOTU_v25_03476)、
    (589) Bacteroides faecis/thetaiotaomicron (ref_mOTU_v25_01657)、
    (590) Lactococcus lactis (ref_mOTU_v25_01300)、
    (591) Clostridiales Family XIII (ext_mOTU_v26_16402)、
    (592) Lachnospiraceae species incertae sedis (ext_mOTU_v26_26654)、
    (593) Mogibacterium timidum (ref_mOTU_v25_04269)、
    (594) [Eubacterium] hallii (ref_mOTU_v25_03632)、
    (595) Coprococcus sp. (ref_mOTU_v25_01683)、
    (596) Bacteroides species incertae sedis (ext_mOTU_v26_26291)、
    (597) Eggerthellaceae species incertae sedis (ext_mOTU_v26_15442)、
    (598) uncultured Flavonifractor sp. (ref_mOTU_v25_07315)、
    (599) Clostridium species incertae sedis (ext_mOTU_v26_17379)、
    (600) Collinsella aerofaciens (ref_mOTU_v25_03626)
    請求項7または8に記載の検査方法。
  29.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としての腎障害の発症の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(601)~(620)に示す腸内細菌である、
    (601) Clostridiales species incertae sedis (meta_mOTU_v25_13006)、
    (602) Streptococcus anginosus (ref_mOTU_v25_00569)、
    (603) Streptococcus intermedius/constellatus (ref_mOTU_v25_00572)、
    (604) Actinomyces marseillensis/pacaensis (ref_mOTU_v25_03846)、
    (605) Firmicutes bacterium CAG:114 (ref_mOTU_v25_07728)、
    (606) Methanobrevibacter smithii (ref_mOTU_v25_03695)、
    (607) Clostridiales sp. (ref_mOTU_v25_03661)、
    (608) Megamonas funiformis/rupellensis (ref_mOTU_v25_02318)、
    (609) Streptococcus anginosus (ref_mOTU_v25_00570)、
    (610) Bacteroides species incertae sedis (ext_mOTU_v26_18132)、
    (611) Blautia massiliensis (ref_mOTU_v25_03342)、
    (612) Collinsella species incertae sedis (ext_mOTU_v26_17328)、
    (613) Bacteroides plebeius (ref_mOTU_v25_05069)、
    (614) Rothia dentocariosa (ref_mOTU_v25_04800)、
    (615) Clostridiales species incertae sedis (meta_mOTU_v25_12635)、
    (615) Anaerotruncus colihominis (ref_mOTU_v25_03438)、
    (617) Staphylococcaceae sp. (ref_mOTU_v25_04195)、
    (618) Streptococcus oralis (ref_mOTU_v25_00289)、
    (619) Alistipes finegoldii (ref_mOTU_v25_03682)、
    (620) Firmicutes species incertae sedis (meta_mOTU_v25_14410)
    請求項7または8に記載の検査方法。
  30.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記合併症としてのD-dimerの1.0mg/ml以上への上昇の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(641)~(660)に示す腸内細菌である、
    (641) Bifidobacterium pseudocatenulatum (ref_mOTU_v25_02700)、
    (642) Bacteroides species incertae sedis (ext_mOTU_v26_17504)、
    (643) Akkermansia muciniphila (ref_mOTU_v25_03591)、
    (644) uncultured Eubacterium sp (meta_mOTU_v25_13063)、
    (645) Catabacter hongkongensis (ref_mOTU_v25_06126)、
    (646) Lactobacillus plantarum (ref_mOTU_v25_00930)、
    (647) Rothia dentocariosa (ref_mOTU_v25_04800)、
    (648) Slackia exigua (ref_mOTU_v25_01958)、
    (649) Enterococcus faecium/durans (ref_mOTU_v25_00323)、
    (650) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、
    (651) Streptococcus oralis/pseudopneumoniae (ref_mOTU_v25_00296)、
    (652) Bacteroides cellulosilyticus/fragilis (ref_mOTU_v25_01597)、
    (653) Blautia obeum/wexlerae (ref_mOTU_v25_02154)、
    (654) Bacteroides sp. (ref_mOTU_v25_03475)、
    (655) Hungatella hathewayi (ref_mOTU_v25_03436)、
    (656) Staphylococcus aureus (ref_mOTU_v25_00340)、
    (657) [Ruminococcus] gnavus (ref_mOTU_v25_01594)、
    (658) Lactobacillus casei/paracasei (ref_mOTU_v25_01406)、
    (659) Butyricicoccus sp. (ref_mOTU_v25_02967)、
    (660) Alistipes species incertae sedis (meta_mOTU_v25_12829)
    請求項7または8に記載の検査方法。
  31.  前記目的変数は、前記罹患者について観察された、COVID-19に罹患した際の、前記長期入院の有無であり、
     前記特徴量は、前記腸内菌叢の細菌組成データから抽出された下記(701)~(720)に示す腸内細菌である、
    (701) Ruthenibacterium lactatiformans (ref_mOTU_v25_04716)、
    (702) Blautia massiliensis (ref_mOTU_v25_03342)、
    (703) Streptococcus anginosus/intermedius (ref_mOTU_v25_00567)、
    (704) Massilioclostridium coli (ref_mOTU_v25_10237)、
    (705) Eggerthella lenta (ref_mOTU_v25_00719)、
    (706) Clostridiales Family XIII (ext_mOTU_v26_16402)、
    (707) Clostridiales sp. (ref_mOTU_v25_04568)、
    (708) Firmicutes sp. (ref_mOTU_v25_02743)、
    (709) Bacteroides plebeius (ref_mOTU_v25_05069)、
    (710) Eisenbergiella tayi (ref_mOTU_v25_03446)、
    (711) Phascolarctobacterium succinatutens (ref_mOTU_v25_03700)、
    (712) Granulicatella species incertae sedis (ext_mOTU_v26_19463)、
    (713) Faecalibacterium prausnitzii (ref_mOTU_v25_06112)、
    (714) Veillonella dispar (ref_mOTU_v25_01940)、
    (715) Actinomyces graevenitzii (ref_mOTU_v25_04054)、
    (716) Pseudoflavonifractor species incertae sedis (meta_mOTU_v25_14283)、
    (717) Bacteroides faecis/thetaiotaomicron (ref_mOTU_v25_01657)、
    (718) Dorea longicatena (ref_mOTU_v25_03692)、
    (719) [Clostridium] leptum (ref_mOTU_v25_03688)、
    (720) Blautia obeum/wexlerae (ref_mOTU_v25_02154)
    請求項7または8に記載の検査方法。
  32.  被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを取得する取得部と、
     取得された前記腸内菌叢の遺伝子組成データおよび前記血液代謝産物データの少なくとも一方を、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測し、または、取得された前記腸内菌叢の細菌組成データを、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する予測部と、
    を備える、
    検査装置。
  33.  コンピュータに、
     被験者から採取された生体試料中の、腸内菌叢の遺伝子組成データ、血液代謝産物データ、および、腸内菌叢の細菌組成データの少なくとも1つを取得する取得ステップと、
     取得された前記腸内菌叢の遺伝子組成データおよび前記血液代謝産物データの少なくとも一方を、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測し、または、取得された前記腸内菌叢の細菌組成データを、予め準備しておいた予測モデルに入力することで、COVID-19に罹患した際の症状発現リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを予測する予測ステップと、
    を実行させる、
    検査プログラム。
PCT/JP2022/048200 2022-03-22 2022-12-27 Covid-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム Ceased WO2023181572A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2024509773A JPWO2023181572A1 (ja) 2022-03-22 2022-12-27

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2022045879 2022-03-22
JP2022-045879 2022-03-22

Publications (1)

Publication Number Publication Date
WO2023181572A1 true WO2023181572A1 (ja) 2023-09-28

Family

ID=88100918

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2022/048200 Ceased WO2023181572A1 (ja) 2022-03-22 2022-12-27 Covid-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム

Country Status (2)

Country Link
JP (1) JPWO2023181572A1 (ja)
WO (1) WO2023181572A1 (ja)

Non-Patent Citations (6)

* Cited by examiner, † Cited by third party
Title
BUYUKOZKAN MUSTAFA, ALVAREZ-MULETT SERGIO, RACANELLI ALEXANDRA C., SCHMIDT FRANK, BATRA RICHA, HOFFMAN KATHERINE L., SARWATH HINA,: "Integrative Metabolomic and Proteomic Signatures Define Clinical Outcomes in Severe COVID-19", MEDRXIV, 11 February 2022 (2022-02-11), XP093094885, Retrieved from the Internet <URL:https://www.medrxiv.org/content/10.1101/2021.07.19.21260776v2.full.pdf> [retrieved on 20231025], DOI: 10.1101/2021.07.19.21260776 *
LI SIJIA, YANG SIYUAN, ZHOU YUZHENG, DISOMA CYROLLAH, DONG ZIJUN, DU ASHUAI, ZHANG YONGXING, CHEN YONG, HUANG WEILIANG, CHEN JUNRU: "Microbiome Profiling Using Shotgun Metagenomic Sequencing Identified Unique Microorganisms in COVID-19 Patients With Altered Gut Microbiota", FRONTIERS IN MICROBIOLOGY, vol. 12, XP093094881, DOI: 10.3389/fmicb.2021.712081 *
LIU YIJUN, ZHANG HONGYANG, TANG XIAOJUN, JIANG XUEJUN, YAN XIAOJUAN, LIU XIZHAO, GONG JIANG, MEW KENLEY, SUN HAO, CHEN XIUFENG, ZO: "Distinct Metagenomic Signatures in the SARS-CoV-2 Infection", FRONTIERS IN CELLULAR AND INFECTION MICROBIOLOGY, vol. 11, XP093094880, DOI: 10.3389/fcimb.2021.706970 *
ROBERTS, IVAYLA ET AL.: "Untargeted metabolomics of COVID-19 patient serum reveals potential prognostic markers of both severity and outcome", METABOLOMICS, vol. 18, 2021, pages 6, XP037647119, DOI: 10.1007/sll306-021-01859-3> *
SCHULT DAVID, REITMEIER SANDRA, KOYUMDZHIEVA PLAMENA, LAHMER TOBIAS, MIDDELHOFF MORITZ, ERBER JOHANNA, SCHNEIDER JOCHEN, KAGER JUL: "Gut bacterial dysbiosis and instability is associated with the onset of complications and mortality in COVID-19", GUT MICROBES, LANDES BIOSCIENCE, UNITED STATES, vol. 14, no. 1, 31 December 2022 (2022-12-31), United States , XP093094888, ISSN: 1949-0976, DOI: 10.1080/19490976.2022.2031840 *
SINDELAR MIRIAM, STANCLIFFE ETHAN, SCHWAIGER-HABER MICHAELA, ANBUKUMAR DHANALAKSHMI S., ADKINS-TRAVIS KAYLA, GOSS CHARLES W., O’HA: "Longitudinal metabolomics of human plasma reveals prognostic markers of COVID-19 disease severity", CELL REPORTS MEDICINE, vol. 2, no. 8, 1 August 2021 (2021-08-01), pages 100369, XP093094887, ISSN: 2666-3791, DOI: 10.1016/j.xcrm.2021.100369 *

Also Published As

Publication number Publication date
JPWO2023181572A1 (ja) 2023-09-28

Similar Documents

Publication Publication Date Title
Wahl et al. Epigenome-wide association study of body mass index, and the adverse outcomes of adiposity
Shen et al. Metagenomic sequencing of bile from gallstone patients to identify different microbial community patterns and novel biliary bacteria
Ma et al. Metagenome analysis of intestinal bacteria in healthy people, patients with inflammatory bowel disease and colorectal cancer
Galkin et al. Human microbiome aging clocks based on deep learning and tandem of permutation feature importance and accumulated local effects
CN107075563B (zh) 用于冠状动脉疾病的生物标记物
Margiotta et al. Gut microbiota composition and frailty in elderly patients with chronic kidney disease
Hoyles et al. Molecular phenomics and metagenomics of hepatic steatosis in non-diabetic obese women
Qin et al. Alterations of the human gut microbiome in liver cirrhosis
Bonder et al. The influence of a short-term gluten-free diet on the human gut microbiome
Kicic et al. Assessing the unified airway hypothesis in children via transcriptional profiling of the airway epithelium
Palmer et al. Concordance between gene expression in peripheral whole blood and colonic tissue in children with inflammatory bowel disease
Sun et al. The gut microbiota heterogeneity and assembly changes associated with the IBD
CN107075453B (zh) 冠状动脉疾病的生物标记物
Tang et al. Integrated analysis of biopsies from inflammatory bowel disease patients identifies SAA1 as a link between mucosal microbes with TH17 and TH22 cells
CN114438165B (zh) 针对稳定型冠心病的急性冠脉综合征风险评估标志物及应用
CN112509701A (zh) 急性冠脉综合征的风险预测方法及装置
Sinha et al. Maternal antibiotic prophylaxis during cesarean section has a limited impact on the infant gut microbiome
Jollet et al. Insight into the role of gut microbiota in Duchenne muscular dystrophy: an age-related study in mdx mice
Knudsen et al. The lower airways microbiome and antimicrobial peptides in idiopathic pulmonary fibrosis differ from chronic obstructive pulmonary disease
Ezzeldin et al. Current understanding of human metaproteome association and modulation
Zuffa et al. A multi-organ Murine metabolomics atlas reveals molecular dysregulations in Alzheimer’s Disease
Shimizu et al. Dysbiosis of gut microbiota in patients with severe COVID‐19
Shaw et al. Assessing the colonic microbiota in children: effects of sample site and bowel preparation
Danhaive et al. Pulmonary hypertension in developmental lung diseases
WO2023181572A1 (ja) Covid-19に罹患した際の症状発現リスク、重症化リスク、合併症の発症リスク、及び長期入院リスクよりなる群から選択される少なくとも1つを検査する方法、検査装置、および、検査プログラム

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22933699

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2024509773

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 22933699

Country of ref document: EP

Kind code of ref document: A1