EP4630813A2 - Improved method for scalable untargeted metabolomic workflow - Google Patents
Improved method for scalable untargeted metabolomic workflowInfo
- Publication number
- EP4630813A2 EP4630813A2 EP23901396.4A EP23901396A EP4630813A2 EP 4630813 A2 EP4630813 A2 EP 4630813A2 EP 23901396 A EP23901396 A EP 23901396A EP 4630813 A2 EP4630813 A2 EP 4630813A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- sample
- features
- samples
- relevant
- data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H01—ELECTRIC ELEMENTS
- H01J—ELECTRIC DISCHARGE TUBES OR DISCHARGE LAMPS
- H01J49/00—Particle spectrometers or separator tubes
- H01J49/0027—Methods for using particle spectrometers
- H01J49/0036—Step by step routines describing the handling of the data generated during a measurement
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/10—Signal processing, e.g. from mass spectrometry [MS] or from PCR
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H10/00—ICT specially adapted for the handling or processing of patient-related medical or healthcare data
- G16H10/40—ICT specially adapted for the handling or processing of patient-related medical or healthcare data for data related to laboratory analysis, e.g. patient specimen analysis
Definitions
- One aspect is a method of non-targeted determination of unique biological metabolites in individual samples of a sample set, each individual sample comprising a composition of chemical constituents, the method comprising: a. analyzing reference features obtained from a reference sample to identify non-relevant reference features; b. filtering the reference features by removing the non-relevant reference features, thereby producing a set of relevant reference features that characterize the unique biological metabolites; and, c. applying the non-relevant and/or relevant reference features to sample features obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample.
- the method may comprise conducting the steps of analyzing, filtering and applying on each individual sample in the sample set, thereby determining the unique biological metabolites in all individuals of the sample set.
- the reference features are updated by re- analyzing the reference sample to obtain updated reference features the reference features are replaced with the updated reference features.
- the number of individual samples in the plurality of samples may be less than the entire number of individual samples in the sample set.
- the plurality of samples may comprise at least 2, samples, optionally at least 3 samples, optionally at least 5 samples, optionally at least 10 samples, optionally at least 20 samples, optionally at least 50 samples, optionally at least 100 samples, or optionally at least 150 samples.
- the plurality of samples my comprise at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 90%, or at least 95% of the samples in the sample set.
- the reference sample may comprise an aliquot from each of a plurality of the individual samples.
- the reference features may be obtained by a method comprising subjecting the reference sample to a separation technique and mass spectrometry.
- the separation technique may comprise chromatography, which may comprise gas chromatography or liquid chromatography.
- the non-relevant reference features may characterize artifacts, background chemical constituents and/or non-unique biological metabolites.
- the background chemical constituents may comprise contaminants.
- the non-unique biological metabolites may comprise adducts and/or fragment ions.
- the reference features may comprise reference separation data and reference mass data generated from the reference sample.
- the step of applying may comprise: identifying in the sample features those sample features corresponding to non-relevant reference features and removing those sample features that correspond to the non-relevant reference features to produce a set of relevant sample features; and, using the relevant sample features to determine the unique biological metabolites in the sample.
- the step of applying may comprise: identifying in the sample features those sample features corresponding to relevant reference features to produce a set of relevant sample features; and, using the relevant sample features to determine the unique biological metabolites in the sample.
- the method may comprise conducting the steps of obtaining and applying to each sample in the sample set, thereby determining the composition of the chemical constituents in all individual samples in the sample set.
- the method may comprise identifying one or more of the chemical constituents by comparing the relevant sample features to a library of information comprising features characterizing chemical entities.
- One aspect comprises a method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a.
- One aspect of the disclosure comprises a method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a. using a separation technique and mass spectrometry to produce reference features comprising separation data and mass data from a reference sample; b. collecting and storing the reference features; c.
- the step of using the filtered reference features may comprise: obtaining sample features from an individual sample; identifying sample features that correspond to the filtered reference features to produce a list of relevant sample features that characterize unique chemical constituents, thereby determining the composition of chemical constituents in the sample.
- a separation apparatus for separating the chemical constituents and producing separation-related features characterizing the chemical constituents
- a mass spectrometer for performing mass spectrometry on portions of the separated chemical constituents and producing mass spectrometry-related features characterizing the chemical constituents
- a first module for receiving, collecting and/or storing the separation-related features and/or the mass spectrometry -related features
- a user interface coupled to the module for making the separation-related features and/or the mass spectrometry -related features available in human- accessible form.
- the system may comprise a library of features characterizing chemical entities, the features produced using the separation apparatus and the mass spectrometer, wherein the library of features comprises separation features and mass spectrometry features characterizing identified chemical entities.
- the separation apparatus may comprise an apparatus for performing chromatography, which may comprise liquid chromatography or gas chromatography. Liquid chromatography may comprise HPLC or UPLC.
- the separation apparatus may comprise an electrophoresis apparatus, which may be coupled to the mass spectrometer.
- the separation-related features may comprise peak retention time, peak intensity, and/or peak width.
- the mass-spectrometry-related features may comprise mass or an m/z value.
- FIGS. 2 a & b show total ion chromatograms.
- FIG. 3 illustrates a pipeline for handling metabolomics data.
- Polar and lipid metabolites are extracted from plasma samples into 96-well plates for LC/MS analysis.
- a pooled sample is prepared for feature detection, MS/MS acquisition, and use as a QC sample.
- Untargeted metabolomics analysis is performed on all samples. After detecting features from the pooled sample, background features and degeneracies are filtered.
- FIGS. 4a & 4b show data pre-processing of a single batch.
- FIGS.5a-5c show retention time deviations in lipid metabolites.
- FIGS. 6a-6e shows a comparison between individual and pooled samples. a) Venn diagram showing the breakdown of features detected in at least one individual sample (red) versus in at least one pooled sample (green). A higher fraction of features missed in the pooled sample had no hits returned when searching the Human Metabolome Database and KEGG. b-c) Histograms showing the fraction of samples where a feature was detected.
- FIGS. 7a-7c show targeted extraction of peak areas versus traditional global peak integration.
- a-b) The pie charts represent the percentage of all measurements that are missing values (grey) when targeted extraction of peak areas with Skyline is performed (a) or when traditional peak integration within XCMS is performed (b). In (a) the missing value percentage is too small to be visible on the pie chart.
- FIGS.8a-8e illustrate correcting for batch effects in metabolomics data.
- PCA Principal components analysis
- a) Principal components analysis (PCA) of unnormalized lipid metabolic profiles shows strong batch effects. Each dot represents a unique sample. Dots are colored according to their corresponding batch number.
- PCA Principal components analysis
- FIG.9a-9e show correcting for batch effects in metabolomics data.
- PCA Principal components analysis
- Normalization score is the change in coefficient of variation (CV) in the research samples (relative to the unnormalized data) divided by the change in CV for the QC samples. A higher score indicates a reduction of technical variation.
- Violin plots showing the CV distribution of all compounds in the QC samples for each evaluated batch-correction algorithm. The lipid metabolite counterpart to these data is shown in FIG.8.
- FIGS. 10a & 10b show that QC samples enable discrimination of batch and biological effects. Metabolite intensities can change as a function of analysis batch. If batches are biased by biologically different sample groups, then the application of non-QC based batch normalization can remove biological variation in addition to technical variation because the different types of variation cannot be easily distinguished. However, QC samples enable technical shifts caused by batch number (such as DGTS 17:0) to be differentiated from biological differences between batches (such as inosine).
- the plots in (a) show the unnormalized metabolite intensities as a function of run order for QC samples (left) and research samples (right). In (b) the metabolite intensities after random forest correction are shown for QC and research samples for the same compounds.
- FIGS.11a-11d show internal standard variability.
- a-b Violin plots showing the distribution of coefficients of variation (CV) for the internal standards across all samples within a batch or across all batches (1-22) for both unnormalized (a) and random forest corrected data (b).
- c-d Principal components analysis (PCA) of internal standard intensities across all samples in the unnormalized (c) and random forest corrected data (d). Each dot is a sample. Samples are colored by batch.
- Violin plots showing the distribution of CVs for the internal standards across all samples for the 14 batch correction methods evaluated in this study.
- FIGS.12a-12d show that random forest normalization reduces batch effects in QC samples.
- PCA Principal components analysis
- FIG.15 shows that metabolites associated with geographic location are not reflective of subject age. Principal components analysis (PCA) of normalized metabolic profiles (polar and lipid metabolites) shows no age-dependent pattern in geographically associated metabolites. Each dot represents a unique sample.
- PCA Principal components analysis
- the present disclosure relates to an improved method for conducting a metabolomics workflow. More specifically, the present disclosure relates to a metabolomics workflow that reduces the computational burden required to analyze the abundance of compounds present in a sample, thereby allowing the process to be scaled for analysis of a larger number of samples than is possible with currently available techniques.
- the disclosed process achieves this scalability using a reference sample that represents the chemical complexity of an entire sample set. Analysis of the refence sample allows detection and identification of thousands of features, most of which do not characterize unique biological metabolites, within a single sample.
- the features may then be filtered to remove non-relevant features, which characterize non-relevant chemical constituents, yielding a list of relevant features.
- This list of relevant features may then be used to identify relevant features characterizing unique-biological metabolites within each individual sample, thereby reducing the computational burden necessary to analyze the sample set.
- a method of the disclosure may generally be practiced by analyzing a reference sample to obtain biologically relevant reference features and using these relevant reference features to identify relevant features in one or more individual samples in the sample set.
- the relevant features may also be quantitated, thereby producing a metabolomic fingerprint of the biological metabolites in the individual sample. In some relevant features and/or the metabolic fingerprint may be used to identify the individual biological metabolites present in the individual sample.
- the relevant features may be assessed in each individual sample in the sample set, serially or sequentially, to determine the metabolomic fingerprint for many or all the individual samples in the sample set.
- a metabolite refers to one or more metabolites.
- the terms “a”, “an”, “one or more” and “at least one” can be used interchangeably.
- the terms “comprising”, “including” and “having” can be used interchangeably.
- the term “comprising” may be replaced with “consisting” or with “consisting essentially of” in particular aspect, as desired.
- the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements or use of a “negative” limitation.
- One aspect of the disclosure is a method of non-targeted determination of unique biological metabolites in individual samples of a sample set, each individual sample comprising a composition of chemical constituents, the method comprising analyzing reference features obtained from a reference sample to identify non- relevant reference features; filtering the reference features by removing the non-relevant reference features, thereby producing a set of relevant reference features that characterize the unique biological metabolites; and, applying the non-relevant and/or relevant reference features to sample features obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample.
- non-targeted determination means characterizing the chemical composition of a complex sample containing chemical constituents without prior knowledge of one or more chemical constituents in the sample.
- chemical constituents refers to the molecules present within a sample.
- Chemical constituents encompasses any molecule present within a sample including but not limited to, metabolites, including endogenous metabolites, exogenous metabolites, unique metabolites, non-unique metabolites and contaminants.
- metabolite and “biological metabolite”, may be used interchangeably and refer to the set of small molecules comprising substrates, intermediates, and products of metabolism.
- Metabolites of the disclosure are generally less than about 2000 kDa in size, optionally less than about 1500 kDa in size. Metabolites include both endogenous metabolites (i.e., produced by an individual) and exogenous metabolites (i.e., drugs, environmental toxins, etc.). Examples of endogenous metabolites include, but are not limited to, organic acids, fatty acids, triglycerides, cholesterol, phospholipids, sugars, vitamins, and co-factors. It is understood by those of skill in the art that analysis of a metabolite using a technique such as mass spectrometry results in the production of numerous species of metabolite ions.
- the different species result from such things as fragmentation of the metabolite, incorporation into the metabolite of isotopes such as carbon- 13 or nitrogen-15, and binding of an ion species to such things salts or solvent, to form adducts. This may result in the production of several signals, each signal coming from a different species of the metabolite.
- one single ion species is chosen to represent each metabolite.
- a unique metabolite refers to a select, monoisotopic, single ion species of an intact metabolite.
- the term “unique metabolite” excludes fragmented species of metabolites, metabolites comprising naturally abundant stable isotopes, such as carbon-13 or nitrogen-15, and ion species of the metabolite other than the selected ion species, such as other adducts of the metabolite.
- a non-unique metabolite refers to a fragment of a metabolite, a non- monoisotopic form of the selected metabolite and additional adducts of the selected ion species of the metabolite.
- the term back “background” refers to features characterizing contaminants and artifacts.
- contaminant refers to chemicals present in the sample but that originate from outside of the sample (e.g., from the equipment used to analyze the sample).
- a tube holding a sample may contain plasticizers that may leech into the sample causing the sample to contain plasticizer.
- Plasticizer in the sample represents a contaminate.
- an “artifact” refers to features that do not characterize real chemicals, but that arise from electronic noise in the mass spectrometer or errors in downstream peak detection software. Methods of identifying contaminants and artifacts are known to those skilled in the art and are also disclosed herein.
- the term “feature” refers to the one or more points of data, such as a peak retention time or mass-to-charge (m/z) value, that characterize a chemical constituent.
- a feature obtained using a separation technique such as chromatography (e.g., HPLC)
- retention time rt
- mass spectrometry may comprise a m/z value.
- a feature obtained using a separation technique (e.g., chromatography) and mass spectrometry may comprise both a rt and a m/z value.
- a “relevant feature” is a feature that characterizes a unique metabolite.
- a “non-relevant feature” is a feature that characterizes non-unique metabolites, contaminants, and that are artifacts.
- Features (e.g., reference features, sample features, etc.) of the disclosure are obtained from samples (e.g., reference samples, individual samples) of the disclosure using various methods.
- features are obtained by subjecting a sample to a separation technique. Any separation that suitably separates chemical constituents within the sample may be used. Examples of separation techniques include, but are not limited to, chromatography and electrophoresis, including capillary electrophoresis.
- Suitable methods of chromatography include, but are not limited to, gas chromatography (GC), high performance liquid chromatography (HPLC), or ultra-performance liquid chromatography (UPLC).
- features are obtained by subjecting the sample to mass spectrometry. In certain aspects, features are obtained by subjecting the sample to chromatography, such as HPLC or UPLC, and mass spectrometry. Accordingly, in certain aspects, a feature comprises a rt and a m/z value.
- biological sample refers to a sample obtained from an individual, including a sample of biological tissue or fluid origin obtained in vivo or in vitro. Such samples may include, but are not limited to, blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluids, dialysates, lavage fluids, and fluids derived from organs or tissue.
- sample refers to a biological sample obtained from a single individual.
- sample set refers to a defined plurality of individual samples.
- Examples include, but are not limited to, humans and other primates, including non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cattle, sheep, pigs, seals, goats and horses; domestic mammals such as dogs and cats; laboratory animals including rodents such as mice, rats and guinea pigs; birds, including domestic, wild and game birds such as chickens, turkeys and other gallinaceous birds, ducks, geese, and the like.
- references samples used in methods of the disclosure should capture the complexity of the chemical constituents present in the sample set.
- any suitable reference sample that sufficiently captures such complexity may be used.
- the refence sample may a sample designed to represent “normal” human plasma.
- the reference sample may be produced by combining aliquots from a plurality of individual samples in the sample set. Any number of aliquots may be combined although it will be understood that the larger the number of aliquots combined, the more accurate the result.
- the refence sample comprises aliquots from at least about 10%, at least about 20%, at least about 30% ⁇ at least about 40% ⁇ at least about 50% ⁇ at least about 60% ⁇ at least about 70% ⁇ at least about 80% ⁇ at least about 90% ⁇ or from 100% of the individual samples in the sample set.
- the refence sample comprises aliquots from at least 2, at least about 5, at least about 10, at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 500, at least about 750, at least about 1000, at least about 500, or at least about 2000 individual samples in the sample set.
- reference features are obtained from the refence sample.
- the reference features may have been obtained from the sample at a point in time significantly prior to (e.g., days or weeks) than the time at which analysis of the refence features is conducted.
- the entities obtaining the reference features from the reference sample and conducting the analysis may be, but need not be, the same entity.
- reference features may be stored, such as in a feature matrix, and retrieved at a later time for analysis.
- obtaining the reference features from the refence sample may be part of the disclosed method so that obtainment and analysis of the features is conducted relatively simultaneously.
- the refence features are obtained by subjecting the reference sample to a separation technique and/or mass spectrometry.
- the refence features are obtained by subjecting the reference sample to a separation technique and mass spectrometry to produce a reference feature comprising a rt and a m/z value.
- Filtering of the reference features may comprise identifying and removing non-relevant reference features to produce a set of relevant reference features (e.g., a feature matrix, feature table, feature list, etc.) that characterize unique biological metabolites.
- the identified non-relevant reference features and/or the relevant reference features may then be used to analyze data obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample.
- Such a method is advantageous as it restricts the computation necessary to identify both the relevant and non-relevant features to the reference sample, thereby reducing the computational burden on the entire sample set of individual samples.
- using relevant reference features to analyze data from an individual sample may comprise applying the relevant reference features to sample features.
- applying the relevant reference features to sample features may comprise comparing the relevant reference features to the sample features, identifying those sample features corresponding to relevant refence features, and using those samples features that correspond to relevant reference features to determine the composition of unique biological metabolites in the individual sample.
- corresponding features are features from two different samples that comprise the same data points (e.g., the same rt, the same m/z value, or the same rt and the same m/z value).
- applying the relevant reference features to sample features may comprise limiting the analysis of data obtained from an individual sample to that data having retention times and/or m/z values of the relevant reference features.
- applying the non-relevant reference features to sample features may comprise ignoring (i.e., excluding from further analysis) sample features corresponding to non-relevant refence features.
- the remaining set of relevant sample features characterize the unique biological metabolites in the sample from which the sample features were obtained and may be used to identify the composition of unique biological metabolites in the individual sample.
- the instant disclosure has discussed applying the non- relevant, and relevant, reference features to sample features from one individual sample. However, it will be understood by those of skill in the art that because the reference sample reflects the complexity of the entire sample set, the aforementioned processes of applying the refence features to sample features may be iterated with each individual sample in the sample set.
- the result of such iteration is the determination of the composition of unique metabolites in every individual sample in the sample set.
- measurements may drift due to environmental factors (e.g., temperature, humidity, etc.) affecting the measurement instruments.
- environmental factors e.g., temperature, humidity, etc.
- Such shift may cause a change in the features obtained for a particular metabolite, making it difficult to compare and align features across samples analyzed at disparate times.
- iRT indexed retention time
- the reference features may be updated by periodically subjecting the reference sample to the separation technique and mass spectrometry, thereby obtaining updated reference features from the reference sample.
- the reference sample is re- analyzed to update the refence features after a plurality of individual samples is analyzed.
- the plurality of individual samples comprises at least 2, at least 3, at least 5, at least 10, at least 20, at least 50, at least 100 or at least 150 samples.
- Samples [0056] Blood samples were collected at participants’ homes in dipotassium ethylenediaminetetraacetic acid (K2- EDTA) collection tubes, which were immediately placed on a frozen gel pack. After shipment to the laboratory via courier or postal service, the samples were then centrifuged to isolate plasma, which was subsequently stored at -80 °C. Pooled samples were prepared from a subset of plasma samples to serve as quality control (QC) samples, and for use in peak list formation and metabolite identification (see Supporting Information). The QC sample was prepared from 58 samples of the first analysis batch and thus subjecting samples to an additional freeze-thaw cycle was avoided.
- K2- EDTA dipotassium ethylenediaminetetraacetic acid
- Polar metabolite extracts were directly analyzed (without any drying step) via hydrophilic interaction liquid chromatography (HILIC) coupled to HRMS in negative mode. For both LC/MS analyses, samples were randomized. Additionally, LC/MS/MS data were acquired to aid metabolite identification. [0059] Even though positive-mode and negative-mode data for both lipid and polar metabolite extracts was not collected, doing so would be beneficial when resources permit. Blank samples were injected at the beginning and end of each worklist and used for background peak detection/removal. [0060] Representative total ion chromatograms for blank, study, and QC samples for both lipid and polar metabolite extracts are shown in FIGS.2a & 2b.
- the R and Python scripts used to perform the peak detection analysis are available on GitHub (https://github.com/e-stan/metabolomics_workflow) and include the values of all parameters utilized.
- a peak list for the lipid metabolites was directly generated based on identifications from Lipid Annotator. Any workflow or software can be used to generate peaks lists.
- Metabolite identification [0064] Identification of polar metabolites was supported by matching the accurate mass and MS/MS fragmentation data to an in-house MS/MS library created from authentic reference standards and online MS/MS libraries with DecoID software. For online database searching, the top hit for each feature with a dot-product similarity of greater than 80 was considered as the putative identification.
- MSI identification levels are given in Table S1 and Table S2 for polar and lipid metabolites, respectively. Code and scripts used to perform the automated portion of the metabolite identification workflow are available on GitHub. Lipid iterative MS/MS data were annotated with the Lipid Annotator software (Agilent Technologies), and lipid identifications were provided as sum compositions because insufficient information was available to deduce specific fatty acid compositions. Lipid identifications were subject to the same manual curation as applied to the polar metabolite data. Any workflow or software can be used for compound identification.
- Extracting peak areas [0066] Following the generation of a peak list and metabolite identification, all data files were analyzed in Skyline (version 20.1.0.155) batch per batch to obtain peak areas. The m/z values of the metabolite target lists were used to extract peak areas under consideration of retention times or indexed retention times (iRT) (see Supporting Information). Due to the data being acquired over several months, 14 different batch correction approaches were tested for peak area normalization (see Supporting Information). Additionally, a report containing the acquisition times of all samples was exported to be used for batch correction.
- iRT indexed retention times
- Polar metabolites were separated on a SeQuant® ZIC®-pHILIC column (100 x 2.1 mm, 5 ⁇ m, polymer, Merck-Millipore) including a ZIC®-pHILIC guard column (2.1 mm x 20 mm, 5 ⁇ m). The use of an inline filter prior to the guard column is recommended.
- the column compartment temperature was maintained at 40 ⁇ C and the flow rate was set to 250 ⁇ L ⁇ min-1.
- the mobile phases consisted of A: 95% water, 5% acetonitrile, 20 mM ammonium bicarbonate, 0.1% ammonium hydroxide solution (25% ammonia in water), 2.5 ⁇ M medronic acid, and B: 95% acetonitrile, 5% water, 2.5 ⁇ M medronic acid.
- Medronic acid was used in mobile phase B (mainly acetonitrile) in this and previous studies.
- the method may also be practiced by adding 5 ⁇ M medronic acid to the aqueous mobile phase A, but no medronic acid in B.
- the m/z range was 50-1700. Data were acquired under continuous reference mass correction m/z 119.0363 and 966.0007. Samples were randomized prior to analysis. In addition, a quality control (QC) sample was injected after every 12th sample to monitor signal stability of the instrument. The HILIC column was equilibrated with three blank and four QC injections prior to starting the actual samples. Prior to every batch, the mass spectrometer was calibrated and before starting to inject the actual samples, the TIC, internal standard intensities, and mass accuracy of the QC samples was compared to previous batches. More details on best practices for system suitability testing can be found elsewhere.
- QC quality control
- LC/MS analysis of lipid metabolites [0065] An aliquot of 4 ⁇ L of lipid extract was subjected to LC/MS analysis by using an Agilent 1290 Infinity II LC-system coupled to an Agilent 6545 Q-TOF mass spectrometer with a dual Agilent Jet Stream electrospray ionization source. Lipids were separated on an Acquity UPLC® HSS T3 column (2.1 x 150 mm, 1.8 ⁇ m) including an Acquity UPLC® HSS T3 VanGuard Pre-Column (2.1 x 5mm, 1.8 ⁇ m) at a temperature of 60 ⁇ C and a flow rate of 250 ⁇ L ⁇ min-1.
- the mobile phases consisted of A: 60% acetonitrile, 40% water, 0.1% formic acid, 10 mM ammonium formate, 2.5 ⁇ M medronic acid, and B: 90% 2-propanol, 10% acetonitrile, 0.1% formic acid, 10 mM ammonium formate (dissolved in 1 mL water).
- the following linear gradient was used: 0-2 min, 30% B; 17 min, 75% B; 20 min, 85%; 23-26 min, 100% B; 26, 30% B followed by a re- equilibration phase of 5 min.
- Lipids were detected in positive ion mode at a scan rate of 2 spectra per second with the following source parameters: gas temperature 250 ⁇ C, drying gas flow 11 L ⁇ min-1, nebulizer pressure 35 psi, sheath gas temperature 300 ⁇ C, sheath gas flow 12 L ⁇ min-1, VCap 3000 V, nozzle voltage 500 V, Fragmentor 160 V, Skimmer 65 V, Oct 1 RF Vpp 750 V.
- the m/z range was 50-1700. Data were acquired under continuous reference mass correction at m/z 121.0509 and 922.0890 in positive ion mode. Samples were randomized before analysis. In addition, a QC sample was injected after every 12th sample to monitor signal stability of the instrument.
- the RP column was equilibrated with three blank and three QC injections prior to starting the actual samples. The same system suitability testing described for the polar metabolite analysis was used for the lipid metabolite analysis.
- Preparation of pooled samples [0067] A pool of samples from the first batch was prepared by mixing 330 ⁇ L of all samples that contained at least 500 ⁇ L (58 samples). Next, 160 aliquots of 110 ⁇ L each were frozen at -80 °C, so that two 50 ⁇ L pooled QC samples could be included in every batch. This pooled QC not only served as a reference sample across batches, but it was also used for feature detection and metabolite identification.
- Metabolite identifications were supported by the pooled QC samples and 8 additional pooled samples that were prepared by pooling 10 ⁇ L aliquots of 10 samples of 8 random batches.
- the eight additional pooled samples cover the four different locations and a wide age range.
- Three pooled samples contained only samples from Denmark, whereas the other five were a mix of the three different US locations.
- the average age ranged from 69.1 to 87.9 years, with an average standard deviation of 13.3 years. Across those pooled samples, 52% were from male and 48% from female participants. MS/MS data were acquired in negative mode (see details below) on 9 total pooled samples.
- MS/MS spectra for polar and lipids metabolites were acquired by using an iterative data dependent acquisition (iDDA) approach in the MassHunter Acquisition Software (Version 10.1.48, Agilent Technologies) on an Agilent 6545 QTOF. The same source settings as for MS1 data acquisition were used. MS/MS spectra were acquired at a scan rate of 3 spectra/s with different intensity thresholds and collision energies of 10, 20, and 40 V to increase identification rates.
- iDDA can be readily implemented by creating consecutive acquisition methods for the blank and study samples in the iterative workflow. Thus, exclusion lists no longer need to be manually generated after each run. Details on how to achieve iDDA on the Agilent system can be found in an application note.
- MS/MS data for polar metabolites were acquired on an Orbitrap ID-X Tribrid mass spectrometer (Thermo Scientific).
- a Vanquish Horizon UHPLC system was interfaced with the mass spectrometer via electrospray ionization in negative mode with a spray voltage of 2.8 kV.
- RF lens value was 60%.
- Data were acquired in data dependent acquisition (DDA) mode by using the built-in deep scan option (AcquireX) with a mass range of 67-900 m/z and 120K resolution for MS1 scans. MS/MS scans were acquired at 15K resolution from three distinct pooled samples.
- DDA data dependent acquisition
- AcquireX built-in deep scan option
- MS/MS scans were acquired at 15K resolution from three distinct pooled samples.
- iRT uses the retention time of selected “indexing” compounds to adjust the retention times of other compounds.
- iRT were implemented for each lipid class, selecting 2-3 lipid metabolites per class as indexing compounds (see Table 3). In selecting indexing compounds, lipids were chosen that were chromatographically well separated from other species and that had high intensity.
- the first step of the process is to compute the iRT of each lipid in a class. The calculation relies on the retention times of each lipid as measured from a single sample.
- iRT uses the retention time of the first ( ⁇ ⁇ ) and last ( ⁇ ⁇ ) eluting lipid in a class to convert a measured retention time for a compound ( ⁇ ⁇ ) to its corresponding iRT ( ⁇ ⁇ ).
- ⁇ ⁇ (1)
- the iRT of the indexing compounds ( ⁇ ) and the observed retention times for the indexing compounds in the sample are used to calculate a linear regression of iRT to retention time by minimizing the error between observed and fit retention times, according to Equation 2.
- each metabolite was normalized by computing the mean of the metabolite intensity in the flanking QC samples for a particular research sample and multiplying the metabolite’s intensity in this research sample by the mean of the metabolite’s intensity across all QC samples and dividing by the mean of the flanking QC samples.
- each batch of samples was QC normalized individually followed by an inter-batch correction with ComBat.
- SVR linear regression, and random forest, the respective model was fit for each metabolite by using the batch and run-order position of the QC samples against the deviation in the metabolite’s intensity from the mean QC intensity.
- the process reduces the data burden of untargeted metabolomics such that informatics tools typically applied to targeted studies can be leveraged to profile research samples efficiently and rapidly, without the need to subject each sample to computationally intensive analyses (e.g., peak detection, correspondence determination, peak grouping, metabolite identification, etc.).
- a schematic of the workflow is shown in FIG.3.
- the disclosed workflow was used to analyze a subset of ⁇ 2,000 human plasma samples from the Long Life Family Study (LLFS), conducted by the National Institute on Aging at the National Institutes of Health (NIH), Bethesda, Maryland, U.S.
- LLFS Long Life Family Study
- peak areas were extracted from the research and QC samples in a batch-by-batch fashion by using Skyline .
- the Skyline command-line interface can be used to automatically generate a document for each batch.
- retention times were stable for all samples within a batch (FIGS.4a & 4b).
- retention-time bounds were set by inspection of the QC samples within each batch, and these bounds were applied to all samples by importing peak boundaries for each sample.
- a Python script was used to generate the peak boundary import file in the current work, this is no longer necessary when using the “Synchronize Integration” function in Skyline 21.2, which was released after the instant data processing was performed.
- the list which was comprised of 5,894 total features, included features that were only detected from a single research sample.
- a total of 3,241 features (both identified and unknowns) from the list were detected in at least one replicate of the pooled sample (FIG. 6a).
- the m/z values of all features were searched against endogenous metabolites in the Human Metabolome Database and the Kyoto Encyclopedia of Genes and Genomes. In total, 40.4% of the features detected in the pooled sample had at least one hit in the databases.
- missing values must be removed from the data to facilitate downstream processing.
- missing values were infrequent ( ⁇ 0.03% of all measurements) and most likely arose from metabolites at concentrations below the limit of detection of the instrument rather than random metabolite dropout during peak detection.
- missing values were imputed with the half minimum approach.
- the disclosed method focuses on processing data from a small number of pooled samples that are created by mixing aliquots of individual samples from the study. By initially limiting data processing to only the pooled samples, the disclosed method leverages standard informatics tools in untargeted metabolomics that have been optimized for small sample sizes.
- the pooled samples are intended to capture the totality of unique compounds across the entire study cohort but, in practice, unique compounds from individual samples are missed.
- An analysis of the data shows that unique compounds are most often missed when they are only present in a few participant samples at low concentrations, which causes them to be diluted below the limit of detection during pooling.
- the results indicate that signals missing from the analysis of pooled samples are more likely to originate from rare exogenous compounds (e.g., chemicals from unique hygiene products or specific environments). While rare exogenous compounds may certainly be biologically interesting, their analysis will require the development of new peak detection, alignment, and annotation algorithms that can be scaled to thousands of samples without loss of functionality or accuracy.
- the present disclosure enables high-throughput applications of untargeted metabolomics at the population scale, such as observed in the LLFS.
- the disclosed workflow collects separation and mass spectrometry features for all individual samples in the sample set.
Landscapes
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
The present disclosure relates to an improved workflow method for metabolic analyses of a large number of samples in a sample set. The improved method utilizes a reference sample that represents the chemical complexity of the sample set. Analysis of the reference sample identifies the relevant and non-relevant features, which may then be used to focus analysis of individual samples on specific portions of the data obtained from each individual sample, thereby reducing the computational burden of analyzing all the samples in the sample set. Also disclosed are systems for practicing the disclosed methods.
Description
IMPROVED METHOD FOR SCALABLE UNTARGETED METABOLOMIC WORKFLOW CROSS-REFERENCE TO RELATED APPLICATIONS [0001] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial Number 63/386,087, entitled “IMPROVED METHOD FOR SCALABLE UNTARGETED METABOLOMIC WORKFLOW,” filed December 5, 2022, and the contents of which are incorporated herein in their entirety. BACKGROUND OF THE INVENTION [0002] The goal of precision medicine is to identify subgroups within the population for whom strategies to prevent, diagnose, and treat disease states can be uniquely tailored. It is expected that studies of large, diverse, and longitudinal cohorts will be critical to accelerate progress toward this end over the next decade. Indeed, research projects having hundreds of thousands of participants are already emerging with the FinnGen, UK Biobank, and All of Us cohorts. Given the key role of metabolism in health and diagnostics, the application of untargeted metabolomics to big cohorts promises to inform the practice of precision medicine in ways that are highly complementary to other technologies such as genomics. Accordingly, many metabolomic methods have been developed (see, for example, U.S. Patent No. 7884318, U.S. Patent No.8428881, and U.S. Patent No.10267777, all of which are incorporated herein by reference in their entireties). At this time, however, standard data-processing workflows used in untargeted metabolomics are not amenable to such large sample sizes. [0003] When performing untargeted metabolomics with liquid chromatography/mass spectrometry (LC/MS), a typical biological specimen yields thousands of signals having unique m/z values and retention times, often referred to as “features”. After features are detected from each individual sample, it must be determined which represent the same analyte across all of the LC/MS runs, a process known as alignment or correspondence determination. Finally, the intensity of every feature from each aligned sample must be assessed. A complication is that experimental drift is compound specific and introduces non-linear measurement errors of variable magnitude throughout an experiment. Considering the number of features in a conventional untargeted metabolomics experiment,
it is therefore impractical to process and interpret the data by manually inspecting all of the features at the global scale. [0004] Over the last two decades, several software platforms such as mzMine, XCMS, and MS-DIAL have emerged for automated processing of untargeted metabolomics data. These informatics tools are highly effective and widely used for the analysis of small cohorts, but they were not designed to support projects with thousands of samples. While methods for feature detection can generally be applied to each data file separately, making analysis of large sample numbers cumbersome but feasible, algorithms for correspondence determination are not as readily scaled. Methods such as Obiwarp within XCMS were designed to process all of the data files on one computing workstation at the same time. As the number of samples being evaluated grows, the amount of memory required for data processing increases and can eventually prevent the analysis from being completed. The specific number of samples that can be supported depends upon the user’s available computing power and software, but a recent report indicated that programs for untargeted metabolomics reproducibly crash when processing 250 data files or more. Although approaches, such as SLAW, have developed modified algorithms with lower computational overhead, user-friendly tools for processing large sample-sets remain limited. [0005] Thus, there is a need for improved metabolomic workflows that are scalable for analysis of a large number (e.g., >1,000) of samples and the correspondingly large datasets. The present disclosure addresses this need and provides other benefits as well. BRIEF DESCRIPTION OF THE INVENTION [0006] One aspect is a method of non-targeted determination of unique biological metabolites in individual samples of a sample set, each individual sample comprising a composition of chemical constituents, the method comprising: a. analyzing reference features obtained from a reference sample to identify non-relevant reference features; b. filtering the reference features by removing the non-relevant reference features, thereby producing a set of relevant reference features that characterize the unique biological metabolites; and,
c. applying the non-relevant and/or relevant reference features to sample features obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample. [0007] In certain aspects, the method may comprise conducting the steps of analyzing, filtering and applying on each individual sample in the sample set, thereby determining the unique biological metabolites in all individuals of the sample set. In certain aspects, after the steps of analyzing, filtering, and applying have been conducted on a plurality of individual samples in the sample set, the reference features are updated by re- analyzing the reference sample to obtain updated reference features the reference features are replaced with the updated reference features. The number of individual samples in the plurality of samples may be less than the entire number of individual samples in the sample set. The plurality of samples may comprise at least 2, samples, optionally at least 3 samples, optionally at least 5 samples, optionally at least 10 samples, optionally at least 20 samples, optionally at least 50 samples, optionally at least 100 samples, or optionally at least 150 samples. The plurality of samples my comprise at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 90%, or at least 95% of the samples in the sample set. In certain aspects, the reference sample may comprise an aliquot from each of a plurality of the individual samples. The reference features may be obtained by a method comprising subjecting the reference sample to a separation technique and mass spectrometry. The separation technique may comprise chromatography, which may comprise gas chromatography or liquid chromatography. In certain aspects, the non-relevant reference features may characterize artifacts, background chemical constituents and/or non-unique biological metabolites. The background chemical constituents may comprise contaminants. The non-unique biological metabolites may comprise adducts and/or fragment ions. The reference features may comprise reference separation data and reference mass data generated from the reference sample. [0008] In certain aspects, the step of applying may comprise: identifying in the sample features those sample features corresponding to non-relevant reference features and removing those sample features that correspond to the non-relevant reference features to produce a set of relevant sample features; and, using the relevant sample features to determine the unique biological metabolites in the sample. In certain aspects, the
step of applying may comprise: identifying in the sample features those sample features corresponding to relevant reference features to produce a set of relevant sample features; and, using the relevant sample features to determine the unique biological metabolites in the sample. [0009] The method may comprise conducting the steps of obtaining and applying to each sample in the sample set, thereby determining the composition of the chemical constituents in all individual samples in the sample set. The method may comprise identifying one or more of the chemical constituents by comparing the relevant sample features to a library of information comprising features characterizing chemical entities. [0010] One aspect comprises a method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a. obtaining reference features comprising separation data and mass data generated from a reference sample; b. filtering the reference features by eliminating separation data and mass data relating to non-unique chemical constituents to produce filtered reference features comprising filtered separation data and filtered mass data; c. using the filtered reference features to characterize the presence and/or quantity of unique chemical constituents in each individual sample of the sample set. [0011] One aspect of the disclosure comprises a method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a. using a separation technique and mass spectrometry to produce reference features comprising separation data and mass data from a reference sample; b. collecting and storing the reference features; c. analyzing the stored reference features to identify non-relevant reference features comprising separation data and mass data characterizing non- unique chemical constituents; d. removing the non-relevant reference features comprising separation data and mass data characterizing non-unique chemical constituents to produce a
filtered data set comprising filtered reference features comprising separation data and mass data characterizing unique chemical constituents; e. using the filtered reference features to characterize the presence and/or quantity of unique chemical constituents in each sample of the sample set. [0012] The step of using the filtered reference features may comprise: obtaining sample features from an individual sample; identifying sample features that correspond to the filtered reference features to produce a list of relevant sample features that characterize unique chemical constituents, thereby determining the composition of chemical constituents in the sample. [0013] One aspect of the disclosure if a system for conducting the method of any one of claims 1-21, the system comprising: a. a separation apparatus for separating the chemical constituents and producing separation-related features characterizing the chemical constituents; b. a mass spectrometer for performing mass spectrometry on portions of the separated chemical constituents and producing mass spectrometry-related features characterizing the chemical constituents; c. a first module for receiving, collecting and/or storing the separation-related features and/or the mass spectrometry -related features; and, d. a user interface coupled to the module for making the separation-related features and/or the mass spectrometry -related features available in human- accessible form. [0014] In certain aspects, the system may comprise a library of features characterizing chemical entities, the features produced using the separation apparatus and the mass spectrometer, wherein the library of features comprises separation features and mass spectrometry features characterizing identified chemical entities. The separation apparatus may comprise an apparatus for performing chromatography, which may comprise liquid chromatography or gas chromatography. Liquid chromatography may comprise HPLC or UPLC. The separation apparatus may comprise an electrophoresis apparatus, which may be coupled to the mass spectrometer. The separation-related features may comprise peak
retention time, peak intensity, and/or peak width. The mass-spectrometry-related features may comprise mass or an m/z value. BRIEF DESCRIPTION OF THE DRAWINGS [0015] FIG. 1 illustrates solid phase extraction of plasma by using CAPTIVA 96 well plates (Agilent Technologies). One aliquot of plasma is subjected to solid-phase extraction with a CAPTIVA EMR 96-well plate. After elution of the polar fraction (steps 1-8), the lipids still bound to the solid-phase extraction material are eluted into a second collection plate, dried to completeness, and reconstituted for RP separation followed by HRMS analysis. The polar metabolite extract is subjected to HILIC separation followed by HRMS analysis without further manipulation. Created with BioRender.com [0016] FIGS. 2 a & b show total ion chromatograms. Shown are representative total ion chromatograms for a blank (black), a study sample (purple), and a QC sample (green) for lipid (a) and polar (b) metabolites. [0017] FIG. 3 illustrates a pipeline for handling metabolomics data. Polar and lipid metabolites are extracted from plasma samples into 96-well plates for LC/MS analysis. A pooled sample is prepared for feature detection, MS/MS acquisition, and use as a QC sample. Untargeted metabolomics analysis is performed on all samples. After detecting features from the pooled sample, background features and degeneracies are filtered. The remaining features are subjected to metabolite identification with DecoID and Lipid Annotator, and the returned putative identifications are manually curated. The peak areas for these metabolites are extracted from the research samples by using Skyline. Retention-time shifts are manually corrected per batch for polar metabolites and automatically adjusted by using indexed retention times for lipid compounds. The peak areas are imputed and normalized to remove missing values and batch effects from the data. The final output contains the metabolite information (name, m/z, retention time) and normalized metabolite intensities for each research and QC sample. [0018] FIGS. 4a & 4b show data pre-processing of a single batch. a) Overlay of the total ion chromatograms (TICs) of the QC samples injected after every 12th sample. b) Visualization of the peak boundaries and apex (shown as a black line) enable confirmation of stable retention times after importing data files into Skyline. If two compounds, such as leucine and isoleucine, are not fully baseline separated, peaks might not
be integrated consistently. Peak boundaries can be adjusted and imported for all data files. Baseline signal from QTOF instruments will result in a reported retention time from Skyline, even if no peak is present (e.g., blank samples). [0019] FIGS.5a-5c show retention time deviations in lipid metabolites. (a- b) retention time (RT) deviations (in minutes) versus run order for LPC 0:0/18:2 (a), and TG 48:1 (b). Each dot represents a sample. Dots are colored by analysis batch. c) Histogram of all retention time deviations across all samples. [0020] FIGS. 6a-6e shows a comparison between individual and pooled samples. a) Venn diagram showing the breakdown of features detected in at least one individual sample (red) versus in at least one pooled sample (green). A higher fraction of features missed in the pooled sample had no hits returned when searching the Human Metabolome Database and KEGG. b-c) Histograms showing the fraction of samples where a feature was detected. In b, the distribution for features detected only in individual samples is shown. In c, the distribution for features detected in both individual and pooled samples is shown. d-e) Histograms showing the log10 transformed mean intensities of features detected in only individual samples (d) or detected in both individual and pooled samples (e). [0021] FIGS. 7a-7c show targeted extraction of peak areas versus traditional global peak integration. a-b) The pie charts represent the percentage of all measurements that are missing values (grey) when targeted extraction of peak areas with Skyline is performed (a) or when traditional peak integration within XCMS is performed (b). In (a) the missing value percentage is too small to be visible on the pie chart. c) Boxplots showing the coefficients of variation (CV) when peaks are integrated with targeted extraction in Skyline (blue) or with traditional processing in XCMS (red). [0022] FIGS.8a-8e illustrate correcting for batch effects in metabolomics data. a) Principal components analysis (PCA) of unnormalized lipid metabolic profiles shows strong batch effects. Each dot represents a unique sample. Dots are colored according to their corresponding batch number. b) Comparison of 14 different batch-correction algorithms on the lipid metabolic profiles. The normalization score is the change in coefficient of variation (CV) in the research samples (relative to the unnormalized data) divided by the change in CV for the QC samples. A higher score indicates a reduction of technical variation. c) PCA plot of random-forest normalized lipid metabolic profiles shows reduced clustering by batch. d) Intensity of CE 16:0 as a function of run order for both unnormalized (top) and random forest corrected data (bottom). e) Violin plots showing the CV distribution of all compounds
in the QC samples for each evaluated batch-correction algorithm. The polar metabolite counterpart to these data is shown in FIG.9. [0023] FIGS.9a-9e show correcting for batch effects in metabolomics data. a) Principal components analysis (PCA) of unnormalized polar metabolic profiles shows clustering by batch. Each dot represents a unique sample. Dots are colored according to batch number. b) Comparison of 14 different batch-correction algorithms on the polar metabolic profiles. Normalization score is the change in coefficient of variation (CV) in the research samples (relative to the unnormalized data) divided by the change in CV for the QC samples. A higher score indicates a reduction of technical variation. c) PCA plot of random forest normalized polar metabolic profiles. While samples are still clustered by batch, this is due to biological differences between batches in the research samples. QC samples are not clustered by batch after correction (see FIG.11). d) Intensity of gingerol as a function of run order for both unnormalized (top) and random forest corrected data (bottom). e) Violin plots showing the CV distribution of all compounds in the QC samples for each evaluated batch-correction algorithm. The lipid metabolite counterpart to these data is shown in FIG.8. [0024] FIGS. 10a & 10b show that QC samples enable discrimination of batch and biological effects. Metabolite intensities can change as a function of analysis batch. If batches are biased by biologically different sample groups, then the application of non-QC based batch normalization can remove biological variation in addition to technical variation because the different types of variation cannot be easily distinguished. However, QC samples enable technical shifts caused by batch number (such as DGTS 17:0) to be differentiated from biological differences between batches (such as inosine). The plots in (a) show the unnormalized metabolite intensities as a function of run order for QC samples (left) and research samples (right). In (b) the metabolite intensities after random forest correction are shown for QC and research samples for the same compounds. Dots are colored by analysis batch. [0025] FIGS.11a-11d show internal standard variability. a-b) Violin plots showing the distribution of coefficients of variation (CV) for the internal standards across all samples within a batch or across all batches (1-22) for both unnormalized (a) and random forest corrected data (b). c-d) Principal components analysis (PCA) of internal standard intensities across all samples in the unnormalized (c) and random forest corrected data (d). Each dot is a sample. Samples are colored by batch. e) Violin plots showing the distribution
of CVs for the internal standards across all samples for the 14 batch correction methods evaluated in this study. [0026] FIGS.12a-12d show that random forest normalization reduces batch effects in QC samples. Principal components analysis (PCA) of unnormalized (a-b) metabolic profiles of QC samples shows strong batch effects in polar (a) and lipid metabolites (b). After random forest correction (c-d), QC samples are not clustered by batch for polar (c) and lipid (d) metabolites. Each dot represents a QC sample. Dots are colored according to batch number. [0027] FIGS. 13a-13d show that metabolic profiles are reflective of geographic location. a) Principal components analysis (PCA) of normalized metabolic profiles (polar and lipid metabolites) shows clustering based on United States (BU = Boston, NY = New York City, PT = Pittsburg) and Denmark (DK, Odense) field sites. Each dot represents a unique sample. Dots are colored according to geographic location. b-c) Lipid (b) and polar (c) metabolites associated with geographic location (|FC| > 2, p < 0.05, One- way ANOVA). d) Age distribution for samples from the different field sites. Data shown are median ± interquartile range, NDHB, N,N-diethyl-4-hydroxybenzamide; CMPF, 3-carboxy- 4-methyl-5-propyl-2-furanpropanoic acid. [0028] FIGS.14a-14d show analysis of unknown metabolites. a) Heatmap showing the relative metabolite intensity across all samples for the 3421 unknown metabolites profiled in this study (rows = metabolites, columns = samples). b) Histogram showing the coefficient of variation (CV) across the QC samples for all unknown metabolites after random forest batch correction. c) Heatmap of the 29 unknown metabolites significantly associated with field site. d) Boxplot showing one exemplary unknown metabolite that shows significantly lower intensity in Denmark (DK) than the US sites. The column colors in (a) and (c) show the field site for each sample (BU = Boston, NY = New York City, PT = Pittsburg, DK = Odense). [0029] FIG.15 shows that metabolites associated with geographic location are not reflective of subject age. Principal components analysis (PCA) of normalized metabolic profiles (polar and lipid metabolites) shows no age-dependent pattern in geographically associated metabolites. Each dot represents a unique sample. Dots are colored according to subject age.
DETAILED DESCRIPTION OF THE INVENTION [0030] The present disclosure relates to an improved method for conducting a metabolomics workflow. More specifically, the present disclosure relates to a metabolomics workflow that reduces the computational burden required to analyze the abundance of compounds present in a sample, thereby allowing the process to be scaled for analysis of a larger number of samples than is possible with currently available techniques. The disclosed process achieves this scalability using a reference sample that represents the chemical complexity of an entire sample set. Analysis of the refence sample allows detection and identification of thousands of features, most of which do not characterize unique biological metabolites, within a single sample. The features may then be filtered to remove non-relevant features, which characterize non-relevant chemical constituents, yielding a list of relevant features. This list of relevant features may then be used to identify relevant features characterizing unique-biological metabolites within each individual sample, thereby reducing the computational burden necessary to analyze the sample set. Thus, a method of the disclosure may generally be practiced by analyzing a reference sample to obtain biologically relevant reference features and using these relevant reference features to identify relevant features in one or more individual samples in the sample set. The relevant features may also be quantitated, thereby producing a metabolomic fingerprint of the biological metabolites in the individual sample. In some relevant features and/or the metabolic fingerprint may be used to identify the individual biological metabolites present in the individual sample. In some aspects, the relevant features may be assessed in each individual sample in the sample set, serially or sequentially, to determine the metabolomic fingerprint for many or all the individual samples in the sample set. [0031] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the claims. [0032] It must be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, a metabolite refers to one or more metabolites. As such, the terms "a", "an", "one or more" and "at least one" can be used interchangeably.
[0033] Similarly, the terms "comprising", "including" and "having" can be used interchangeably. As used herein, the term “comprising” may be replaced with “consisting” or with “consisting essentially of” in particular aspect, as desired. [0034] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements or use of a "negative" limitation. [0035] Various terms relating to aspects of the present disclosure are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art, unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein. [0036] Unless otherwise expressly stated, it is in no way intended that any method or aspect set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not specifically state in the claims or descriptions that the steps are to be limited to a specific order, it is in no way intended that an order be inferred, in any respect. This holds for any possible non-expressed basis for interpretation, including matters of logic with respect to arrangement of steps or operational flow, plain meaning derived from grammatical organization or punctuation, or the number or type of aspects described in the specification. [0037] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub- combinations are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein. [0038] As used herein, the term “about” means that the recited numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical value is used, unless indicated otherwise by the
context, the term “about” means the numerical value can vary by ±10% and remain within the scope of the disclosed embodiments. [0039] Publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. [0040] One aspect of the disclosure is a method of non-targeted determination of unique biological metabolites in individual samples of a sample set, each individual sample comprising a composition of chemical constituents, the method comprising analyzing reference features obtained from a reference sample to identify non- relevant reference features; filtering the reference features by removing the non-relevant reference features, thereby producing a set of relevant reference features that characterize the unique biological metabolites; and, applying the non-relevant and/or relevant reference features to sample features obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample. [0041] As used herein, non-targeted determination means characterizing the chemical composition of a complex sample containing chemical constituents without prior knowledge of one or more chemical constituents in the sample. The phrase “chemical constituents” refers to the molecules present within a sample. “Chemical constituents” encompasses any molecule present within a sample including but not limited to, metabolites, including endogenous metabolites, exogenous metabolites, unique metabolites, non-unique metabolites and contaminants. [0042] As used herein, “metabolite” and “biological metabolite”, may be used interchangeably and refer to the set of small molecules comprising substrates, intermediates, and products of metabolism. Metabolites of the disclosure are generally less than about 2000 kDa in size, optionally less than about 1500 kDa in size. Metabolites include both endogenous metabolites (i.e., produced by an individual) and exogenous metabolites (i.e., drugs, environmental toxins, etc.). Examples of endogenous metabolites include, but are not limited to, organic acids, fatty acids, triglycerides, cholesterol, phospholipids, sugars, vitamins, and co-factors. It is understood by those of skill in the art that analysis of a metabolite using a technique such as mass spectrometry results in the production of numerous species of metabolite ions. The different species result from such things as
fragmentation of the metabolite, incorporation into the metabolite of isotopes such as carbon- 13 or nitrogen-15, and binding of an ion species to such things salts or solvent, to form adducts. This may result in the production of several signals, each signal coming from a different species of the metabolite. In order to perform methods of the disclosure, preferably, one single ion species is chosen to represent each metabolite. As used herein, a unique metabolite refers to a select, monoisotopic, single ion species of an intact metabolite. As such, the term “unique metabolite” excludes fragmented species of metabolites, metabolites comprising naturally abundant stable isotopes, such as carbon-13 or nitrogen-15, and ion species of the metabolite other than the selected ion species, such as other adducts of the metabolite. A non-unique metabolite refers to a fragment of a metabolite, a non- monoisotopic form of the selected metabolite and additional adducts of the selected ion species of the metabolite. [0043] The term back “background” refers to features characterizing contaminants and artifacts. The term “contaminant” refers to chemicals present in the sample but that originate from outside of the sample (e.g., from the equipment used to analyze the sample). For example, a tube holding a sample may contain plasticizers that may leech into the sample causing the sample to contain plasticizer. Plasticizer in the sample represents a contaminate. As used herein, an “artifact” refers to features that do not characterize real chemicals, but that arise from electronic noise in the mass spectrometer or errors in downstream peak detection software. Methods of identifying contaminants and artifacts are known to those skilled in the art and are also disclosed herein. [0044] As used herein, the term “feature” refers to the one or more points of data, such as a peak retention time or mass-to-charge (m/z) value, that characterize a chemical constituent. For example, a feature obtained using a separation technique, such as chromatography (e.g., HPLC), may comprise retention time (rt). Likewise, a feature obtained using mass spectrometry may comprise a m/z value. In another example, a feature obtained using a separation technique (e.g., chromatography) and mass spectrometry may comprise both a rt and a m/z value. A “relevant feature” is a feature that characterizes a unique metabolite. A “non-relevant feature” is a feature that characterizes non-unique metabolites, contaminants, and that are artifacts. [0045] Features (e.g., reference features, sample features, etc.) of the disclosure are obtained from samples (e.g., reference samples, individual samples) of the disclosure using various methods. In certain aspects, features are obtained by subjecting a
sample to a separation technique. Any separation that suitably separates chemical constituents within the sample may be used. Examples of separation techniques include, but are not limited to, chromatography and electrophoresis, including capillary electrophoresis. Suitable methods of chromatography include, but are not limited to, gas chromatography (GC), high performance liquid chromatography (HPLC), or ultra-performance liquid chromatography (UPLC). In certain aspects, features are obtained by subjecting the sample to mass spectrometry. In certain aspects, features are obtained by subjecting the sample to chromatography, such as HPLC or UPLC, and mass spectrometry. Accordingly, in certain aspects, a feature comprises a rt and a m/z value. [0046] As used herein, the term "biological sample" refers to a sample obtained from an individual, including a sample of biological tissue or fluid origin obtained in vivo or in vitro. Such samples may include, but are not limited to, blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluids, dialysates, lavage fluids, and fluids derived from organs or tissue. An “individual sample” refers to a biological sample obtained from a single individual. The term “sample set” refers to a defined plurality of individual samples. [0047] The terms “individual”, “subject”, and “patient” are well-recognized in the art and are herein used interchangeably to refer to any human or other animal. Examples include, but are not limited to, humans and other primates, including non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cattle, sheep, pigs, seals, goats and horses; domestic mammals such as dogs and cats; laboratory animals including rodents such as mice, rats and guinea pigs; birds, including domestic, wild and game birds such as chickens, turkeys and other gallinaceous birds, ducks, geese, and the like. The terms individual, subject, and patient by themselves, do not denote a particular age, sex, race, and the like. Thus, individuals of any age, whether male or female, are intended to be covered by the present disclosure and include, but are not limited to the elderly, adults, children, babies, infants, and toddlers. Likewise, the methods of the present invention can be applied to any race, including, for example, Caucasian (white), African- American (black), Native American, Native Hawaiian, Hispanic, Latino, Asian, and European. [0048] Reference samples used in methods of the disclosure should capture the complexity of the chemical constituents present in the sample set. Thus, any suitable reference sample that sufficiently captures such complexity may be used. For example, in a
sample set comprising individual samples obtained from blood (e.g., plasma), the refence sample may a sample designed to represent “normal” human plasma. One example of such a reference sample is Standard Reference Material (SRM) 1950 (Metabolites in Human Plasma), sold by MilliporeSigma (SKU # NIST1950). Alternatively, the reference sample may be produced by combining aliquots from a plurality of individual samples in the sample set. Any number of aliquots may be combined although it will be understood that the larger the number of aliquots combined, the more accurate the result. In certain aspects, the refence sample comprises aliquots from at least about 10%, at least about 20%, at least about 30%¸at least about 40%¸at least about 50%¸at least about 60%¸at least about 70%¸at least about 80%¸at least about 90%¸ or from 100% of the individual samples in the sample set. In certain aspects, the refence sample comprises aliquots from at least 2, at least about 5, at least about 10, at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 500, at least about 750, at least about 1000, at least about 500, or at least about 2000 individual samples in the sample set. [0049] As stated above, reference features are obtained from the refence sample. In certain aspects, the reference features may have been obtained from the sample at a point in time significantly prior to (e.g., days or weeks) than the time at which analysis of the refence features is conducted. Moreover, the entities obtaining the reference features from the reference sample and conducting the analysis may be, but need not be, the same entity. For example, reference features may be stored, such as in a feature matrix, and retrieved at a later time for analysis. In certain aspects, obtaining the reference features from the refence sample may be part of the disclosed method so that obtainment and analysis of the features is conducted relatively simultaneously. In certain aspects, the refence features are obtained by subjecting the reference sample to a separation technique and/or mass spectrometry. In certain aspects, the refence features are obtained by subjecting the reference sample to a separation technique and mass spectrometry to produce a reference feature comprising a rt and a m/z value. [0050] Filtering of the reference features may comprise identifying and removing non-relevant reference features to produce a set of relevant reference features (e.g., a feature matrix, feature table, feature list, etc.) that characterize unique biological metabolites. The identified non-relevant reference features and/or the relevant reference features may then be used to analyze data obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample. Such a method
is advantageous as it restricts the computation necessary to identify both the relevant and non-relevant features to the reference sample, thereby reducing the computational burden on the entire sample set of individual samples. In certain aspects, using relevant reference features to analyze data from an individual sample may comprise applying the relevant reference features to sample features. In such aspect, applying the relevant reference features to sample features may comprise comparing the relevant reference features to the sample features, identifying those sample features corresponding to relevant refence features, and using those samples features that correspond to relevant reference features to determine the composition of unique biological metabolites in the individual sample. As used herein, with regard to comparing features, “corresponding features” are features from two different samples that comprise the same data points (e.g., the same rt, the same m/z value, or the same rt and the same m/z value). In certain aspects, applying the relevant reference features to sample features may comprise limiting the analysis of data obtained from an individual sample to that data having retention times and/or m/z values of the relevant reference features. [0051] In certain aspects, applying the non-relevant reference features to sample features may comprise ignoring (i.e., excluding from further analysis) sample features corresponding to non-relevant refence features. The remaining set of relevant sample features characterize the unique biological metabolites in the sample from which the sample features were obtained and may be used to identify the composition of unique biological metabolites in the individual sample. [0052] Heretofore, the instant disclosure has discussed applying the non- relevant, and relevant, reference features to sample features from one individual sample. However, it will be understood by those of skill in the art that because the reference sample reflects the complexity of the entire sample set, the aforementioned processes of applying the refence features to sample features may be iterated with each individual sample in the sample set. The result of such iteration is the determination of the composition of unique metabolites in every individual sample in the sample set. [0053] It will be understood by those skilled in the art that when performing an iterated analytic process, measurements may drift due to environmental factors (e.g., temperature, humidity, etc.) affecting the measurement instruments. Such shift may cause a change in the features obtained for a particular metabolite, making it difficult to compare and align features across samples analyzed at disparate times. The instant inventors have
discovered that, surprisingly, this problem may be addressed by applying the concept of an indexed retention time (iRT) in the separation step. Thus, in aspects in which the identification and/or application of features steps are repeated, the reference features may be updated by periodically subjecting the reference sample to the separation technique and mass spectrometry, thereby obtaining updated reference features from the reference sample. Preferably, several (e.g., 2, 3, 4, 5, or 6) chemical constituents having a high detection intensity and that are easily distinguishable from other chemical constituents in the separation step peaks are chosen as indexing compounds. The change in retention times of these compounds in previous “runs” and later runs can easily be determined, and this difference applied to all retention times. In certain aspects, the reference sample is re- analyzed to update the refence features after a plurality of individual samples is analyzed. In certain aspects, the plurality of individual samples comprises at least 2, at least 3, at least 5, at least 10, at least 20, at least 50, at least 100 or at least 150 samples. [0054] This written description uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal languages of the claims. EXAMPLES [0055] Samples [0056] Blood samples were collected at participants’ homes in dipotassium ethylenediaminetetraacetic acid (K2- EDTA) collection tubes, which were immediately placed on a frozen gel pack. After shipment to the laboratory via courier or postal service, the samples were then centrifuged to isolate plasma, which was subsequently stored at -80 °C. Pooled samples were prepared from a subset of plasma samples to serve as quality control (QC) samples, and for use in peak list formation and metabolite identification (see Supporting Information). The QC sample was prepared from 58 samples of the first analysis batch and thus subjecting samples to an additional freeze-thaw cycle was avoided. Generating a pooled sample from all subject specimens in the study was unfeasible because some samples were not available when the analysis started (blood draws are still ongoing).
QC aliquots were stored at -80 °C. Additionally, SPLASH Lipidomix (Avanti Polar Lipids), a deuterium labeled lipid mix designed for human plasma analysis, was used as an internal standard in QC samples for lipid metabolite analysis. The QC sample was injected after every 12th research sample. [0057] Metabolite extraction, LC/MS, and LC/MS/MS analysis [0058] To maximize coverage while minimizing sample-preparation time, a solid-phase extraction (SPE) was performed to isolate polar and lipid metabolites. One batch of plasma samples at a time was thawed on ice and vortexed. Each batch typically consisted of 92 research samples, 2 QC samples, and 2 blanks. An aliquot of each sample was then transferred into a 96-well SPE plate. A two-step extraction with either acetonitrile and methanol or methyl tert-butyl ether and methanol was used to obtain polar and lipid metabolites in separate fractions (FIG.1). Using SPE eliminates the need for a centrifugation to remove the protein fraction of the samples. The lipid extract was dried under a nitrogen flow and reconstituted prior to analysis via reversed-phase (RP) chromatography coupled to high-resolution mass spectrometry (HRMS) in positive mode. Polar metabolite extracts were directly analyzed (without any drying step) via hydrophilic interaction liquid chromatography (HILIC) coupled to HRMS in negative mode. For both LC/MS analyses, samples were randomized. Additionally, LC/MS/MS data were acquired to aid metabolite identification. [0059] Even though positive-mode and negative-mode data for both lipid and polar metabolite extracts was not collected, doing so would be beneficial when resources permit. Blank samples were injected at the beginning and end of each worklist and used for background peak detection/removal. [0060] Representative total ion chromatograms for blank, study, and QC samples for both lipid and polar metabolite extracts are shown in FIGS.2a & 2b. Although substantial signal is observed for the blank samples, features whose intensities were not at least three times higher in study samples than the blank were removed from downstream analysis. A detailed description of the metabolite extraction and LC/MS analysis can be found in the Supporting Information. [0061] Generating peak lists [0062] A peak list for polar metabolites was generated by combining the results of centwave7 peak detection (within XCMS), background subtraction, and adduct selection (CAMERA) from six pooled samples of distinct sample subsets. Additionally, a
similar workflow was performed within the AcquireX software on three additional pooled samples, and the unfiltered features from this analysis were combined with the XCMS results. The R and Python scripts used to perform the peak detection analysis are available on GitHub (https://github.com/e-stan/metabolomics_workflow) and include the values of all parameters utilized. A peak list for the lipid metabolites was directly generated based on identifications from Lipid Annotator. Any workflow or software can be used to generate peaks lists. [0063] Metabolite identification [0064] Identification of polar metabolites was supported by matching the accurate mass and MS/MS fragmentation data to an in-house MS/MS library created from authentic reference standards and online MS/MS libraries with DecoID software. For online database searching, the top hit for each feature with a dot-product similarity of greater than 80 was considered as the putative identification. These results were further filtered by using in-house retention times, predicted retention times from a method similar to ReTip, and manual curation to remove noise peaks, interfered peaks, and incorrect identifications. MSI identification levels are given in Table S1 and Table S2 for polar and lipid metabolites, respectively. Code and scripts used to perform the automated portion of the metabolite identification workflow are available on GitHub. Lipid iterative MS/MS data were annotated with the Lipid Annotator software (Agilent Technologies), and lipid identifications were provided as sum compositions because insufficient information was available to deduce specific fatty acid compositions. Lipid identifications were subject to the same manual curation as applied to the polar metabolite data. Any workflow or software can be used for compound identification. [0065] Extracting peak areas [0066] Following the generation of a peak list and metabolite identification, all data files were analyzed in Skyline (version 20.1.0.155) batch per batch to obtain peak areas. The m/z values of the metabolite target lists were used to extract peak areas under consideration of retention times or indexed retention times (iRT) (see Supporting Information). Due to the data being acquired over several months, 14 different batch correction approaches were tested for peak area normalization (see Supporting Information). Additionally, a report containing the acquisition times of all samples was exported to be used for batch correction.
[0058] Supporting Information [0059] Metabolite extraction [0060] Participant plasma, which had been collected into dipotassium ethylenediaminetetraacetic acid (K2- EDTA) collection tubes and stored at -80 °C upon collection, was thawed on ice. A 50 µL aliquot was transferred onto the solid-phase extraction (SPE) CAPTIVA-EMR Lipid 96-well plate (Agilent Technologies) before addition of 200 µL of acetonitrile:methanol (1:1) (v/v) containing 12.5 µM internal standard (consisting of uniformly 13C and 15N labeled amino acids, MSK-A2-1.2, Cambridge Isotope Laboratories, Inc). While ideally isotopically labeled internal standards for a wide variety of compounds would be used, resource constraints only allowed for an isotopically labeled amino acid mix to be added to the extraction solvent. While this is sufficient for assessing technical variation, the internal standards cannot be used for quantitation in this study. The samples were mixed for 1 min at 360 rpm on an orbital shaker at room temperature prior to a 10 min incubation period at 4 ^C. Afterwards, 150 µL of 2:2:1 acetonitrile:methanol:water (v/v/v) was added to the samples. The samples were mixed on an orbital shaker (360 rpm) for an additional 10 min at room temperature. The samples were then eluted into a 96-deepwell collection plate using a positive pressure manifold (Biotage PRESSURE+ 96). A second elution was performed with 100 µL of 2:2:1 acetonitrile:methanol:water (v/v/v). Polar eluates were stored at -80 ^C until LC/MS analysis until lipid data acquisition was complete (2.5 months). [0061] Lipid metabolites still bound to the SPE material were then released into a second elution plate, in two elution steps applying 2x 500 µL 1:1 methyl tert-butyl ether:methanol (v/v) onto the SPE cartridge and a positive pressure manifold. The combined eluates were dried under a stream of nitrogen (Biotage SPE Dry Evaporation System) at room temperature and reconstituted with 200 µL 1:12-propanol:methanol (v/v) prior to LC/MS analysis. The same SPE lot was used for all 2,000 study samples. [0062] LC/MS analysis of polar metabolites [0063] An aliquot of 4 µL of polar metabolite extract was subjected to LC/MS analysis by using an Agilent 1290 Infinity II liquid-chromatography (LC) system coupled to an Agilent 6545 Quadrupole-Time-of-Flight (Q-TOF) mass spectrometer with a dual Agilent Jet Stream electrospray ionization source. Polar metabolites were separated on a SeQuant® ZIC®-pHILIC column (100 x 2.1 mm, 5 µm, polymer, Merck-Millipore) including a ZIC®-pHILIC guard column (2.1 mm x 20 mm, 5 µm). The use of an inline
filter prior to the guard column is recommended. The column compartment temperature was maintained at 40 ^C and the flow rate was set to 250 µL ^min-1. The mobile phases consisted of A: 95% water, 5% acetonitrile, 20 mM ammonium bicarbonate, 0.1% ammonium hydroxide solution (25% ammonia in water), 2.5 µM medronic acid, and B: 95% acetonitrile, 5% water, 2.5 µM medronic acid. Medronic acid was used in mobile phase B (mainly acetonitrile) in this and previous studies. However, the method may also be practiced by adding 5 µM medronic acid to the aqueous mobile phase A, but no medronic acid in B. The following linear gradient was applied: 0 to 1 min, 90% B; 12 min, 35% B; 12.5 to 14.5 min, 25% B; 15 min, 90% B followed by a re-equilibration phase of 4 min at 400 µL ^min-1 and 2 min at 250 µL ^min-1. Polar metabolites were detected at a scan rate of 1 spectrum per second in negative ion mode with the following source parameters: gas temperature 200 ^C, drying gas flow 10 L ^min-1, nebulizer pressure 44 psi, sheath gas temperature 300 ^C, sheath gas flow 12 L ^min-1, VCap 3000 V, nozzle voltage 2000 V, Fragmentor 100 V, Skimmer 65 V, Oct 1 RF Vpp 750 V. The m/z range was 50-1700. Data were acquired under continuous reference mass correction m/z 119.0363 and 966.0007. Samples were randomized prior to analysis. In addition, a quality control (QC) sample was injected after every 12th sample to monitor signal stability of the instrument. The HILIC column was equilibrated with three blank and four QC injections prior to starting the actual samples. Prior to every batch, the mass spectrometer was calibrated and before starting to inject the actual samples, the TIC, internal standard intensities, and mass accuracy of the QC samples was compared to previous batches. More details on best practices for system suitability testing can be found elsewhere. [0064] LC/MS analysis of lipid metabolites [0065] An aliquot of 4 µL of lipid extract was subjected to LC/MS analysis by using an Agilent 1290 Infinity II LC-system coupled to an Agilent 6545 Q-TOF mass spectrometer with a dual Agilent Jet Stream electrospray ionization source. Lipids were separated on an Acquity UPLC® HSS T3 column (2.1 x 150 mm, 1.8 µm) including an Acquity UPLC® HSS T3 VanGuard Pre-Column (2.1 x 5mm, 1.8 µm) at a temperature of 60 ^C and a flow rate of 250 µL ^min-1. The mobile phases consisted of A: 60% acetonitrile, 40% water, 0.1% formic acid, 10 mM ammonium formate, 2.5 µM medronic acid, and B: 90% 2-propanol, 10% acetonitrile, 0.1% formic acid, 10 mM ammonium formate (dissolved in 1 mL water). The following linear gradient was used: 0-2 min, 30% B; 17 min, 75% B; 20 min, 85%; 23-26 min, 100% B; 26, 30% B followed by a re-
equilibration phase of 5 min. Lipids were detected in positive ion mode at a scan rate of 2 spectra per second with the following source parameters: gas temperature 250 ^C, drying gas flow 11 L ^min-1, nebulizer pressure 35 psi, sheath gas temperature 300 ^C, sheath gas flow 12 L ^min-1, VCap 3000 V, nozzle voltage 500 V, Fragmentor 160 V, Skimmer 65 V, Oct 1 RF Vpp 750 V. The m/z range was 50-1700. Data were acquired under continuous reference mass correction at m/z 121.0509 and 922.0890 in positive ion mode. Samples were randomized before analysis. In addition, a QC sample was injected after every 12th sample to monitor signal stability of the instrument. The RP column was equilibrated with three blank and three QC injections prior to starting the actual samples. The same system suitability testing described for the polar metabolite analysis was used for the lipid metabolite analysis. [0066] Preparation of pooled samples [0067] A pool of samples from the first batch was prepared by mixing 330 µL of all samples that contained at least 500 µL (58 samples). Next, 160 aliquots of 110 µL each were frozen at -80 °C, so that two 50 µL pooled QC samples could be included in every batch. This pooled QC not only served as a reference sample across batches, but it was also used for feature detection and metabolite identification. [0068] Metabolite identifications were supported by the pooled QC samples and 8 additional pooled samples that were prepared by pooling 10 µL aliquots of 10 samples of 8 random batches. The eight additional pooled samples cover the four different locations and a wide age range. Three pooled samples contained only samples from Denmark, whereas the other five were a mix of the three different US locations. The average age ranged from 69.1 to 87.9 years, with an average standard deviation of 13.3 years. Across those pooled samples, 52% were from male and 48% from female participants. MS/MS data were acquired in negative mode (see details below) on 9 total pooled samples. Peak detection to generate a peak list for DecoID was performed on six pooled samples by using XCMS and three pooled samples by using AcquireX. [0069] Acquiring MS/MS data [0070] MS/MS spectra for polar and lipids metabolites were acquired by using an iterative data dependent acquisition (iDDA) approach in the MassHunter Acquisition Software (Version 10.1.48, Agilent Technologies) on an Agilent 6545 QTOF. The same source settings as for MS1 data acquisition were used. MS/MS spectra were acquired at a scan rate of 3 spectra/s with different intensity thresholds and collision
energies of 10, 20, and 40 V to increase identification rates. iDDA can be readily implemented by creating consecutive acquisition methods for the blank and study samples in the iterative workflow. Thus, exclusion lists no longer need to be manually generated after each run. Details on how to achieve iDDA on the Agilent system can be found in an application note. [0071] Additionally, MS/MS data for polar metabolites were acquired on an Orbitrap ID-X Tribrid mass spectrometer (Thermo Scientific). A Vanquish Horizon UHPLC system, with the same chromatographic conditions as described above, was interfaced with the mass spectrometer via electrospray ionization in negative mode with a spray voltage of 2.8 kV. The following source conditions were used: sheath gas flow 50 arbitrary units (Arb), auxiliary gas flow 10 Arb, sweep gas flow 1 Arb, ion transfer tube temperature 300 ^C, and vaporizer temperature 200 ^C. The RF lens value was 60%. Data were acquired in data dependent acquisition (DDA) mode by using the built-in deep scan option (AcquireX) with a mass range of 67-900 m/z and 120K resolution for MS1 scans. MS/MS scans were acquired at 15K resolution from three distinct pooled samples. [0072] Indexed retention times [0073] To adjust for sample-specific retention time drifts in lipid metabolites, indexed retention times (iRT) were implemented. iRT uses the retention time of selected “indexing” compounds to adjust the retention times of other compounds. iRT were implemented for each lipid class, selecting 2-3 lipid metabolites per class as indexing compounds (see Table 3). In selecting indexing compounds, lipids were chosen that were chromatographically well separated from other species and that had high intensity. [0074] The first step of the process is to compute the iRT of each lipid in a class. The calculation relies on the retention times of each lipid as measured from a single sample. As shown in Equation 1, iRT uses the retention time of the first (^^^^^) and last (^^^^^) eluting lipid in a class to convert a measured retention time for a compound (^^^) to its corresponding iRT (^^^^). ^^^^ = (1)
[0075] Next, for each subsequent sample, the iRT of the indexing compounds (^) and the observed retention times for the indexing compounds in the sample
are used to calculate a linear regression of iRT to retention time by minimizing the error between observed and fit retention times, according to Equation 2.
[0076] While others have used Lowess regression to convert iRTs to retention times, it was found that linear regression was able to accurately estimate retention times from iRT values, suggesting that retention-time drifts are linear within each class of lipids. [0077] Lastly, the adjusted retention times (^^^^^) are calculated for all non-indexed compounds in the lipid class by using the linear regression and the calculated iRTs for these compounds, according to Equation 3.
[0078] This process was repeated for each lipid class and the correction was carried out for each sample. The iRT adjustment is implemented in Python and the relevant Google Colab link can be found on GitHub. [0079] Batch-correction evaluation [0080] Fourteen batch-correction methods were evaluated in this study: L1 (“constant sum”), L2 (“constant length”, median, QC normalization, ComBat, combined ComBat and QC normalization, support vector regression (SVR), linear regression, random forest, probabilistic quotient normalization (PQN), WaveICA, EigenMS, gaussian (normal) quantile transform, and uniform quantile transform. In L1, L2, and median normalization, the L1 norm, L2 norm, or median of each sample’s metabolic profile was computed and was used to normalize the metabolite intensities for the sample. In QC normalization, each metabolite was normalized by computing the mean of the metabolite intensity in the flanking QC samples for a particular research sample and multiplying the metabolite’s intensity in this research sample by the mean of the metabolite’s intensity across all QC samples and dividing by the mean of the flanking QC samples. In combined ComBat and QC normalization, each batch of samples was QC normalized individually followed by an inter-batch correction with ComBat. In SVR, linear regression, and random forest, the respective model was fit for each metabolite by using the batch and run-order position of the QC samples against the deviation in the metabolite’s intensity from the mean QC
intensity. The fitting procedure for these models was carried out by using Scikit-learn with default parameters. After fitting, the deviation was predicted for all QC and research samples and this deviation was subtracted from the intensity of each sample. For all correction methods except QC normalization, metabolite intensities are log2 transformed prior to normalization. For QC normalization, metabolite intensities are log2 transformed after normalization. Python implementations of the normalization algorithms, accompanying Google Colab links, and the code for performing the comparative analysis are available on GitHub. [0081] To assess the performance of each method, a metric was computed from the coefficients of variation (CV) in both the unnormalized and normalized data. First, in the unnormalized data, the CV of each metabolite amongst the QC and research samples was separately computed. Then, these values were again computed in the normalized data. The research CV values in the normalized data were divided by the research CV values in the unnormalized data, and the QC CV values in the normalized data were divided by the QC CV values in the unnormalized data. These two quantities were then divided by each other (change in research CV divided by the change in QC CV) to give the normalization score for each metabolite. The mean normalization score across all metabolites was used to score the overall performance of the method. CV values were computed on non-log2 transformed metabolite intensities. As a secondary performance metric for evaluating batch correction methods on the polar metabolite data, the intensity CVs were calculated, both inter- and intra-batch, of the internal standards spiked into each sample. Since these CV values were calculated across all samples, not just QC samples, this approach enables potential “overfitting” of the batch correction methods to be detected. [0082] Analysis of unknown metabolites [0083] Of the 6,036 features detected from the pooled samples, 3,421 high- quality features were kept after manual inspection of the extracted ion chromatograms. The peak areas for these high-quality features were then extracted with Skyline-daily (version 21.2.1.485), imputed with the half-minimum approach, and normalized with the random forest algorithm. Features whose CV value across the QC samples was greater than 10% were removed from downstream analysis. Additionally, to remove redundancy in the data, features that had a Pearson correlation across all samples of greater than 0.95 with another unknown feature or identified metabolite were excluded from statistical analysis. The
remaining features were tested for association with field site by using a one-way ANOVA with a Bonferroni correction. [0084] Results and Discussion [0085] The presented strategy to analyze large cohorts with untargeted metabolomics is based on the observation that most of the features in an experiment do not correspond to unique metabolites of biological relevance. Rather than attempting to evaluate each of these features in every sample, a small number of pooled reference samples was used to annotate features of interest. The process reduces the data burden of untargeted metabolomics such that informatics tools typically applied to targeted studies can be leveraged to profile research samples efficiently and rapidly, without the need to subject each sample to computationally intensive analyses (e.g., peak detection, correspondence determination, peak grouping, metabolite identification, etc.). A schematic of the workflow is shown in FIG.3. As a demonstration, the disclosed workflow was used to analyze a subset of ~2,000 human plasma samples from the Long Life Family Study (LLFS), conducted by the National Institute on Aging at the National Institutes of Health (NIH), Bethesda, Maryland, U.S. A description of each step of the workflow is provided below, with further details in the Experimental Section and Supporting Information. [0086] Sample preparation and data acquisition [0087] Before LC/MS data were acquired for the LLFS sample set, the 2,005 plasma samples were organized into 22 batches (mostly 92 samples per batch). Longitudinal samples from the same participant were included within the same batch. Polar and lipid metabolites were extracted from plasma samples into 96-well plates by using solid- phase-extraction. A pooled sample was prepared from the first batch of the LLFS sample set to serve as a QC sample. All samples were analyzed via LC/MS. Additional pooled samples (see Supporting Information) were prepared and analyzed via LC/MS/MS for metabolite identification. [0088] Data processing for pooled samples and subsequent data extraction [0089] The data collected from pooled samples were subjected to a standard processing workflow for untargeted metabolomics. Namely, feature detection, grouping, filtering of background and degeneracy, and MS/MS-based compound identification was applied to form a peak list of putatively identified metabolites that are suitable for multi- omics integration. The peak list of identified polar and lipid metabolites can be found in Table 1 and Table 2, respectively. To demonstrate that this workflow is compatible with
analysis of unknowns at the scale typically seen in untargeted metabolomics studies, more than 3,000 unidentified features were added to the peak list. [0090] After generating the peak list, peak areas were extracted from the research and QC samples in a batch-by-batch fashion by using Skyline . The Skyline command-line interface can be used to automatically generate a document for each batch. For polar metabolites separated with HILIC, retention times were stable for all samples within a batch (FIGS.4a & 4b). Thus, retention-time bounds were set by inspection of the QC samples within each batch, and these bounds were applied to all samples by importing peak boundaries for each sample. Although a Python script was used to generate the peak boundary import file in the current work, this is no longer necessary when using the “Synchronize Integration” function in Skyline 21.2, which was released after the instant data processing was performed. [0091] Retention-time values were generally stable across all experiments. Even though lipid metabolites tended to show more drift than polar metabolites, only five samples from all of the lipid runs had retention-time deviations greater than 0.25 min. Among those compounds that showed retention-time drifts, lysophosphatidylcholines and lysophosphatidylethanolamines were most pronounced in specific samples within each batch, which may be attributed to matrix effects (FIGS.5a-5c). Accordingly, the concept of iRT, which was initially established for proteomics and recently integrated into lipidomics, was applied. In brief, two to three compounds per lipid class that are of high intensity and easily distinguished from nearby peaks were chosen as indexing compounds. When importing all data files into Skyline in centroid mode, the peaks for the indexing compounds were consistently picked correctly. Incorrect peak assignments were corrected by inspection of the retention time replicate comparison plots and manual adjustment of peak boundaries to ensure that the apex of all indexing peaks was within the integration boundaries. The retention times for the indexing compounds were then used in a Python script to adjust and generate peak boundaries for all compounds in all samples (further details are provided in the Supporting Information). Upon importing those boundaries, correct integration was manually verified before peak areas were exported. It was found that this strategy can efficiently correct RT shifts of over one minute. In total, after peak area extraction in Skyline, 172 identified polar metabolites, 188 identified lipid metabolites, and 3,421 unidentified features from the polar data were profiled across 2,001 research samples and 197 QC samples in the LLFS sample set. Of note, 4 research samples were excluded from downstream
analysis because of uncharacteristically low signal abundance. Extracting peak areas with the process described above takes an experienced researcher 20-30 minutes per batch of samples analyzed. The major time investment is defining and curating the peak list that is formed from analysis of the pooled sample. In the case of LLFS, defining a peak list took several weeks and included acquisition of extensive MS/MS data, manual removal of artifact peaks and interferences, and review of metabolite identifications. The time required to form this peak list will depend on the study and the number of non-biological and redundant signals filtered. [0092] Comparing metabolite coverage from a pooled sample to individually measured samples [0093] Limiting data processing to pooled samples greatly reduces the computational burden of untargeted metabolomics, but risks that low-abundance compounds only present in a small number of research samples will be missed. To determine the number of features potentially missed by only performing data processing on pooled samples, a subset of 58 research samples from LLFS was evaluated. For this analysis, nine replicate injections of a pooled sample were created by mixing small aliquots of each of the 58 research samples. Then, feature detection, background subtraction, and selection for M +/- H ions on each of the 58 research samples, the 9 pooled samples, and 4 blank samples was performed. Peak detection was achieved with centwave7, and features with peak areas less than 10,000 in a particular sample were classified as undetected in that sample. Background subtraction was accomplished by removing features having an intensity in the blank sample that was lower than 1/3 the intensity of the research or pooled sample. M +/- H ions were selected by applying the CAMERA software package9. [0094] After applying the disclosed data-processing workflow, a list compiling all of the features that were detected from each of the 58 research samples was made. The list, which was comprised of 5,894 total features, included features that were only detected from a single research sample. A total of 3,241 features (both identified and unknowns) from the list were detected in at least one replicate of the pooled sample (FIG. 6a). To assess the biological relevance of the features not detected in the pooled sample, the m/z values of all features were searched against endogenous metabolites in the Human Metabolome Database and the Kyoto Encyclopedia of Genes and Genomes. In total, 40.4% of the features detected in the pooled sample had at least one hit in the databases. Of the features missed in the pooled sample, only 32.2% had a database hit, suggesting that these
missed features were more likely to be exogenous metabolites (e.g., drugs, environmental toxins, cosmetics). Additionally, features not detected in the pooled sample were detected on average in less than 10% of the individual samples (FIG.6b). In contrast, features detected in the pooled samples were present in the majority of the individual samples (FIG. 6c). Features not detected in the pooled samples were also an order of magnitude lower in abundance (FIGS.6d & 6e). [0095] These data reveal that, while there is a reduction in the number of detected features when using a pooled sample, the missing features are above the limit of detection in only a small subset of research samples. It has been suggested previously that features not detected in at least 70% of the samples be excluded from the untargeted metabolomics analysis, irrespective of the data-processing workflow. With this threshold, the number of features missed by only subjecting a pooled sample to data processing would be reduced to just 52 (<1% of total features). Another major complication of pursuing features that cannot be detected in the pooled samples is that they are challenging to normalize with respect to batch effects and technical variation. Thus, even though features detected in a small number of samples might be of biological interest, more sensitive methods will need to be developed to evaluate them. At the present time, no matter which data- processing workflow is applied, these features are not well suited for large-scale studies. [0096] Data post-processing [0097] After extracting peak areas, missing values must be removed from the data to facilitate downstream processing. In the workflow using pooled samples, due to the targeted extraction of peak areas, missing values were infrequent (< 0.03% of all measurements) and most likely arose from metabolites at concentrations below the limit of detection of the instrument rather than random metabolite dropout during peak detection. Thus, missing values were imputed with the half minimum approach. When comparing the frequency of missing values produced with targeted extraction and conventional XCMS- based processing, it was found that 9.2% of all peak areas for a single batch of the LLFS samples were missing values when XCMS was employed. In contrast, less than 0.001% of all measurements were missing values when targeted extraction was applied to the same sample set (FIGS. 7a-c). Additionally, the technical variation in the output peak areas was 5x lower with targeted extraction of peak intensities when compared to XCMS (FIG.7c). [0098] Given the size of the LLFS study, the raw data had to be acquired over several months. When combining the extracted data from each group of 92 research
samples, strong batch effects were observed as can be seen in FIG. 8a and FIG. 9a for identified lipid and polar metabolites, respectively. Unfortunately, samples from human subjects for large studies cannot always be randomized into analysis batches because of practical limitations such as longitudinal sample collection or funding timeline constraints. As a result, differences between batches may be due to technical or biological variation (FIGS. 10a & 10b). For example, in the case of the LLFS sample set, the last six batches consisted of samples primarily collected in Denmark, while the other batches consisted largely of samples collected in the United States. [0099] To discriminate between biological and technical variation, identical QC samples across all batches are critical to evaluating and guiding batch-correction methods. Here, 14 different normalization algorithms were tested for their capability to minimize technical variation while keeping biological variance intact (details of this comparison are provided in the Supporting Information). The results of the analysis revealed that the performance of many of the methods is dataset dependent (FIG.8b and FIG.9b), but a random forest-based batch correction algorithm outperformed the other evaluated approaches in both the lipid and polar metabolite data. As a secondary validation, the internal standard variability in the polar metabolite data was also compared. It was again observed that the random forest-based correction reduced both intra- and inter-batch variability (FIGS. 11a-d). When comparing the performance of random forest to the other evaluated methods, it was found that QC, ComBat, and QC+ComBat performed similar to random forest in reducing internal standard CV values, with ComBat+QC achieving the lowest variability (FIG.11e). Overall, although random forest performed well on the data, it may be useful to test different batch-correction methods before selecting the algorithm to apply to a dataset. Such an evaluation can be performed by adapting the code written for analysis, which is available on GitHub. After correction, previously batch-affected metabolites no longer showed batch-dependent intensity drifts, and technical variation within the data was reduced (FIGS. 8c-e, FIGS.9c-e). Although some batch-associated clustering remained in the polar metabolic profiles, when examining only the QC samples (FIGS.12a-d), no such clustering occurred. These findings indicate that biological differences between batches are driving the observed patterns, rather than technical variability. After batch correction, the metabolic profiles can then be subjected to downstream statistical analysis to identify interesting biological patterns. [0100] Metabolic profiles cluster by geographic location
[0101] Collection of untargeted metabolomics data for the entire LLFS cohort is still underway, but it was desired to demonstrate that the described workflow results in metabolic profiles containing biologically relevant information. As such, differences in the profiles of polar and lipid metabolites that could be attributed to the geographic origin of the samples were investigated. The LLFS samples were collected in four distinct field sites: Boston, MA; Pittsburg, PA; New York City, NY; and Odense, Denmark. Given the differences in dietary habits between the United States and Denmark, it was surmised that the plasma metabolic profiles would be different between samples collected in the United States and Denmark. In fact, when analyzing the unnormalized metabolic profiles of the identified metabolites profiled in these data, 72 metabolites had maximum absolute fold changes greater than two between the field sites and p-values of less than 0.05 (one-way ANOVA). Given that samples from the field sites were not uniformly distributed across the sample batches, however, technical variation may artificially introduce separation to their metabolic profiles. Indeed, after performing random forest batch correction, only 45 metabolites met the same statistical cutoffs, demonstrating the importance of batch correction to interpreting metabolomics data. These 45 metabolites led to pronounced clustering between the United States and Denmark samples (FIG. 13a). Several of the differentiating lipid metabolites also contain multiple unsaturations (e.g., CE 22:5, DG 36:4, LPC 20:5, PC 37:5, PC 38:6, PC 40:7, TG 56:7, TG 58:7, TG 58:8, TG 60:11, TG 60:12) that may reflect differential abundance of omega-3 and omega-6 polyunsaturated fatty acids (DHA, EPA, linoleic acid, etc.) derived from the diet (FIG. 13b). Scandinavian countries have been shown to have elevated levels of circulating omega-3 fatty acids when compared to the United States and other countries with westernized food habits, which is potentially due to higher per capita consumption of fish and shellfish. Of the polar metabolites, the most striking difference was inosine (FIG.13c). Inosine is a known dietary metabolite with high concentrations in cow milk. This result indicates a difference in the amount of dairy consumption between individuals in the United States and Denmark, as has been suggested before. [0102] Although the identified metabolites were the main focus of the analysis for LLFS, where metabolomics data will be linked to corresponding gene and protein measurements, the existence of differences in unknowns between the sample groups was investigated. After filtering (see Supporting Information) and performing statistical analysis of >3,000 unknowns from the polar metabolite extracts profiled in this study, 29
unique unknowns that were associated with field site location (|FC| > 2, p < 0.05, One-way ANOVA) were found, as shown in FIGS.14a-d. The differences between fields sites were dramatic, but participants from Denmark were also on average younger than those from the United States (p <0.0001) and this could have contributed to the observed metabolic profiles (FIG.13d). Notably, however, a principal component analysis of those samples colored by age did not show an age-dependent pattern within United States or Denmark samples (FIG. 15), suggesting that the differences cannot solely be attributed to age. An additional potential confounding factor is that the time between blood collection and centrifugation varied between samples and field sites. In a previous study, it was shown that this delay can cause alterations in metabolite levels. Further statistical analysis of this dataset will be performed that considers covariates and genetic relationships between LLFS participants once data collection for LLFS is complete. [0103] Conclusions [0104] The trend in the omic sciences is to evaluate increasingly large sample sizes approaching tens of thousands to hundreds of thousands of specimens. Larger sample groups increase statistical power and may enable a study to distinguish between sub- populations within a group that have a unique treatment effect, which is the vision of precision medicine. The challenge of using exceptionally large sample cohorts, on the other hand, is the burden of collecting and processing high volumes of data. Applications of untargeted metabolomics to large sample cohorts have been limited because standard software programs are not compatible with population-based studies. This work describes a workflow designed to perform untargeted metabolomics on >12,000 human plasma samples from the LLFS cohort. [0105] To reduce to the computational burden of performing global processing of all HRMS data files within the cohort, the disclosed method focuses on processing data from a small number of pooled samples that are created by mixing aliquots of individual samples from the study. By initially limiting data processing to only the pooled samples, the disclosed method leverages standard informatics tools in untargeted metabolomics that have been optimized for small sample sizes. This not only facilitates feature detection, but also leads to a considerable data reduction from removal of adducts, background signals, and other degeneracies. Analysis of the pooled sample does require a significant time investment, but it can be completed in parallel with data acquisition for the research samples. Further, because peak areas for metabolites detected in the pooled sample
can be extracted from the research samples as data are generated, the rate limiting factor in the described workflow is data acquisition rather than data processing. This contrasts with conventional processing workflows that require all data to be acquired prior to analysis and curation. [0106] In preparation of the reference sample, care should be taken to ensure that pooled samples cover different sample groups and represent the biological diversity present in the study, such as age and gender. The pooled samples are intended to capture the totality of unique compounds across the entire study cohort but, in practice, unique compounds from individual samples are missed. An analysis of the data shows that unique compounds are most often missed when they are only present in a few participant samples at low concentrations, which causes them to be diluted below the limit of detection during pooling. The results indicate that signals missing from the analysis of pooled samples are more likely to originate from rare exogenous compounds (e.g., chemicals from unique hygiene products or specific environments). While rare exogenous compounds may certainly be biologically interesting, their analysis will require the development of new peak detection, alignment, and annotation algorithms that can be scaled to thousands of samples without loss of functionality or accuracy. It should be noted that, even when conventional informatics workflows are used for untargeted metabolomics, features only detected in a low percentage of samples are usually discarded because of the difficulties in correcting for technical drift. [0107] When processing untargeted metabolomics data with conventional workflows, retention times are aligned by using correspondence algorithms such as Obiwarp. Here, sample specific retention-time drifts were corrected using the application of iRT. This procedure requires some manual intervention from the user to ensure correct peak detection of indexing compounds and to adjust batch-to-batch variations in retention time. The instant disclosure demonstrates that the iRT approach corrected for retention-time drifts in lipid metabolites, but it was not necessary for polar metabolites because intra-batch retention-time drifts in these data were not observed. Should retention-time drifts be observed for polar metabolites, the iRT approach could easily be extended by determining groups of compounds that have correlated retention-time drifts across samples and using them for indexing. [0108] In large-scale longitudinal studies such as LLFS, it may not be feasible to fully randomize the order in which specimens are analyzed by LC/MS. Correcting for the technical variation between batches of samples is therefore required to measure biological differences accurately. Here, 14 different approaches for batch correction were
evaluated and it was found that accounting for batch effects by using a random forest normalization to estimate drift in each reference chemical was the most effective. It should be noted that while this study used a limited number of internal standards, incorporating additional compounds that cover a broader range of chemical classes may be beneficial for assessing the possibility of metabolite degradation and for evaluating the overall quality of batch correction results. [0109] In summary, the present disclosure enables high-throughput applications of untargeted metabolomics at the population scale, such as observed in the LLFS. As a further benefit, the disclosed workflow collects separation and mass spectrometry features for all individual samples in the sample set. Thus, as new computational resources become available, methods disclosed herein may be applied using more advanced technology.
15060-1625 (020387/WO) PCT Table 1 l
15060-1625 (020387/WO) PCT
15060-1625 (020387/WO) PCT
15060-1625 (020387/WO) PCT
15060-1625 (020387/WO) PCT
Table 2
15060-1625 (020387/WO) PCT
15060-1625 (020387/WO) PCT
CT
CT
CT
15060-1625 (020387/WO) PCT
CT
CT
15060-1625 (020387/WO) PCT Table 3
Claims
WHAT IS CLAIMED IS: 1. A method of non-targeted determination of unique biological metabolites in individual samples of a sample set, each individual sample comprising a composition of chemical constituents, the method comprising: a. analyzing reference features obtained from a reference sample to identify non-relevant reference features; b. filtering the reference features by removing the non-relevant reference features, thereby producing a set of relevant reference features that characterize the unique biological metabolites; and, c. applying the non-relevant and/or relevant reference features to sample features obtained from an individual sample to determine the composition of unique biological metabolites in the individual sample. 2. The method of claim 1, comprising conducting the steps of analyzing, filtering and applying on each individual sample in the sample set, thereby determining the unique biological metabolites in all individuals of the sample set. 3. The method of claim 2, wherein after the steps of analyzing, filtering and applying have been conducted on a plurality of individual samples in the sample set, updating the reference features by re-analyzing the reference sample to obtain updated reference features and replacing the reference features with the updated reference features, wherein the number of individual samples in the plurality of samples is less than the entire number of individual samples in the sample set. 4. The method of claim 3, wherein the plurality of samples comprises at least 2, samples, optionally at least 3 samples, optionally at least 5 samples, optionally at least 10 samples, optionally at least 20 samples, optionally at least 50 samples, optionally at least 100 samples, or optionally at least 150 samples. 5. The method of any one of claims 1-4, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples. 6. The method of any one of claims 1-5, wherein the reference features are obtained by a method comprising subjecting the reference sample to a separation technique and mass spectrometry.
7. The method of any one of claims 1-6, wherein the non-relevant reference features characterize artifacts, background chemical constituents and non- unique biological metabolites. 8. The method of claim 7, wherein the background chemical constituents comprise contaminants. 9. The method of claim 7 or 8, wherein the non-unique biological metabolites comprise adducts and/or fragment ions. 10. The method of any one of claims 6-9, wherein the separation technique comprises chromatography. 11. The method of claim 10, wherein the chromatography comprises gas chromatography or liquid chromatography. 12. The method of any one of claims 6-11, wherein the reference features comprise reference separation data and reference mass data generated from the reference sample. 13. The method of any one of claims 1-12, wherein the step of applying comprises: a. identifying in the sample features those sample features corresponding to non-relevant reference features and removing those sample features that correspond to the non-relevant reference features to produce a set of relevant sample features; and, b. using the relevant sample features the determine the unique biological metabolites in the sample. 14. The method of any one of claims 1-12, wherein the step of applying comprises: a. identifying in the sample features those sample features corresponding to relevant reference features to produce a set of relevant sample features; and, b. using the relevant sample features to determine the unique biological metabolites in the sample. 15. The method of any one of claims 1-14, comprising conducting the steps of obtaining and applying to each sample in the sample set, thereby determining the composition of the chemical constituents in all individual samples in the sample set. 16. The method of any one of claims 1-15, comprising identifying one or more of the chemical constituents by comparing the relevant sample features to a library of information comprising features characterizing chemical entities.
17. A method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a. obtaining reference features comprising separation data and mass data generated from a reference sample; b. filtering the reference features by eliminating separation data and mass data relating to non-unique chemical constituents to produce filtered reference features comprising filtered separation data and filtered mass data; c. using the filtered reference features to characterize the presence and/or quantity of unique chemical constituents in each individual sample of the sample set. 18. The method of claim 17, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples. 19. A method of non-targeted determination of a composition of chemical constituents in individual samples in a sample set, comprising: a. using a separation technique and mass spectrometry to produce reference features comprising separation data and mass data from a reference sample; b. collecting and storing the reference features; c. analyzing the stored reference features to identify non-relevant reference features comprising separation data and mass data characterizing non- unique chemical constituents; d. removing the non-relevant reference features comprising separation data and mass data characterizing non-unique chemical constituents to produce a filtered data set comprising filtered reference features comprising separation data and mass data characterizing unique chemical constituents; e. using the filtered reference features to characterize the presence and/or quantity of unique chemical constituents in each sample of the sample set. 20. The method of claim 19, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples. 21. The method of claim 19 or 20, wherein the step of using the filtered reference features comprises: a. obtaining sample features from an individual sample;
b. identifying sample features that correspond to the filtered reference features to produce a list of relevant sample features that characterize unique chemical constituents, thereby determining the composition of chemical constituents in the sample. 22. A system for conducting the method of any one of claims 1-21, the system comprising: a. a separation apparatus for separating the chemical constituents and producing separation-related features characterizing the chemical constituents; b. a mass spectrometer for performing mass spectrometry on portions of the separated chemical constituents and producing mass spectrometry-related features characterizing the chemical constituents; c. a first module for receiving, collecting and/or storing the separation-related features and/or the mass spectrometry -related features; and, d. a user interface coupled to the module for making the separation-related features and/or the mass spectrometry -related features available in human- accessible form. 23. The system of claim 22, comprising a library of features characterizing chemical entities, the features produced using the separation apparatus and the mass spectrometer, wherein the library of features comprises separation features and mass spectrometry features characterizing identified chemical entities. 24. The system of claim 22 or 23, wherein separation apparatus comprises an apparatus for performing chromatography. 25. The system of claim 24, wherein the chromatography comprises liquid chromatography. 26. The system of claim 25, wherein the liquid chromatography comprises HPLC and UPLC 27. The system of claim 22 or 23, wherein separation apparatus comprises an electrophoresis apparatus. 28. The system of any one of claims 22-24, wherein the separation apparatus is coupled to the mass spectrometer. 29. The system of any one of claims 22-28, wherein the separation-related features comprise peak retention time, peak intensity, and/or peak width.
30. The system of any one of claims 22-29, wherein the mass-spectrometry-related features comprise mass or an m/z value.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263386087P | 2022-12-05 | 2022-12-05 | |
| PCT/US2023/082308 WO2024123681A2 (en) | 2022-12-05 | 2023-12-04 | Improved method for scalable untargeted metabolomic workflow |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4630813A2 true EP4630813A2 (en) | 2025-10-15 |
Family
ID=91380079
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23901396.4A Pending EP4630813A2 (en) | 2022-12-05 | 2023-12-04 | Improved method for scalable untargeted metabolomic workflow |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4630813A2 (en) |
| JP (1) | JP2026502062A (en) |
| AU (1) | AU2023391474A1 (en) |
| WO (1) | WO2024123681A2 (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN120820662B (en) * | 2025-09-17 | 2025-11-18 | 岛津企业管理(中国)有限公司 | Method for synchronously analyzing and detecting metabolites and lipids in aquatic products |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2004040407A2 (en) * | 2002-10-25 | 2004-05-13 | Liposcience, Inc. | Methods, systems and computer programs for deconvolving the spectral contribution of chemical constituents with overlapping signals |
| ES2452025T3 (en) * | 2005-06-30 | 2014-03-31 | Biocrates Life Sciences Ag | Apparatus for analyzing a metabolite profile |
| EP2270699A1 (en) * | 2009-07-02 | 2011-01-05 | BIOCRATES Life Sciences AG | Method for normalization in metabolomics analysis methods with endogenous reference metabolites |
| CN111989747A (en) * | 2018-04-05 | 2020-11-24 | 伊耐斯克泰克-计算机科学与技术系统工程研究所 | Spectrophotometry and apparatus for predicting quantification of a component in a sample |
-
2023
- 2023-12-04 WO PCT/US2023/082308 patent/WO2024123681A2/en not_active Ceased
- 2023-12-04 AU AU2023391474A patent/AU2023391474A1/en active Pending
- 2023-12-04 JP JP2025532598A patent/JP2026502062A/en active Pending
- 2023-12-04 EP EP23901396.4A patent/EP4630813A2/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024123681A2 (en) | 2024-06-13 |
| AU2023391474A1 (en) | 2025-06-19 |
| JP2026502062A (en) | 2026-01-21 |
| WO2024123681A3 (en) | 2024-07-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Muthubharathi et al. | Metabolomics: small molecules that matter more | |
| Courant et al. | Basics of mass spectrometry based metabolomics | |
| US8068987B2 (en) | Method and system for profiling biological systems | |
| Lindon et al. | Metabonomics in pharmaceutical R & D | |
| Dunn et al. | Molecular phenotyping of a UK population: defining the human serum metabolome | |
| US7653496B2 (en) | Feature selection in mass spectral data | |
| Harlina et al. | Possibilities of liquid chromatography mass spectrometry (LC-MS)-based metabolomics and lipidomics in the authentication of meat products: a mini review | |
| Wanichthanarak et al. | Accounting for biological variation with linear mixed-effects modelling improves the quality of clinical metabolomics data | |
| Arulvasan et al. | High-quality identification of volatile organic compounds (VOCs) originating from breath | |
| Southam et al. | Characterization of monophasic solvent-based tissue extractions for the detection of polar metabolites and lipids applying ultrahigh-performance liquid chromatography–mass spectrometry clinical metabolic phenotyping assays | |
| Jiang et al. | An integrated metabonomic and proteomic study on Kidney-Yin Deficiency Syndrome patients with diabetes mellitus in China | |
| Wang et al. | LC-MS-based metabolomics discovers purine endogenous associations with low-dose salbutamol in urine collected for antidoping tests | |
| Deda et al. | GC-MS-based metabolic phenotyping | |
| Stojiljkovic et al. | Evaluation of horse urine sample preparation methods for metabolomics using LC coupled to HRMS | |
| Shi et al. | MS based foodomics: An edge tool integrated metabolomics and proteomics for food science | |
| AU2023391474A1 (en) | Improved method for scalable untargeted metabolomic workflow | |
| Wishart et al. | Metabolomics | |
| Kartsova et al. | Current role of modern chromatography with mass spectrometry and nuclear magnetic resonance spectroscopy in the investigation of biomarkers of endometriosis | |
| Peng et al. | A metabonomic analysis of serum from rats treated with ricinine using ultra performance liquid chromatography coupled with mass spectrometry | |
| Rowles III et al. | Decoding cancer across scales with metabolomics | |
| Sunyer-Caldú et al. | Screening of Biological Samples with HRMS to Evaluate the External Human Chemical Exposome | |
| Gonzalez-Riano et al. | Advanced lipidomics using UHPLC-ESI-QTOF-MS/MS reveals novel lipids in hibernating syrian hamsters | |
| Lei et al. | SpecLipIDA: a pseudotargeted lipidomics approach for polyunsaturated fatty acids in milk | |
| Liu et al. | Analysis of the lipidomic profile of vegetable oils and animal fats and changes during aging by UPLC-Q-exactive orbitrap mass spectrometry | |
| Çelebier et al. | Recent developments in CE-MS based metabolomics |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250625 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |