EP4695611A1 - Label-free multiplex proteotyping of microbial isolates - Google Patents
Label-free multiplex proteotyping of microbial isolatesInfo
- Publication number
- EP4695611A1 EP4695611A1 EP24717246.3A EP24717246A EP4695611A1 EP 4695611 A1 EP4695611 A1 EP 4695611A1 EP 24717246 A EP24717246 A EP 24717246A EP 4695611 A1 EP4695611 A1 EP 4695611A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- sample
- fraction
- fractions
- peptide
- organisms
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
- G01N33/6803—General methods of protein analysis not limited to specific proteins or families of proteins
- G01N33/6848—Methods of protein analysis involving mass spectrometry
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K1/00—General methods for the preparation of peptides, i.e. processes for the organic chemical preparation of peptides or proteins of any length
- C07K1/14—Extraction; Separation; Purification
- C07K1/16—Extraction; Separation; Purification by chromatography
- C07K1/20—Partition-, reverse-phase or hydrophobic interaction chromatography
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/02—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving viable microorganisms
- C12Q1/04—Determining presence or kind of microorganism; Use of selective media for testing antibiotics or bacteriocides; Compositions containing a chemical indicator therefor
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
- G01N33/6803—General methods of protein analysis not limited to specific proteins or families of proteins
- G01N33/6842—Proteomic analysis of subsets of protein mixtures with reduced complexity, e.g. membrane proteins, phosphoproteins, organelle proteins
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N2570/00—Omics, e.g. proteomics, glycomics or lipidomics; Methods of analysis focusing on the entire complement of classes of biological molecules or subsets thereof, i.e. focusing on proteomes, glycomes or lipidomes
Definitions
- tandem mass spectrometry-based proteotyping has been found to be a valuable complementary methodology to quickly classify atypical organisms, which has been shown to be applicable even with mixtures of microorganisms [2] [3], This approach is well adapted to culturomics [4], where large numbers of microorganisms are isolated by multiplying growth conditions and varying culture media.
- tandem mass spectrometry-based proteotyping is based on low molecular-weight peptides rather than whole proteins, this methodology is much more sensitive than wholecell MALDI-TOF MS. Furthermore, it provides very precise taxonomical identification through the identification of thousands of peptide sequences. However, it can present a certain number of difficulties depending on the sample’s type and composition, especially when several samples have to be analyzed. To gain throughput of tandem mass spectrometry coupled to reverse phase chromatography, multiplexing could be in principle achieved by labeling the proteins from the organisms present in the different samples with distinguishable isobaric mass tags. This method would be in principle very expensive and requires a number of manipulations at the sample preparation stage.
- the cost and speed of tandem mass spectrometry is not optimized for successful applications involving a great number of samples containing numerous different organisms.
- the next expected step would be to improve the throughput of MS/MS proteotyping by optimizing the costs and the duration of the mass spectrometry measurement per sample. For example, a method enabling to analyze multiple samples in a single nanoLC-MS/MS run without the need of isobaric mass tag reagent would be ideal, as it would save instrument time, therefore spare time and money.
- the present invention solves this need, by providing a multiplexing method involving the analysis of several samples in a single tandem mass spectrometry run, resulting in the analysis of at least 21 samples done in about 60 minutes, without any need for the use of labeling reagents.
- the examples below prove the performance of the approach for the simultaneous analysis of 21 bacterial isolates with a single 60-min gradient nanoLC- MS/MS run. This approach significantly decreases the per-isolate cost of identification, and is highly sensitive and reproducible.
- the present invention concerns an innovative label-free multiplexing method for the identification of microorganisms using tandem mass spectrometry, based on off-line reverse-phase fractionation of individual peptidomes. Multiplexing is achieved by mixing fractions of staged hydrophobicity, so that each sample can be mapped to specific elution times.
- multiplexing is achieved by mixing fractions of staged hydrophobicity, so that each sample can be mapped to specific elution times.
- up to 21 different samples have been analyzed by reverse-phase nanoLC-MS/MS in a single run, and the 21 different microorganisms present in these 21 different samples have been appropriately identified by phylopeptidomic proteotyping.
- these 21 microorganisms were identified in a single 60-min analytical run of a tandem mass spectrometer coupled to reverse phase chromatography, resulting in a rate of one microorganism identified per 3 min of mass spectrometry time, without any need of labeling reagents.
- its capacity to simultaneously analyze 21 isolates within a single 60-min gradient nanoLC-MS/MS run resulted in significant time and cost savings compared to 21 separate nanoLC-MS/MS analyses.
- Taxon-specific peptides are sequences that are uniquely found in a given taxon, giving specificity to the identification.
- Taxon-to-Spectrum Matches (TSMs) entities correspond to the number of spectra assigned to a given taxon and thus represent a proxy of the abundance of the identified taxa.
- TSMs Taxon-to-Spectrum Matches
- the method of the invention allows to identify and/or quantify the organisms present in X biological samples, wherein X is an integer superior or equal to two (2). It comprises the following successive steps: a) Providing one peptide sample PSi from each of said X biological samples, wherein i is an integer comprised between 1 and X, b) Separating said peptide samples PSi into at least N fractions Fj according to their hydrophobicity level, each fraction Fj being characterized by its hydrophobicity level Hj and/or another elution characteristic Ej, wherein N is an integer superior or equal to X, and wherein j is an integer comprised between 1 and N, c) Pooling at least one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej for each peptide sample PSi, so as to generate a final sample containing, for each PSi, only one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej, d)
- the method of the invention uses mass spectrometry (or MS).
- MS mass spectrometry
- This analytical technique is classically used in the identification of complex mixtures of chemical compounds and molecules of biological origin and in particular in proteomics.
- the basic principle of MS analysis is the measurement of the mass/electrical charge ratio of ionic species that have been created by ionization of the sample to be analyzed and to which electric and magnetic fields are applied.
- the MS technique implemented in the context of the present invention is more particularly tandem mass spectrometry or “MS/MS”.
- a particular stable ion from the ionic species generated in a first MS step (called “parent” or “precursor”) is selected and then decomposed or fragmented whereby ions from this decomposition or fragmentation (called “daughter” or “product”) are generated before being analyzed via a second MS step.
- parent or “precursor”
- daughter or “product”
- MS/MS provides an extremely interesting analytical tool especially for the analysis of complex samples, by allowing the generation of thousands of mass spectra or MS/MS spectra that are attributable to peptides, which in turn can be used to identify the proteins present in the sample using protein databases.
- the MS/MS analysis is implemented on peptides obtained from all the proteins initially contained in the sample to be characterized.
- peptide designates any kind of peptides that can be analysed by MS/MS, e.g., any peptide containing at least 6 amino acids and having a mass compatible with the MS/MS devices.
- samples that naturally contain such peptides for example, peptides found in cellular extracts prepared from an organism or peptides found in the extracellular medium after cultivation of the organism.
- the samples contain cells or organisms that should be degraded so as to produce the peptides that can be analyzed by mass spectrometry.
- the samples can be centrifugated or filtrated, resuspended in buffer, and lysed by chemical or physical processes, by conventional means.
- the peptide sample useful in the method of the invention can be obtained by any method, such as those disclosed in US 6,558,946, in WO 2012/083150, or in WO2014187983, which are incorporated herein by reference.
- chemical processes such as trypsinisation are herein preferred.
- each peptide sample contains roughly the same amount of peptides.
- the purpose of this first step a) is to ensure that, for each of the X biological samples to be analysed, at least one sample containing a sufficient amount - and preferably a similar amount - of peptides is provided or generated. This sufficient amount is for example comprised between 0.2ng and lOng of peptides per sample. Accordingly, at least X peptide samples will be analysed in the following steps of the method of the invention. These peptide samples will be herein referred to as “PSi”, for “Peptide Sample” number “i”, where i is an integer comprised between 1 and X.
- the label-free multiplexing concept herein developed relies on “prefractionation”, i.e., on a step where the peptide samples PSj from each biological sample to be characterized are first resolved by peptide separation, so as to provide peptide fractions of different hydrophobicity levels. More precisely, this prefractionation step involves that each peptide sample PSi is separated into at least N peptide fractions (N being superior or equal to X), the N fractions of the same PSi being referred to as “Fj”, where “j” is an integer comprised between 1 and N. Each and every fraction Fj is then characterized by its hydrophobicity level Hj and/or its elution characteristic Ej.
- This peptide separation can be performed by any means enabling to separate peptides according to their hydrophobicity level.
- the peptide separation in step b) is performed by hydrophobicity filtering, for example on hydrophobic membranes having particular characteristics so as to let particular peptides elute, depending on their hydrophobicity.
- the elution characteristic Ei which can characterize the fraction Fj will be the number of elution steps performed with increasing percentage of solvent with hydrophobic characteristics that have been performed to achieve its recovery.
- the peptide separation in step b) is performed by reversephase chromatography, i.e., with the same separation mode as the peptide separation applied during step d) of LC-MS/MS analysis.
- a HPLC system as proposed in the example part below.
- the inventors used a HPLC system ZORBAX StableBond C18 column (Agilent) with particle size 5 pm, average pore size 300 A, length 15 cm and internal diameter 4.6 mm, equipped with an automatic collector, allowing the collection of fractions in 96-well plates. Numerous samples could be fractionated with automated procedures, to produce and store the fractions for each peptide sample.
- steps b) and d) columns that do have distinct characteristics in terms of length, internal diameter, porosity, sorbents, and type of stationary phase.
- the skilled person well knows which columns can be used depending on the content of the sample to analyze.
- the elution characteristic Ei that can characterize the hydrophobicity of a fraction Fj to be considered is for example its elution time.
- This elution time can be any time (e.g., 10”, 20”, 30”, L, 2’, 5’, etc.), depending on the number of fractions to be performed (e.g. 1 min if 30 fractions have to be performed in 30 min).
- the peptide separation in step b) is performed by reversephase HPLC with an elution time of about 30 seconds.
- the peptide separation in step b) of the method of the invention is therefore performed by reverse phase spin columns, magnetic beads, Solid-Phase Extraction columns, or stage tips, preferably by reverse phase spin columns as disclosed in the example II below.
- final sample a particular mix, herein called “final sample”, is produced.
- This final sample will be analyzed by MS/MS in the next step.
- This final sample has to be carefully constituted, so that the method of the invention enables to identify the organisms present in each starting samples with accuracy.
- the final sample should contain fractions of all the initial samples. These fractions are obtained from step b) with distinct hydrophobicity characteristics so as to make it possible to distinguish them. More precisely, for the method of the invention to be performant and reliable, this final sample should contain only one fraction for each define hydrophobicity level. In other terms, all the fractions Fj added in the final sample should have a different hydrophobicity level.
- the final sample contains at least one peptide fraction of each biological sample.
- the final sample contains at least one peptide fraction of each biological sample.
- This step c) therefore consists in pooling several fractions Fj together, said fractions having been carefully chosen so as to avoid to combine two fractions (from different PSi or from the same PSi) having either the same elution characteristics and/or the same hydrophobicity levels.
- step c) consists in pooling at least one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej for each peptide sample PSi, so as to generate a final sample containing, for each PSi, only one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej.
- step c) consists in pooling at least one fraction Fj from each peptide sample PSi, each fraction Fj being characterized by its hydrophobicity level Hj and/or another elution characteristic Ej, so as to generate a final sample wherein all the fractions Fj added in the final sample have a different hydrophobicity level Hj and/or elution characteristic Ej.
- the final sample contains:
- any fraction Fj of PSi with any Fj fraction of PS2, with any fraction Fj of PS3, etc. provided that the hydrophobicity levels and/or elution characteristics of each fraction Fj is /are different. If all the fractions of PSihave been obtained by means of the same separating means, then it is foreseen that fraction Fi of PSi will have the same hydrophobicity levels and/or elution characteristic as the fractions Fi of the other PSi. That’s why, in this case, the final sample of step c) will contain only one fraction Fi from one PSi, only one fraction F2 from one PSi attire only one fraction F3 from one PSi, and more generally only Fj from one PSi, etc.
- each fraction of the same PSi has different hydrophobicity levels and/or elution characteristics.
- a duplicate of one of the isolates having different chromatographic elution times was included as 11 th or 21 rst fraction in the final mixes Mi l and M21. This is even recommended so as to ensure internal controls and therefore the accuracy of the method of the invention (the presence of the two distinct fractions of the same sample being confirmed by the corresponding mass spectra leading to the same organisms).
- the quantity of peptides in each eluted fraction of step b) is evaluated (for example, by optical density measurement at 205 nm, 220 nm or 280 nm, or colorimetric Lowry or Ninhydrin method, or by any other reliable means) or empirically established with preliminary assays so that the fractions mixed in the final sample contain the same equivalent quantities of peptides.
- the label -free multiplexing concept was validated in the present study using two mixtures of peptide fractions, Ml 1 and M21 - comprising 11 and 21 fractions, respectively. These mixtures were subjected to a 60-min gradient analysis on a Q-Exactive HF tandem mass spectrometer incorporating a high-field Orbitrap analyzer. Microorganisms could be identified at the species level from acquisition data recorded over just 1 min thanks to the high number of informative MS/MS spectra acquired with this system. The equivalent of 21 hydrophobicity -resolved microorganisms were identified in a single 60-min tandem mass spectrometry acquisition, resulting in a rate of identification of one microorganism every 3 min of mass spectrometry.
- next-generation instruments such as the Exploris 480 (Thermo) or TIMS-TOF Pro -Brucker) can deliver denser datasets for an equal run time, thus potentially allowing a larger number of samples to be multiplexed, or the final chromatography to be performed at higher speed.
- This step of the invention therefore involves any one of the following means to measure information on peptides that can help deciphering their sequence: liquid chromatography, mass spectrometry, liquid chromatography/mass spectrometry, high performance liquid chromatography, ultra-high performance liquid chromatography, Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry/mass spectrometry, Biological Aerosol Mass Spectrometry, ion mobility/mass spectrometry or ion mobility/mass spectrometry/mass spectrometry, or tandem mass spectrometry being performed in data dependent mode or data independent mode, or nanopore-based reading of peptide sequence.
- MALDI Matrix-assisted laser desorption/ionization
- Step d) of the method of the invention consists in injecting the final sample obtained in step c) into a chromatographic system coupled to a tandem mass spectrometer for resolving the peptides present in the final sample according to their hydrophobicity. Then, it is possible to establish the charge z of the peptide ion, its mass, and the mass spectra of the fragment of eluted peptides after their fragmentation.
- step d) is performed by using a HPLC system coupled to tandem mass spectrometer, preferably a nanoLC device coupled to a MS/MS tandem mass spectrometer.
- the number of short windows is preferably the integer X or any multiple of X (e.g., 1.5X, 2X, 3X, 5X, 10X, etc.).
- the inventors used a 60-min HPLC gradient coupled to a tandem MS/MS analysis and the resulting dataset was subdivided in 60 acquisition windows of 1 min.
- step d) is performed by subjecting the final sample of step c) to a 60-min HPLC gradient coupled to a tandem MS/MS analysis, wherein the resulting MS/MS spectra dataset has been divided into interpretable sub-datasets corresponding to short windows of elution times, for example of less than 2 minutes, typically of about 1 minute or of about 30 seconds.
- Step e) consists in analysing the MS/MS spectra obtained in step d) (in particular those contained in the smaller sub-datasets) by comparing them with a database of known peptide sequence data and matching each peptide sequence of said MS/MS spectrum to one or more organisms present in said database.
- Unipept- based proteotyping associates each peptide identified with a taxon via the UniProtKB database based on the lowest common ancestor approach [5], and is thus based on taxonspecific peptides.
- TCUP uses the attribution of the spectra to peptides / proteins and then to a taxon using a dedicated proteome database corresponding to a limited number of taxa potentially present in the sample [6].
- it is possible to use any of these taxonomic workflows comprising databases which have been fully annotated and attributed to specific taxa.
- the matching of the peptide sequences from the sample to taxa is performed by comparing each of these sequences to sequences present in public and/or private sequence databases in which sequences are annotated with taxon information, using techniques of sequence identity searches, and more generally of sequence similarity searches such as BLAST. This leads to a list of taxa (which may be any taxonomical grouping) and these are sorted in the order of the number of matching sequences they comprise, from the maximum value to the least. All these workflows are based on taxonspecific peptides.
- Taxon-to-Spectrum Matches TSM
- sipePEP taxon-specific peptide sequences
- taxon means a group of one (or more) populations of organism(s), which a taxonomist adjudges to be a unit.
- each of the spectra obtained in step d) is attributed to a peptide sequence encompassed into one or more database sequences and hence assigned to a taxon or a series of taxa, with the total number of matches per taxon being recorded.
- TSM spectrum to spectrum match
- taxon-specific peptide sequence(s) or “spePEP(s)” designates the peptide sequences that are specific to a taxon, as determined by peptide sequence sequence search on the whole database of theoretical proteomes. As described below, the number of spePEPs will be used to identify the organisms present in the initial biological samples.
- This last step consists in identifying the organisms present in each of the X biological samples, by attributing the organisms identified in step e) to each initial biological sample. This is done by using the specific elution characteristics at which each organism was identified.
- TSMs and spePEPs for the microbial species identified can be extracted from independent searches and normalized relative to the total number of bacterial TSMs and spePEPs per sample, respectively.
- This step preferably requires calculating, for each given acquisition window, an index which is herein called “Species Proteotyping Index” or “SPi” for each identified taxon and each fraction:
- TSMs 1 x TSMs + k x spePEP
- SPMs the number of TSMs in said acquisition window
- peerPEP the number of taxon-specific peptide sequences in said acquisition window
- k is a number comprised between 1.5 and 2.5.
- a mathematical modelisation of the SPi index along the retention time allows identifying the SPi maximum or the SPi maxima (in case of multiple identification of the same taxon in various fractions). Once it is calculated, it is possible to identify the elution characteristic of the taxon which corresponds to the elution value for which the maximum value of SPi for each organism is obtained. As shown in the example below, it is in particular possible to fit nonlinear regression curves of SPi and retention time, for example with a software such as GraphPad Prism in order to obtain the SPi maximum value.
- the method of the invention enables not only to identify but also to quantify the relative ratio of organisms present in each of the starting biological samples.
- TSM values of each of the organism identified at a given retention time can be directly compared.
- the absolute quantity of protein biomass of each organism identified at a given retention time can be measured by the specific tandem mass spectrometry signal assigned to each organism.
- specific quantities of isotopically labelled standards can, for example, be subjected to MS/MS analysis together with the peptides from the samples and a standard curve can be constructed from the ion signals generated from these standards. Using this standard curve, the relative abundance of a given ion can be converted to an absolute amount of the corresponding protein in the sample.
- Other alternatives relying on mass spectrometry of specific quantities of standards performed separately in exactly the same conditions can be carried out.
- the method of the invention allows both to generate MS/MS spectra characteristic of the peptides / proteins contained in the sample to be analyzed, and from these spectra to quantify the protein biomass contained in the sample to be analyzed and, via these characteristics, to quantify the relative proportion of organism(s) present in said sample to be analysed or the absolute amount of organism(s) present in said sample to be analysed in referene to standards. It also allows to quantify the amount of particular peptides, for example peptides that are characteristic of antibiotic resistance, virulence factors, or toxins.
- the measurement of peptide abundance can be performed by a method selected from the group comprising a method using spectral counts, extracted Ion Chromatograms, a quantification method based on mass spectrometry data or associated liquidchromatography data, the MS/MS total ion current and methods based on peptide fragments isolation and quantification such as selected reaction monitoring (SRM), multiple reaction monitoring (MRM) or parallel reaction monitoring (PRM).
- SRM selected reaction monitoring
- MRM multiple reaction monitoring
- PRM parallel reaction monitoring
- the method of the invention therefore also contains an additional step h’) of calculating the relative or absolute amount of the detected organisms in each biological sample.
- the quantities of proteins subject to proteolysis have been preferably normalised before step a) or if not, the quantity of peptides in each eluted fraction of step b) has been calculated so that the fractions mixed in the final sample contain the same equivalent quantities of peptides.
- the number X of biological samples that can be analyzed by using the present method is not limited. It is typically superior to two (2), so that multiplexing is effectively achieved. This number X can be extended to 50, 100, 150, 180, 200, 250, 300, 500, 1000 or even more, once the mass spectrometer devices allow longer runs to be performed or are able to analyze more data in shorter time.
- the innovative label-free multiplex strategy described here can thus be applied to any biological sample containing any kind of organisms, e.g., microbial isolates, or part of organisms, e.g. blood or saliva of an animal, leg or antenna of a parasite, or floor extracted from a cereal.
- This sample may even be of unknown nature such as a bioterrorism type sample, an extraterrestrial sample, or a new microbial isolate that has not been yet taxonomically characterized. Generally speaking, it is any sample for which it is desired to characterize by MS/MS the biological material of a protein nature that it contains or by which it is contaminated.
- this sample can be a biological fluid; a plant fluid such as sap, nectar and root exudate; a sample in a culture medium or in a biological culture reactor such as a cell culture of higher eukaryotes, yeasts, fungi, bacteria, viruses or algae; a liquid obtained from an animal or plant tissue; an animal or plant tissue; one or more cells; a cell pellet; a sample in a food matrix; a sample in a chemical reactor; a sample from a wastewater treatment plant; a sample from a composting plant; city water, river water, pond water, lake water, sea water, swimming pool water, water from cooling towers or water from underground sources; a sample from a liquid industrial effluent; wastewater from intensive livestock farming or chemical, pharmaceutical or cosmetic industries; a sample from an air filtration or a coating; a sample from an object such as a piece of fabric, a garment, a sole, a shoe, a tool, a weapon, etc.
- a biological culture reactor such as a cell culture of higher
- a sample of an object such as a fragment of fabric, a garment, a sole, a shoe, a tool, a weapon, etc. may be taken from the following objects: a pharmaceutical product; a cosmetic product; a perfume; a soil sample or a mixture thereof.
- the biological fluid is advantageously selected from the group consisting of blood such as whole blood or anti -coagulated whole blood, blood serum, blood plasma, lymph, saliva, spit, tears, sweat, semen, urine, stool, milk, cerebrospinal fluid, interstitial fluid, isolated bone marrow fluid, mucus or fluid from the respiratory, intestinal or genitourinary tract, cell extracts, tissue extracts and organ extracts.
- the biological fluid can be any fluid naturally secreted or excreted from a human or animal body or any fluid recovered, from a human or animal body, by any technique known to the skilled person such as extraction, sampling or washing. The steps of recovery and isolation of these different fluids from the human or animal body are performed prior to the implementation of the process according to the invention.
- the sample could be any derived products such as organoids, tumoroids, iPSC cells, or CAR-T cells.
- the sample used in the context of the present invention may contain, by its nature or possibly as a result of contamination, one or more cells (identical or different).
- “cell” is understood to mean both a cell of prokaryotic type and a cell of eukaryotic type.
- the cell may be a yeast such as a yeast of the genus Saccharomyces or Candida, a fungus, a parasite, a plant cell or an animal cell such as a mammalian or insect cell.
- the prokaryotic cells are bacteria which can be gram positive or gram negative, or archaea.
- bacteria belonging to the spirochetes and chlamydiae phyla the bacteria belonging to the families of the enterobacteria (such as Escherichia cold), the streptococcaceae (such as Streptococcus'), the microccaceae (such as Staphylococcus), the legionellae, the mycobacteria, the bacillaceae, the cyanobacteria and other.
- enterobacteria such as Escherichia cold
- streptococcaceae such as Streptococcus'
- microccaceae such as Staphylococcus
- the sample used in the present invention may include one (or more) pathogen(s).
- the sample prior to the implementation of the process according to the invention, the sample may be subjected to a pre-treatment to inactivate it.
- a pre-treatment to inactivate it.
- Any technique commonly used to inactivate a sample comprising pathogens can be used as part of this pre-treatment.
- this pre-treatment can consist of heating said sample at 99°C for a period of time between 20 min and 90 min and, in particular, of the order of 60 min (i.e. 60 min ⁇ 10 min) and then allowing it to cool to room temperature (i.e. 22°C ⁇ 3°C).
- this pre-treatment may consist of heating said sample to a temperature above 99°C, in particular above 100°C and in particular of the order of 110°C (i.e. 110°C ⁇ 5°C) for a period of between 5 min and 45 min and, in particular, of the order of 20 min.
- Inactivation of the pathogen can be also performed with other means, such as the use of chemical reagents, or a combination of methods, e.g. chemical reagents and temperature treatment.
- sample is understood to mean both a sample that has undergone this pre-treatment and a sample that is pathogen-free and, in fact, has not undergone such inactivation.
- the organism present in the sample to be analyzed can be e.g., a virus, a bacterium, an archaeon, a yeast, a fungus, an algae, a plant, an animal, a parasite, or any other organism, pathogenic or not, alive or not, taxonomically characterized or not.
- the sample analyzed by the methods of the invention contains only one kind of organism, e.g., one type of virus, one type of bacterium, etc.
- the sample analyzed by the methods of the invention contains more than one kind of organisms. It can e.g., contain several specie of one organism (e.g., several bacterial specie or several yeast strains), or a mixture of several organisms (e.g., bacteria and viruses, or yeast and fungi, or a virus and its animal host).
- several specie of one organism e.g., several bacterial specie or several yeast strains
- a mixture of several organisms e.g., bacteria and viruses, or yeast and fungi, or a virus and its animal host.
- the number X of samples to analyze is so high that the method cannot be performed on each sample separately in a convenient manner by the mass spectrometer currently available (e.g., because their possibility are limited by the duration of the runs), then it is possible to pool a number of samples together and then to apply the method of the invention to the pooled sample. For example, if there are 2000 samples to analyze, one can pool several samples together to reduce the number of samples to multiplex and analyze on the MS/MS run, and depending on the results, reanalyze the samples for which the pool gives an interesting result. This amount can be easily adjusted by the skilled person in view of the number of samples, the need to detect the organisms in each sample or in group samples, the performance of the mass spectrometer devices, etc.
- Figure 1 Workflow for multiplex proteotyping of bacterial isolates using Phylopeptidomics.
- FIG. 1 Distribution of TSMs attributed to P. putida (A), K. aerogenes (B) and R. pickettii (C) in 1-min acquisition windows.
- W window.
- FIG. 3 Number of TSMs attributed at different retention times for a single R. pickettii HPLC fraction. Number of TSMs pointing to the identification of R. pickettii in fraction 1 from R. pickettii for each 1-min acquisition period.
- Figure 4 Scatter plot of Species Proteotyping index (SPi) for R. pickettii and Sagitulla stellata. Proteotyping results of R pickettii for which two fractions contributed to Ml 1 is represented in squares. Proteotyping results of Sagittula stellata for which only one fraction contributed to Ml 1 is represented in dots. Lorentzian fitting are indicated. A and B represents the SPi maximum of the two fractions of R. pickettii in Mi l and C represent the SPi maximum of the fraction of Sagittula stellata.
- Figure 6 Scatter-plot of SPi values of each organisms in 20-sec acquisition time window. SPi values of each organisms composing the mix are represented as function of 20 sec acquisition time.
- the first line is associated to S. stellata, the second line to R. pomeroyi, the third line to O. indolifex, the fourth line to S. cerevisiae, the fifth line to K. aerogenes and the last line to M. tractuosa.
- Figure 8 Scatter plot of Species Proteotyping index of mixtures with different quantities.
- A) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to K. aerogenes with a peptide quantity of 280 ng, the following peaks are associated to AT. tractuosa (60ng), O. indolifex (100 ng), R. pomeroyi (340 ng), S. cerevisiae (180ng) and S. stellata (80ng).
- B) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to S. stellata with a peptide quantity of 140 ng, the following peaks are associated to R. pomeroyi (120 ng), M. tractuosa (40 ng), S. cerevisiae (160 ng), K. aerogenes (160 ng) and O. indolifex (100 ng).
- Figure 9 Scatter plot of Species Proteotyping index of mixtures containing a same organisms in several fractions. These scatter plot show the distribution of SPi values as function of 20-s window acquisition time with a same organism present twice in a mix. The peptide quantity of each fractions was normalized (100 ng). Same organisms are represented with the same color line.
- R. pomeroyi, K. aerogenes, S. cerevisiae, R. pomeroyi, K. aerogenes and S. cerevisiae are associated to each peak in the left to right order.
- B) The SPi values of each organisms are plotted and fit by Lorentzian model.
- S.stellata, R. pomeroyi, O. indolifex, K. aerogenes, O. indolifex and S. cerevisiae are associated to each peak in the left to right order.
- C) The SPi values of each organisms are plotted and fit by Lorentzian model.
- S. cerevisiae, R. pomeroyi, K. aerogenes, S.stellata, K. aerogenes, and R. pomeroyi are associated to each peak in the left to right order.
- D) The SPi values of each organisms are plotted and fit by Lorentzian model.
- R. pomeroyi, K. aerogenes, S. cerevisiae, S.stellata, S. cerevisiae and K. aerogenes are associated to each peak in the left to right order.
- Figure 10 Scatter plot of Species Proteotyping index of mixtures containing a same organisms in following fractions. These scatter plots show the distribution of SPi values as function of 20-s window acquisition time with a same organism present in consecutive fractions within a mix. Large wide peak is an indicator of the presence of a same organism in consecutive fractions. The peptide quantity of each fractions was normalized (100 ng). Same organisms are represented with the same color line. A) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to O. indolifex, the second line to S. cerevisiae, the third line to R pomeroyi and the last line to K. aerogenes.
- Samples were then transferred to 2-mL or 0.5-mL screw cap microtubes (Sarstedt) containing an equimolar mix of 0.1 mm silica beads, 0.1 mm glass beads, and 0.5 mm glass beads.
- Cells were disrupted using a Precellys Evolution bead-beater (Bertin Technologies) operated at 10 000 rpm for 10 x 30-s cycles, with 30-s pauses between each cycle. Beads and debris were removed by centrifugation at 16 000 g for 2 min. Supernatants were transferred into new microtubes and incubated at 99 °C for 5 min. Lysates were stored at -20 °C until use.
- Tryptic peptides were produced by re-suspending the paramagnetic beads in 10 pL of 50 mM NH4HCO3 supplemented with 0.01% Protease Max surfactant (Promega) and 1 pg.pL-1 trypsin gold (Promega). Samples were incubated on the plate for 15 min at 50 °C. After removal of the paramagnetic beads, the resulting 8-pL peptide solution was acidified by addition of trifluoroacetic acid (0.5% final concentration) and the total volume adjusted to 50 pL with aqueous 0.1% TFA. The concentration of tryptic digests was measured using a Pierce colorimetric Peptide Assay, as recommended (Thermo Scientific).
- Peptide digests were fractionated off-line by reverse-phase chromatography on an Agilent 1100 series HPLC system equipped with a UV detector. Peptides were separated on a ZORBAX StableBond Cl 8 column (Agilent) with particle size 5 pm, average pore size 300 A, length 15 cm and internal diameter 4.6 mm. A total of 25 pg of peptides was loaded onto the column, and a 30-min linear gradient was developed from 2.5% to 50% B at a flow rate of 400 pL.min' 1 . Mobile phases A and B were an aqueous solution of 0.1% formic acid (FA) in water, and 0.1% FA in acetonitrile, respectively. Fractions were collected every 30 s for 30 min after a 5-min delay corresponding to the dead volume. An Agilent 1100 G1364C fraction collector was used to collect fractions in 96-well plates. Microplates were stored at 4 °C until use.
- FA formic acid
- the Ml 1 sample was a mixture of eleven 1-min fractions eluted from a 30-min gradient, as presented in Table 1. The fractions included were: Fl; F3; F5; F7; F9; Fl l; F13; F15; F17; F19; F21.
- the M21 sample corresponded to a mixture of twenty-one 30-s fractions - Fl; F2; F3;... : F21 - eluted from a 30-min gradient, as presented in Table 1.
- Table 1 Composition of the concatenated fractions samples making up Mil and M21
- Desalted peptides were separated on a nanoscale PepMap 100 C18 nanoLC column (3 pm, 100 A, 75 pm i.d x 50 cm; Thermo Fisher Scientific) at a flow rate of 0.3 pL.min' 1 , applying a two-slope 60-min gradient as follows: 4-25% B from 0 to 50 min and 25-40% B from 50 to 60 min [10], Mobile phase A consisted of an aqueous solution of 0.1% (v/v) FA in water; phase B consisted of 0.1% FA in 80% acetonitrile. The mass spectrometer was operated in data-dependent mode (DDA) with an MS acquisition scan range of 350 to 1500 m/z.
- DDA data-dependent mode
- the 20 most abundant precursor ions were selected for fragmentation, applying a 10-s dynamic exclusion window and a 1.6-m/z isolation window. Only ions with a 2 + or 3 + charge were selected. HCD fragmentation was carried out with a normalized collision energy of 27 eV.
- ion peak lists were extracted using Mascot Daemon software, version 2.6.0 (Matrix Science), with the following settings: minimum mass (400), maximum mass (5000), grouping tolerance (0), intermediate scans (0), and threshold (1000).
- a Python script was applied to split the molecular ion peak list file into 1-min acquisition periods.
- MS/MS spectra for each of the resulting files were interpreted using Mascot version 2.6.1 (Matrix Science) against the NCBInrS database [18], an in-house assembled subset of the National Center for Biotechnology Information non-redundant (NCBInr) database (downloaded the 3rd of January 2018 as ftp://ftp.ncbi.nlm.nih.gov/blast/db/FASTA/nr.gz). Filtering was applied to the database to retain one representative taxon per species belonging to the Bacteria, Archaea, and Eukaryota superkingdoms. This database contains 50 080 649 protein sequences totaling 20054 069 172 amino acid residues.
- a second query was performed on a database comprising all known annotated genomes from NCBInr associated with the genera identified in the first-round search.
- a third query was performed on an even more reduced database comprising only the annotated genomes of the strains belonging to the species identified in the second query.
- TSMs Taxon- to-Spectrum Matches
- a filter combining the two parameters (lx TSMs + 2x spePEPs) for each 1-min window for each species was used to calculate the retention time (center value) and the r-squared value as an indication of confidence.
- This “Species Proteotyping index” (SPi) filter varied depending on the retention time.
- GraphPad Prism version 6.0 GraphPad software was used to fit nonlinear regression curves of SPi and retention time.
- the label-free multiplexing protocol described here is a new method developed to increase the throughput of tandem mass spectrometry proteotyping of microbial isolates.
- the concept involved the production of peptide mixtures from each individual microbial isolate. These peptide mixtures were then fractionated based on their hydrophobicity, and multiple fractions were subsequently concatenated to associate a specific reverse-phase chromatography retention time with each isolate. After merging the different peptide fractions, a single nanoLC-MS/MS analysis was performed to sequentially generate MS/MS spectra for each organism over a low- to high- hydrophobicity sequence.
- Bacterial isolates can be proteotyped from a short acquisition window, even with a low number of MS/MS spectra
- the tryptic peptides obtained after proteolysis of the soluble proteomes of the three species - Pseudomonas putida, Klebsiella aerogenes, and Ralstonia pickettii - were analyzed by shotgun tandem mass spectrometry on a Q-Exactive HF instrument following a 60-min gradient.
- the resulting datasets comprised 37 139, 35 759, and 27 339 MS/MS spectra, respectively.
- Query of the comprehensive NCBInr database resulted in 21 749, 20 701, and 12 073 PSMs, respectively.
- the ratio of MS/MS spectra interpreted was 59%, 58%, and 44%, respectively.
- the number of unique peptide sequences was 9 706, 9 137, and 7 401, respectively.
- the total number of TSMs extracted at the species level was 21 825 for P. putida, 20 681 for K. aerogenes, and 11 856 for R. pickettii.
- the general profile of the peptides eluted from the reverse-phase chromatography was quite similar for the three peptidomes.
- the P. putida peptidome dataset was split into 60 short windows corresponding to 1 min of acquisition time.
- the peptidome for the pure R. pickettii isolate was resolved into 60 x 30-s fractions by reverse-phase chromatography with a 30-min acetonitrile gradient. One of the fractions was selected, Fl - eluting at 12 min - for analysis by nanoLC-MS/MS with a 60-min acetonitrile gradient. The resulting dataset was subdivided into 60 small subdatasets corresponding to 1-min acquisition windows, and each of these sub-datasets was interpreted for proteotyping. For 11 contiguous fractions, the interpretation undoubtedly pointed to the presence of R. pickettii. These results confirm that a 1-min acquisition window with the Q-Exactive HF instrument provides enough information to proteotype the expected species.
- Figure 3 shows the distribution of the number of TSMs allowing the identification of R. pickettii plotted against the retention time.
- Lorentzian curve fitting was applied to each peak corresponding to the eight other fractions of Mi l to determine the respective proteotyping elution time: Sphingomonas yabuuchiae F3 at 16.3 min, Microbacterium oxydans F5 at 19.3 min, Stenotrophomonas maltophilia F9 at 26.1 min, Serratia marcescens Fl l at 29.7 min, Kineococcus radiotolerans F13 at 33.5 min, Pseudopedobacter saltans F17 at 41.5 min, Pseudomonas aeruginosa F19 at 45.3 min, and Methylobacterium extorquens F21 at 49.3 min.
- the multiplexing capacity of the method of the invention was next studied.
- the M21 chimeric sample was assembled by mixing 21 distinct 30-s fractions from 20 individual peptidomes prepared from isolates. Like for Mi l, two fractions from the R. pickettii isolate were included in this assemblage as a control.
- the peptidome assemblage was once again analyzed by shotgun tandem mass spectrometry with a 60-min gradient, resulting in a dataset comprising 28 883 recorded MS/MS spectra.
- Proteotyping analysis carried out as described for the Ml 1 mixture, revealed the presence of 20 distinct species in the dataset: Bacillus cereus, Bacillus thuringiensis, Deinococcus deserti, Deinococcus proteolyticus, K. radiotolerans, K.
- Table 2 shows the attribution of each organism to its corresponding fraction.
- the first column lists the experimental fraction used to create the mix.
- the 'Expected SPi max’ is a theoretical value calculated using the correlation curve.
- the Tow’ and ‘high’ fraction limits correspond to the range for each fraction.
- the fraction identified must fall within the range.
- the ‘Measured SPi maximum’ corresponds to the value obtained from the nonlinear regression curve, corresponding to the maximum of the SPi.
- the ‘fraction identified’ column corresponds to the theoretical calculation of the fraction number from the SPi maximum measured.
- the theoretical time is calculated as an approximation of the expected retention time. These time values were calculated using the second polynomial equation curve for M21. The low fraction and high fraction limits indicate the theoretical limits of the range for each fraction.
- the Lorentzian nonlinear regression of the SPi values for each microorganism was used to determine the retention time for each species identified. The 21 SPi maximum for the different species agreed perfectly with the chronological order, with values between the theoretical low and high limits.
- tandem mass spectrometry-based proteotyping has an important role to play in the massive identification of isolates, especially environmental microorganisms [11] through deeper characterization of microbiomes [19], or medically-relevant but poorly characterized pathogens [20]. Improving the throughput of this methodology requires robust sample preparation methods that can be applied to any sample [12], and miniaturization of the experimental load upstream of the mass spectrometry step [10], Additional improvements to throughput, such as multiplexing samples to reduce the per-sample costs of mass spectrometry are of considerable interest.
- multiplex isobaric labeling has been successfully used to increase throughput [21], however, the cost of the chemical reagents may hinder its universal adoption. No similar approach to multiplexing has yet been reported for proteotyping.
- the label-free multiplexing concept herein developed relies on prefractionation, where the peptide digest from each isolate to be characterized is first resolved by reverse-phase chromatography. This off-line peptide separation should ideally be performed with the same separation mode and in similar conditions to the on-line separation applied during the final nanoLC-MS/MS analysis.
- HPLC at normal flow was used with a reversephase column with distinct characteristics to the column used with the nanoflow LC- MS/MS system. Although there were notable differences between the two chromatography runs, the elution profile of the fractions was consistent across both systems.
- a HPLC system equipped with an automatic collector was used, allowing fraction collection in 96-well plates.
- the label -free multiplexing concept was validated in the present study using two mixtures of peptide fractions, Ml 1 and M21 - comprising 11 and 21 fractions, respectively. These mixtures were subjected to a 60-min gradient analysis on a Q-Exactive HF tandem mass spectrometer incorporating a high-field Orbitrap analyzer. Microorganisms could be identified at the species level from acquisition data recorded over just 1 min thanks to the high number of informative MS/MS spectra acquired with this system. The methodology could perform equally well whatever the tandem mass spectrometer used, provided the system generates similarly informative MS/MS spectra. Even higher performance can be expected with more recent tandem mass spectrometers that can deliver more MS/MS spectra per unit of time.
- next-generation instruments such as the Exploris 480 can deliver denser datasets for an equal run time, thus potentially allowing a larger number of samples to be multiplexed, or the final chromatography to be performed at higher speed.
- next-generation instruments such as the Exploris 480 can deliver denser datasets for an equal run time, thus potentially allowing a larger number of samples to be multiplexed, or the final chromatography to be performed at higher speed.
- the equivalent of 21 hydrophobicity -resolved microorganisms were identified in a single 60-min tandem mass spectrometry acquisition, resulting in a rate of identification of one microorganism every 3 min of mass spectrometry.
- the five bacteria used in this study were from a commercial source and cultivated as mentioned above.
- S. cerevisiae was from commercial baker’s yeast bought in supermarket that was solubilized as follows: 108 mg in 25 mL of PBS at pH 7.4 (lx) (GIBCO). Pellets cells were obtained after two centrifugations for the removing of residual liquid medium as described above. Resulting pellets were weighed and stored at -20 °C until use.
- Peptides were prepared as previously detailed. Briefly, a specific volume of lx lithium dodecyl sulfate sample loading buffer (Thermo Fisher Scientific) prepared without dyes and supplemented with 5% beta-mercaptoethanol (v/v) was added to each cell pellet (100 pL per 1.7 mg wet biomass). After their disruption, protein lysates were stored at -20 °C until use. Peptide digests were obtained by SP3-based proteolysis performed in a 96 wellplate format as described by Hayoun et al (2020). Lysate (20 pL) was mixed with 40 pg of Sera Mag beads (4pL), formic acid (12 pL) and CH3CN to a final concentration of 85%.
- Peptide digests were fractionated by Cl 8 micro spin columns (Harvard Apparatus) with particle size of 10 pm and average pore size of 300 A as recommended by the supplier.
- TFA trifluoroacetic acid
- a quantity of 40 pg of peptides was loaded onto the column and centrifugated during 1 min at 1000 x g.
- the eluate was load again onto the column and centrifugated.
- a step gradient from 1% to 32% CH3CN in 0.1% formic acid was applied by incrementing 1% at each step.
- Table 3 List of assemblages of five fractions and species for each of their five fractions (from the least hydrophobic to the most hydrophobic).
- Assembled peptide mixtures (M601-M632) were analyzed by nanoLC-MS/MS with an ultimate 3000 nanoLC system (Thermo Fisher Scientific) coupled to a Q-Exactive HF tandem mass spectrometer as described ([16]). Briefly, peptides were desalted on a reverse-phase PepMap 100 C18 p -precolumn. Then, they were separated on a nanoscale PepMap 100 Cl 8 nanoLC column with a flow rate of 0.3 pL.min' 1 following a two-slope 20-min gradient with 4-25% B from 0 to 17 min and 25-32% B from 17 to 20 min.
- Mobile phase A consisted of an aqueous solution of 0.1% (v/v) formic acid in water; phase B consisted of 0.1% formic acid in 100% CH3CN.
- the data-dependent mode (DDA) was conducted with an MS acquisition range of 350 to 1500 m/z. The 20 most abundant precursor ions were selected for fragmentation, applying a 10-s dynamic exclusion window and a 1 .6-/77 z isolation window and 8.3e5 for the intensity threshold.
- Tandem mass spectrometry proteotyping was performed as previously described. Briefly, ion peak lists were extracted using Mascot Daemon software, version 2.6.0 (Matrix Science) generating a MGF file per sample. Then, the mgf file was split into 20- sec acquisition time window by an in-house python script. MS/MS spectra for each of the resulting files were interpreted using Mascot version 2.6.1 (Matrix Science) against the NCBInrS database ([20]). A cascade search for the taxonomical analysis was applied as described above, for each 20 sec splitted files.
- TSMs Taxon-to-Spectrum Matches
- pePEP taxon-specific peptide sequences
- SPi Species Proteotyping index
- a specific fraction is selected, taking care that each fraction is different for the six isolates in order that a specific reverse phase chromatography retention time is associated with each isolate. Then, the six fractions of peptides are mixed together resulting in six pools of peptides with different hydrophobicity characteristics that will eluted sequentially from a reverse phase chromatography. The mixture of peptides is then analyzed by a single run of nanoLC-MS/MS. The interpretation of the MS/MS spectra to identify the organisms present along the gradient per windows of 20 sec is similar to our previous study. Briefly, the file containing the MS/MS spectra recorded along their retention time is sliced into 20 sec portions.
- SPi Species Proteotyping index
- SPi max the maximum of the peak
- FIG. 6 six well-defined elution peaks are clearly distinguished for an experimental mixture that has been analyzed with a nanoLC-MS/MS run operated with a 20 min gradient. In this analytical run six organisms have been identified at the species level: Sagittula slellala. Ruegeria pomeroyi.
- the six peptide fractions that were collected systematically for the six isolates were used to create a diversity of combinations of assemblages.
- a total of 23 different mixtures were created with a peptide quantity for each fraction normalized at 100 ng.
- Each of these 23 multiplexed samples were analyzed with a 20 min nanoLC-MS/MS gradient.
- the 6 expected species were systematically identified and their order along the chromatography was found matching perfectly with the expected order.
- the SPi-max intensities associated to each fraction was relatively constant for the five first fractions, but decreased for the more hydrophobic fraction.
- the integrated signal of the modeled SPi peak is in average for the 23 assemblages 2120 ⁇ 264 TSMs x sec for the first fraction, 2187 ⁇ 438 TSMs x sec for the second fraction, 2687 ⁇ 559 TSMs x sec for the third fraction, 2595 ⁇ 563 TSMs x sec for the fourth fraction, 2241 ⁇ 716 TSMs x sec for the fifth fraction, but only 1132 ⁇ 538 TSMs x sec for the sixth fraction.
- the five first fractions have an average SPi-max of 2366 ⁇ 257 TSMs x sec. This lower level of signal for the sixth fraction is not interfering in the correctness of the identification of the species, nor the correctness of the SPi-max measurement.
- the low taxonomical value of the most hydrophobic fraction is due to i) lower level of MS signals due to poor ionization of hydrophobic peptides, and ii) low number of assigned MS/MS spectra recorded because of poor fragmentation.
- Figure 7 shows a boxplot of the measured retention times for the six peaks and the 23 mixtures.
- the median SPi-max is at 554, 666, 780, 894, 1010, and 1159 sec of chromatography, respectively.
- the range of values are [536-570], [639-682], [764-803], [880-908], [993-1026], and [1143-1170], respectively.
- An average of all the retention time of the 23 mixtures obtained for specific fraction was made to have a reference value.
- the difference between the minimum and the maximum values are used to determine the range of time of retention time for each fraction with the RT low range and the RT high range.
- the multiplex proteotyping method can manage different quantities of peptides
- Mix M624 comprises six peptide fractions: 280 ng of peptides from K. aerogenes (first fraction), 60 ng of M. tractuosa (second fraction), 100 ng of 0. indolifex (third fraction), 340 ng of R. pomeroyi (fourth fraction), 180 ng of S. cerevisiae (fifth fraction), and 80 ng of S. stellata (sixth fraction).
- Figure 8 Panel A
- Six organisms could be detected and their corresponding SPi values were plotted against the elution time. Six peaks were clearly distinguished.
- the SPi- max intensity associated to K. aerogenes (280 ng of peptides) is the double of that associated to M. tractuosa (60 ng).
- the signal of the first species is 3.16 times higher than that of the second, a ratio corresponding exactly to the one expected from the injected quantities. The same was observed for R. pomeroyi with 3.17 times higher than the SPi value of other organisms excepted for K. aerogenes.
- Mix M625 is the assemblage of six peptide fractions: 140 ng of peptides from S.
- FIG. 8 Panel B shows the results of its analysis with the same previous condition. Once again, the six organisms were correctly identified and their elution peaks were clearly distinguished. The order of their respective retention times is in perfect agreement with the composition of the mixture and selected fractions. As indicated in the figure, here the SPi signal is relatively more similar between the peaks but a slight drop of SPi values were noted for M. tractuosa (third fraction) and O. indolifex (sixth fraction) as expected as their quantities are lower than the others with 40 and 100 ng, respectively.
- the retention times associated to each fraction for the two mixtures are 562 ⁇ 7 sec; 673 ⁇ 7 sec; 777 ⁇ 0 sec; 884 ⁇ 4 sec; 1016 ⁇ 5 sec; and 1163 ⁇ 1 sec for the least to the most hydrophobic fractions. These values are perfectly matching the elution windows defined for the mixtures made with fractions with equal amounts of peptides which are indicated in Table 4. Thus, it was observed in these two examples, that the difference of quantities of peptides per fraction does not affect the modeled retention time of the identified organism associated to each fraction.
- the method is robust even if a same organism is present in several fractions in the mix
- M629 comprises only 3 organisms: R. pomeroyi, S. cerevisiae and K. aerogenes. Each of the three organisms contributed to two fractions. As shown in Figure 9 (Panel A), two SPi peaks are clearly distinguished for each of these three organisms. The peaks at 556 and 885 sec corresponds to R. pomeroyi.
- the peaks indicate the presence of S. cerevisiae and the peaks at 678 and 998 sec are from ", aerogenes.
- the last peak (F6) associated to S. cerevisiae shows a decreasing of SPi intensity.
- SPi intensity for the last peak (F6) is at 119 and the average of SPi intensity for last fraction is 202 ⁇ 69.
- M630 contains peptides corresponding to four organisms: K. aerogenes, O. indolifex, R. pomeroyi and S. stellata ( Figure 9, Panel B). Two organisms have two distinct peaks: O. indolifex at 790 and 1017 sec, and R.
- Mix M632 corresponds to peptides from four organisms: R pomeroyi, K. aerogenes, S. cerevisiae, and S. stellata.
- Figure 9 shows that two peaks can be distinguished for two organisms: S. cerevisiae at 786 and 998 sec and K. aerogenes at 679 and 1159 sec. For the remaining two organisms, a single peak was observed at 557 sec for R. pomeroyi and 897 sec for S. stellata.
- the identified organisms and the order of the retention time of the corresponding fractions are in perfect agreement with the composition of this mix.
- the average retention time for each fraction is: 551 ⁇ 6 sec, 668 ⁇ 12 sec, 789 ⁇ 4 sec, 896 ⁇ 7 sec, 1004 ⁇ 9 sec and 1160 ⁇ 4 sec for the least to the most hydrophobic fractions.
- FIG. 10 panel A shows the SPi values of each selected organisms plotted as a function of time, with a point every 20 sec acquisition time window. Four peaks are observed, the first peak being wider than the other peaks. Indeed, its width is 156 sec while the average of the other peaks is 78 ⁇ 8 sec in average. The same observation can be done for the mix M627 represented on Figure 10 (panel B), where three organisms were identified. A single peak was associated to K.
- a selected peptide fraction from each isolate is differentiated from the fractions arising from the other isolates by its hydrophobic characteristics. Because HPLC requires equipment and specific expertise, it is proposed here to fraction the peptidomes with Cl 8 spin columns that would require only a centrifuge. In principle, this approach can be fully automatized and is cost-effective. Here, six fractions differing in hydrophobicity are used for creating a mixture representative of six isolates and analysed into a single 20 min nanoLC-MS/MS run. It was noted that each of the resulting peaks of elution of species- specific peptides are clearly identified. Without any optimisation and automation, the fractionation step took about 1 h for treating in parallel six isolates. The current yield of identification is around 3 min of tandem mass spectrometry per isolate. This results in a rather competitive methodology in terms of costs.
- the best practice to establish a novel analytical method is to test its robustness whatever the conditions.
- the robustness of the methodology was tested by using 23 sets of six isolates with the same organisms but with different combinations. Fractions used for the artificial combinations were normalized with a peptide quantity of 100 ng per added fraction, resulting into mixtures of 600 ng of peptides injected per nanoLC-MS/MS.
- the peptides of each of the six microorganisms present in a mixture can delineate a specific taxon-retention time on the final reverse-phase chromatography. The dispersion of each of these six retention times observed for this set of 23 mixtures is minimal and shows a very good reproducibility.
- the SPi-max value for this organism corresponds exactly to the fraction in the middle, but the width of the peak is enlarged.
- the number of organisms is inferior to the number of introduced fractions, i.e. six here, a simple ratio of the width of peaks and the average of the expected width can confirm the presence of the same organism in neighboring fractions.
- these three parameters should be taken into account to determine the number of fractions containing the same organism.
- two peaks of elution for the organism can be easily distinguished with width in full agreement with the average value. In such case, two SPi-max are obtained.
- the method makes no use of pre-recorded mass spectra databases such as with whole-cell MALDI-TOF mass spectrometry ([1]), but is rather based on the whole known diversity through the current database of annotated full genome sequences. It is thus applicable to any of the organisms already genome sequenced and their relatives, as well as uncharacterized branches of the tree of life because partial conserved information can be always retrieved. Regarding the range of application of the methodology, it was tested whether bacteria or yeasts can be easily identified. Specifically, S. cerevisiae was introduced in several scenarios to prove that the method applies equally well on Bacteria and Eukaryota. S. cerevisiae was always identified. The extraction of peptides from S.
- cerevisiae or from the bacteria is identical as based on an optimized protocol previously developed for complex samples including both types of material ([12]). However, it was observed less signal in terms of abundance (TSMs) and taxon-specificity (SpePEPs) compared to the signals for bacteria while the quantity of peptides was identical. This comes from the dynamic range of the proteomes that differ between Eukaryota and Bacteria on the one hand, and the number of genomes for this specific branch of the tree of life that may lower the Saccharomyces cere ⁇ 'isiae- c ⁇ ⁇ c peptides on the other hand.
- the methodology presented here enables robust and cost-effective multiplexing of samples for microbial identification by tandem mass spectrometry proteotyping.
- the results show how simple the methodology is because the fractionation step before the analysis can be performed with Cl 8 spin columns. Based on this affordable technology, the fractionation of peptidomes can be fully automatized to reduce operations and time to result.
- the current protocol allows the identification of 6 organisms in a single 20-min nanoLC-MS/MS analytical run, but there is room for improvement as the latest generation of tandem mass spectrometers have several fold more capacities than the mass spectrometer used here.
- CMMB Carboxylate-Modified Magnetic Bead
- CIF Isopropanol Gradient Peptide Fractionation
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Chemical & Material Sciences (AREA)
- Molecular Biology (AREA)
- Physics & Mathematics (AREA)
- Organic Chemistry (AREA)
- Immunology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biophysics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biochemistry (AREA)
- Hematology (AREA)
- Urology & Nephrology (AREA)
- General Health & Medical Sciences (AREA)
- Biomedical Technology (AREA)
- Analytical Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Medicinal Chemistry (AREA)
- Microbiology (AREA)
- Biotechnology (AREA)
- General Physics & Mathematics (AREA)
- Wood Science & Technology (AREA)
- Cell Biology (AREA)
- Genetics & Genomics (AREA)
- Food Science & Technology (AREA)
- Zoology (AREA)
- Pathology (AREA)
- General Engineering & Computer Science (AREA)
- Toxicology (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
Abstract
The present invention provides an innovative label-free multiplexing method for the identification of microorganisms using tandem mass spectrometry. This new method is based on off-line reverse-phase fractionation of individual peptidomes, on the subsequent mixing of fractions having staged hydrophobicity and on phylopeptidomic proteotyping resolution. The method of the invention enables to analyze the composition of many samples in only 60-minutes, by using a single-run of mass spectrometry coupled to reverse phase chromatography, without using any labeling reagents. This approach opens new perspectives for the application of high-throughput proteotyping of bacteria using tandem MS such as in large culturomics projects and high-throughput microbial diagnostics.
Description
Label-free multiplex proteotyping of microbial isolates
DESCRIPTION OF THE PRIOR ART
To explore microbial diversity, understand the environment, and screen new catalysts for biotechnological purposes, it is crucial to be able to accurately and rapidly identify the microorganisms present in a sample. Moreover, the ability to identify the presence of known or emergent strains of infectious agents such as bacterial, viral or other diseases causing organisms in a sample is important for the purposes of public health, epidemiology and public safety. Likewise, being able to determine that a given product, such as a processed food or cosmetic preparation comprises only the claimed biological constituents is also a growing concern due to the increasing prevalence of highly processed food stuffs in the food chain and the reliance of consumers on manufacturers and retailers that their products can be trusted.
The routine and reliable identification of the organisms present in numerous samples at once is now possible by means of several multiplex existing methods. To achieve multiplex analyzing, it is for example possible to rely upon the provision of reagents (PCR primers or antibodies) which are specific for a given organism (or at best for a related group of organisms). The problem with these multiplex methods is that the identification of the organisms is based upon the use of a specific detection reagent (or set of reagents) per target organism, what is very expensive and increases the number of manipulations required at the sample preparation stage. Moreover, there are technical limitations due to the number of materials comprised/used in such methods, and it is often not possible to provide an entirely comprehensive assay when the sample contains several organisms.
More recent methods have been developed based on the protein content of the samples that can be analyzed by mass spectrometry. Whole-cell matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) is well-established as a means to achieve this type of identification based on the molecular weights of basic, small polypeptide fingerprints [1], This rapid and inexpensive proteotyping method is ideally suited for known clinical pathogens, but its performance declines when organisms are not closely related to those listed in the experimentally-established profiles stored in the database. To overcome this limitation, tandem mass spectrometry-based proteotyping has been found to be a valuable complementary methodology to quickly classify atypical organisms, which has been shown to be applicable even with mixtures of microorganisms [2] [3], This approach is well adapted to culturomics [4], where large numbers of microorganisms are isolated by multiplying growth conditions and varying culture media.
As tandem mass spectrometry-based proteotyping is based on low molecular-weight peptides rather than whole proteins, this methodology is much more sensitive than wholecell MALDI-TOF MS. Furthermore, it provides very precise taxonomical identification through the identification of thousands of peptide sequences. However, it can present a certain number of difficulties depending on the sample’s type and composition, especially when several samples have to be analyzed. To gain throughput of tandem mass spectrometry coupled to reverse phase chromatography, multiplexing could be in principle achieved by labeling the proteins from the organisms present in the different samples with distinguishable isobaric mass tags. This method would be in principle very expensive and requires a number of manipulations at the sample preparation stage.
As a matter of fact, the cost and speed of tandem mass spectrometry is not optimized for successful applications involving a great number of samples containing numerous
different organisms. To make the methodology accessible to a wide range of users, the next expected step would be to improve the throughput of MS/MS proteotyping by optimizing the costs and the duration of the mass spectrometry measurement per sample. For example, a method enabling to analyze multiple samples in a single nanoLC-MS/MS run without the need of isobaric mass tag reagent would be ideal, as it would save instrument time, therefore spare time and money.
The present invention solves this need, by providing a multiplexing method involving the analysis of several samples in a single tandem mass spectrometry run, resulting in the analysis of at least 21 samples done in about 60 minutes, without any need for the use of labeling reagents. The examples below prove the performance of the approach for the simultaneous analysis of 21 bacterial isolates with a single 60-min gradient nanoLC- MS/MS run. This approach significantly decreases the per-isolate cost of identification, and is highly sensitive and reproducible.
DESCRIPTION OF THE INVENTION
The present invention concerns an innovative label-free multiplexing method for the identification of microorganisms using tandem mass spectrometry, based on off-line reverse-phase fractionation of individual peptidomes. Multiplexing is achieved by mixing fractions of staged hydrophobicity, so that each sample can be mapped to specific elution times. In the proof-of-concept study described in the example below, up to 21 different samples have been analyzed by reverse-phase nanoLC-MS/MS in a single run, and the 21 different microorganisms present in these 21 different samples have been appropriately identified by phylopeptidomic proteotyping. Advantageously, these 21
microorganisms were identified in a single 60-min analytical run of a tandem mass spectrometer coupled to reverse phase chromatography, resulting in a rate of one microorganism identified per 3 min of mass spectrometry time, without any need of labeling reagents. Importantly, its capacity to simultaneously analyze 21 isolates within a single 60-min gradient nanoLC-MS/MS run resulted in significant time and cost savings compared to 21 separate nanoLC-MS/MS analyses.
This method relies on separation of peptidomes by their hydrophobic characteristics, mixture of different fractions of the various isolates, and phylopeptidomics-based proteotyping. This last step relies on the assignation of taxonomical information for all the identified peptides ([3]). Taxon-specific peptides are sequences that are uniquely found in a given taxon, giving specificity to the identification. Taxon-to-Spectrum Matches (TSMs) entities correspond to the number of spectra assigned to a given taxon and thus represent a proxy of the abundance of the identified taxa. The methodology takes into account all the assigned peptides shared either by closely- or distantly-related organisms present in the generic database used to interpret the tandem mass spectrometry spectra at all possible taxonomic ranks.
More generally, the method of the invention allows to identify and/or quantify the organisms present in X biological samples, wherein X is an integer superior or equal to two (2). It comprises the following successive steps: a) Providing one peptide sample PSi from each of said X biological samples, wherein i is an integer comprised between 1 and X, b) Separating said peptide samples PSi into at least N fractions Fj according to their hydrophobicity level, each fraction Fj being characterized by its hydrophobicity
level Hj and/or another elution characteristic Ej, wherein N is an integer superior or equal to X, and wherein j is an integer comprised between 1 and N, c) Pooling at least one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej for each peptide sample PSi, so as to generate a final sample containing, for each PSi, only one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej, d) Injecting the final sample of step c) into a chromatographic system coupled to a tandem mass spectrometer for resolving the peptides by their hydrophobicity, establishing preferably the charge z and the mass of the eluted peptides, and the mass spectra of the eluted peptides after their fragmentation, e) Analysing the MS/MS spectra obtained in step d) by comparing them with a database of known peptide sequence data and matching each peptide sequence of said MS/MS spectrum to one or more organisms present in said database, preferably generating Taxon to Spectrum Matches (TSM) and taxon-specific peptide sequences (spePEP), f) Identification of the organisms present in each of the X biological samples, by classifying the organisms identified in step e) to each initial biological sample, by using the specific elution characteristics at which the organisms were identified.
Each of these steps will now be described in more details.
Step a): providing peptide samples
The method of the invention uses mass spectrometry (or MS). This analytical technique is classically used in the identification of complex mixtures of chemical compounds and molecules of biological origin and in particular in proteomics. The basic principle of MS analysis is the measurement of the mass/electrical charge ratio of ionic species that have
been created by ionization of the sample to be analyzed and to which electric and magnetic fields are applied. The MS technique implemented in the context of the present invention is more particularly tandem mass spectrometry or “MS/MS”. In this technique, a particular stable ion from the ionic species generated in a first MS step (called "parent" or "precursor") is selected and then decomposed or fragmented whereby ions from this decomposition or fragmentation (called "daughter" or "product") are generated before being analyzed via a second MS step. By fine selection of the precursor ions (restricted mass/electrical charge range), only certain ions produced by certain analytes can be introduced into the fragmentation chamber where the collision with atoms of an inert gas produces daughter ions. Since both precursor and daughter ions are reproducibly produced for given ionization/fragmentation conditions, MS/MS provides an extremely interesting analytical tool especially for the analysis of complex samples, by allowing the generation of thousands of mass spectra or MS/MS spectra that are attributable to peptides, which in turn can be used to identify the proteins present in the sample using protein databases.
In the method of the present invention, the MS/MS analysis is implemented on peptides obtained from all the proteins initially contained in the sample to be characterized.
As used herein, the term “peptide” designates any kind of peptides that can be analysed by MS/MS, e.g., any peptide containing at least 6 amino acids and having a mass compatible with the MS/MS devices.
In a particular embodiment, it is possible to use samples that naturally contain such peptides (for example, peptides found in cellular extracts prepared from an organism or peptides found in the extracellular medium after cultivation of the organism).
In another embodiment, the samples contain cells or organisms that should be degraded so as to produce the peptides that can be analyzed by mass spectrometry. In this case, the samples can be centrifugated or filtrated, resuspended in buffer, and lysed by chemical or physical processes, by conventional means. The peptide sample useful in the method of the invention can be obtained by any method, such as those disclosed in US 6,558,946, in WO 2012/083150, or in WO2014187983, which are incorporated herein by reference. For the proteolytic reaction, chemical processes such as trypsinisation are herein preferred.
For facilitating the quantification process, it is recommended to normalize the quantities of proteins subject to proteolysis before step a) so that each peptide sample contains roughly the same amount of peptides. In this aim, it is also possible to normalize the cell biomass (weight and volume of buffer according to the weight).
The purpose of this first step a) is to ensure that, for each of the X biological samples to be analysed, at least one sample containing a sufficient amount - and preferably a similar amount - of peptides is provided or generated. This sufficient amount is for example comprised between 0.2ng and lOng of peptides per sample. Accordingly, at least X peptide samples will be analysed in the following steps of the method of the invention. These peptide samples will be herein referred to as “PSi”, for “Peptide Sample” number “i”, where i is an integer comprised between 1 and X.
In the method of the invention, it is sufficient to use only one peptide sample for each biological sample to be analysed. Several peptide samples of the same biological sample can be nevertheless used as control experiment in order to have several confirmations of the identification.
Step b): separation of the peptide samples
The label-free multiplexing concept herein developed relies on “prefractionation”, i.e., on a step where the peptide samples PSj from each biological sample to be characterized are first resolved by peptide separation, so as to provide peptide fractions of different hydrophobicity levels. More precisely, this prefractionation step involves that each peptide sample PSi is separated into at least N peptide fractions (N being superior or equal to X), the N fractions of the same PSi being referred to as “Fj”, where “j” is an integer comprised between 1 and N. Each and every fraction Fj is then characterized by its hydrophobicity level Hj and/or its elution characteristic Ej.
This peptide separation can be performed by any means enabling to separate peptides according to their hydrophobicity level.
In a first embodiment, the peptide separation in step b) is performed by hydrophobicity filtering, for example on hydrophobic membranes having particular characteristics so as to let particular peptides elute, depending on their hydrophobicity. In this case, the elution characteristic Ei which can characterize the fraction Fj will be the number of elution steps performed with increasing percentage of solvent with hydrophobic characteristics that have been performed to achieve its recovery.
In the example I below, is reported the use of reverse phase separation by HPLC with a Cl 8 reverse phase chromatographic column for resolving the peptidomes into fractions prior to their mass spectrometry analysis.
In a preferred embodiment, the peptide separation in step b) is performed by reversephase chromatography, i.e., with the same separation mode as the peptide separation applied during step d) of LC-MS/MS analysis.
It is in particular advantageous to use a HPLC system, as proposed in the example part below. In this study, the inventors used a HPLC system ZORBAX StableBond C18 column (Agilent) with particle size 5 pm, average pore size 300 A, length 15 cm and internal diameter 4.6 mm, equipped with an automatic collector, allowing the collection of fractions in 96-well plates. Numerous samples could be fractionated with automated procedures, to produce and store the fractions for each peptide sample. Because reversephase chromatography with HPLC is a highly robust and reproducible technique, this sample preparation step resulted in a robust procedure. Indeed, all of the microbial digests analyzed by the inventors presented similar chromatographic profiles during prefractionation, yet each peptide fraction was clearly distinguishable along the acetonitrile gradient during the second chromatography run of step d). Despite the differences in chromatographic conditions (HPLC at normal flow was used in step b) with a reverse-phase column with distinct characteristics as compared with the column used in the nanofl ow LC-MS/MS system of step d)), overlap between adjacent fractions was quite low during the second chromatography run of step d), and the elution profile of the fractions was consistent across both systems.
Thus, as shown in the example below, it is possible to use in steps b) and d) columns that do have distinct characteristics in terms of length, internal diameter, porosity, sorbents, and type of stationary phase. The skilled person well knows which columns can be used depending on the content of the sample to analyze.
When the peptide separation of step b) is performed by HPLC, the elution characteristic Ei that can characterize the hydrophobicity of a fraction Fj to be considered is for example its elution time. This elution time can be any time (e.g., 10”, 20”, 30”, L, 2’, 5’, etc.),
depending on the number of fractions to be performed (e.g. 1 min if 30 fractions have to be performed in 30 min).
In a preferred embodiment, the peptide separation in step b) is performed by reversephase HPLC with an elution time of about 30 seconds.
However, different means allowing similar separation can be used to generate fractions [27]). Magnetic beads ([28]), Solid-Phase Extraction columns ([29]), stage tips ([30]) and spin columns ([31]) have been shown to allow separation of peptides into fractions with defined characteristics without the need for a costly instrument. These approaches thus represent interesting alternatives to make the label free, multiplexing proteotyping approach of the invention more cost effective and user friendly.
As shown in the example II below, it is possible to multiplex samples for tandem mass spectrometry proteotyping by using reverse phase spin columns. It is herein described the multiplexing of six organisms that can be analyzed within 20 min of tandem mass spectrometry after such separation.
In a preferred embodiment, the peptide separation in step b) of the method of the invention is therefore performed by reverse phase spin columns, magnetic beads, Solid-Phase Extraction columns, or stage tips, preferably by reverse phase spin columns as disclosed in the example II below.
Step c): obtention of a final sample to be analysed by MS/MS.
In a next step, the method of the invention requires that a particular mix, herein called “final sample”, is produced. This final sample will be analyzed by MS/MS in the next
step. This final sample has to be carefully constituted, so that the method of the invention enables to identify the organisms present in each starting samples with accuracy.
The final sample should contain fractions of all the initial samples. These fractions are obtained from step b) with distinct hydrophobicity characteristics so as to make it possible to distinguish them. More precisely, for the method of the invention to be performant and reliable, this final sample should contain only one fraction for each define hydrophobicity level. In other terms, all the fractions Fj added in the final sample should have a different hydrophobicity level.
It is also necessary that the final sample contains at least one peptide fraction of each biological sample. However, provided that the above-mentioned requirement of different hydrophobicity levels is complied with, then it is possible to add in the final sample several peptide fractions of the same biological sample in order to have several independent identifications of the same isolates with the objective of increasing the confidence in these identifications.
This step c) therefore consists in pooling several fractions Fj together, said fractions having been carefully chosen so as to avoid to combine two fractions (from different PSi or from the same PSi) having either the same elution characteristics and/or the same hydrophobicity levels.
In a preferred embodiment of the method of the invention, more than one fraction Fj of the same peptide sample PSi is added in the final sample of step c), provided that said fractions Fj have different hydrophobicity levels Hj and/or different elution characteristics Ej.
Accordingly, step c) consists in pooling at least one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej for each peptide sample PSi, so as to generate a final sample containing, for each PSi, only one fraction Fj having a characteristic hydrophobicity level Hj and/or elution characteristic Ej.
In other terms, step c) consists in pooling at least one fraction Fj from each peptide sample PSi, each fraction Fj being characterized by its hydrophobicity level Hj and/or another elution characteristic Ej, so as to generate a final sample wherein all the fractions Fj added in the final sample have a different hydrophobicity level Hj and/or elution characteristic Ej.
In a particular embodiment, after this pooling step c), the final sample contains:
• the fraction Fi of the peptide sample PSi having an hydrophobicity level Hi and an elution characteristic Ei,
• the fraction F2 of the peptide sample PS2 having an hydrophobicity level H2 and an elution characteristic E2,
• the fraction F3 of the peptide sample PS 3 having an hydrophobicity level H3 and an elution characteristic E3,
• (etc...)
• the fraction Fj of the peptide sample PSi having an hydrophobicity level Hj and an elution characteristic Ej, wherein Hi H2 H3 Hj and wherein Ei E2 E3 ^ ... ^ Ej.
It would be also possible to pool any fraction Fj of PSi with any Fj fraction of PS2, with any fraction Fj of PS3, etc. provided that the hydrophobicity levels and/or elution characteristics of each fraction Fj is /are different.
If all the fractions of PSihave been obtained by means of the same separating means, then it is foreseen that fraction Fi of PSi will have the same hydrophobicity levels and/or elution characteristic as the fractions Fi of the other PSi. That’s why, in this case, the final sample of step c) will contain only one fraction Fi from one PSi, only one fraction F2 from one PSi„ only one fraction F3 from one PSi, and more generally only Fj from one PSi, etc.
It is also possible to include in the final sample of step c) several fractions of the same peptide sample PSj, provided that each fraction of the same PSi has different hydrophobicity levels and/or elution characteristics. For example, in the examples below, a duplicate of one of the isolates having different chromatographic elution times was included as 11th or 21rst fraction in the final mixes Mi l and M21. This is even recommended so as to ensure internal controls and therefore the accuracy of the method of the invention (the presence of the two distinct fractions of the same sample being confirmed by the corresponding mass spectra leading to the same organisms).
In another preferred embodiment, the quantity of peptides in each eluted fraction of step b) is evaluated (for example, by optical density measurement at 205 nm, 220 nm or 280 nm, or colorimetric Lowry or Ninhydrin method, or by any other reliable means) or empirically established with preliminary assays so that the fractions mixed in the final sample contain the same equivalent quantities of peptides.
Step d): analysis by MS/MS
The label -free multiplexing concept was validated in the present study using two mixtures of peptide fractions, Ml 1 and M21 - comprising 11 and 21 fractions, respectively. These mixtures were subjected to a 60-min gradient analysis on a Q-Exactive HF tandem mass spectrometer incorporating a high-field Orbitrap analyzer. Microorganisms could be
identified at the species level from acquisition data recorded over just 1 min thanks to the high number of informative MS/MS spectra acquired with this system. The equivalent of 21 hydrophobicity -resolved microorganisms were identified in a single 60-min tandem mass spectrometry acquisition, resulting in a rate of identification of one microorganism every 3 min of mass spectrometry.
This methodology could perform equally well whatever the tandem mass spectrometer used, provided the system generates similarly informative MS/MS spectra. Even higher performance can be expected with more recent tandem mass spectrometers that can deliver more MS/MS spectra per unit of time. For example, next-generation instruments such as the Exploris 480 (Thermo) or TIMS-TOF Pro -Brucker) can deliver denser datasets for an equal run time, thus potentially allowing a larger number of samples to be multiplexed, or the final chromatography to be performed at higher speed.
This step of the invention therefore involves any one of the following means to measure information on peptides that can help deciphering their sequence: liquid chromatography, mass spectrometry, liquid chromatography/mass spectrometry, high performance liquid chromatography, ultra-high performance liquid chromatography, Matrix-assisted laser desorption/ionization (MALDI) mass spectrometry/mass spectrometry, Biological Aerosol Mass Spectrometry, ion mobility/mass spectrometry or ion mobility/mass spectrometry/mass spectrometry, or tandem mass spectrometry being performed in data dependent mode or data independent mode, or nanopore-based reading of peptide sequence.
Step d) of the method of the invention consists in injecting the final sample obtained in step c) into a chromatographic system coupled to a tandem mass spectrometer for
resolving the peptides present in the final sample according to their hydrophobicity. Then, it is possible to establish the charge z of the peptide ion, its mass, and the mass spectra of the fragment of eluted peptides after their fragmentation.
In a preferred embodiment, step d) is performed by using a HPLC system coupled to tandem mass spectrometer, preferably a nanoLC device coupled to a MS/MS tandem mass spectrometer.
As proposed by the inventors in the examples below, it is recommended to divide the resulting MS/MS spectra dataset into interpretable sub-datasets corresponding to short windows of elution times, to focus the interpretation in terms of taxonomical identification per unit of time and be more accurate and selective. The number of short windows is preferably the integer X or any multiple of X (e.g., 1.5X, 2X, 3X, 5X, 10X, etc.). For example, the inventors used a 60-min HPLC gradient coupled to a tandem MS/MS analysis and the resulting dataset was subdivided in 60 acquisition windows of 1 min.
Thus, in a preferred embodiment, step d) is performed by subjecting the final sample of step c) to a 60-min HPLC gradient coupled to a tandem MS/MS analysis, wherein the resulting MS/MS spectra dataset has been divided into interpretable sub-datasets corresponding to short windows of elution times, for example of less than 2 minutes, typically of about 1 minute or of about 30 seconds.
Some years ago, it has been shown that with a specific set-up for injecting samples and preformed gradient, the proteomic analysis of more than 200 samples per day was possible ([32]), thus 7.2 min of tandem mass spectrometry per sample. Here, the objective
being different, the time per isolate was reduced with a classical nanoLC-MS/MS set-up. Combining the present label-free multiplexing approach with such system able to preform gradient could result in even better results. Furthermore, the use of double-barrel liquid chromatography system coupled to the tandem mass spectrometer has been shown to be highly efficient in reducing the mass spectrometry time per sample ([2]; [33]). Such option that can be easily implemented for the present strategy could further reduce the costs of mass spectrometry per sample.
Step e): identification of the peptides contained in the fractions
Step e) consists in analysing the MS/MS spectra obtained in step d) (in particular those contained in the smaller sub-datasets) by comparing them with a database of known peptide sequence data and matching each peptide sequence of said MS/MS spectrum to one or more organisms present in said database.
Several approaches have been proposed for the interpretation of the list of peptides that can be identified in a sample in order to identify the taxa present in a sample: Unipept- based proteotyping associates each peptide identified with a taxon via the UniProtKB database based on the lowest common ancestor approach [5], and is thus based on taxonspecific peptides. Alternatively, TCUP uses the attribution of the spectra to peptides / proteins and then to a taxon using a dedicated proteome database corresponding to a limited number of taxa potentially present in the sample [6], A third approach, based on ProteoClade software, uses also specific peptides to assign peptides to an organism based on predefined database of theoretical proteomes [7], TaxIT applies a two-step analysis, with initial selection of the most relevant species contained in the sample, followed by searches of a reduced database based on the species selected to allow precise identification [8], In the method of the invention, it is possible to use any of these
taxonomic workflows comprising databases which have been fully annotated and attributed to specific taxa. The matching of the peptide sequences from the sample to taxa is performed by comparing each of these sequences to sequences present in public and/or private sequence databases in which sequences are annotated with taxon information, using techniques of sequence identity searches, and more generally of sequence similarity searches such as BLAST. This leads to a list of taxa (which may be any taxonomical grouping) and these are sorted in the order of the number of matching sequences they comprise, from the maximum value to the least. All these workflows are based on taxonspecific peptides.
Recently, phylopeptidomics-based proteotyping was proposed as a means to take advantage of taxon-specific peptides and peptides shared by closely- and distantly-related organisms at all possible taxonomic ranks [3], Based on predictable specific signatures for each microorganism present in the sample when querying a comprehensive database, this approach allows microorganisms to be taxonomically identified even in complex mixtures, and their relative biomasses quantified. The approach was applied to screen isolates from several environments [9], can be amenable to high-throughput [10], and is effective even on poorly-documented branches of the tree of life [11], This approach is thoroughly explained in WO2015/019245, which is herein incorporated by reference.
With this particular approach, Taxon-to-Spectrum Matches (TSM) and taxon-specific peptide sequences (spePEP) can be generated as described in [3], [18], [11], and [23], for example.
In accordance with the present invention "taxon" means a group of one (or more) populations of organism(s), which a taxonomist adjudges to be a unit.
In accordance with the present invention, each of the spectra obtained in step d) is attributed to a peptide sequence encompassed into one or more database sequences and hence assigned to a taxon or a series of taxa, with the total number of matches per taxon being recorded. Once all the spectra from the sample have been tentatively matched, the taxon with the most matches is identified, and the number of matches corresponding to this taxon is defined. In the case of the presence of several taxa, all the taxa with a significant number of matches are identified.
In accordance with the present invention "Taxon to spectrum match(es)" or "TSM(s)" means a match from a MS/MS spectrum query to a given Taxon. As described below, the number of TSMs will be used to identify the taxonomy of the organisms present in the initial biological samples.
In accordance with the present invention "taxon-specific peptide sequence(s)" or "spePEP(s)" designates the peptide sequences that are specific to a taxon, as determined by peptide sequence sequence search on the whole database of theoretical proteomes. As described below, the number of spePEPs will be used to identify the organisms present in the initial biological samples.
Step f): determination of the organisms present in the X samples
This last step consists in identifying the organisms present in each of the X biological samples, by attributing the organisms identified in step e) to each initial biological sample. This is done by using the specific elution characteristics at which each organism was identified.
These elution characteristics are preferably obtained with a polynomial regression model that can be used whatever the changes of technique between step b) and d) that could
introduce hydrophobicity shifts. For example, for each 1-min window, the TSMs and spePEPs for the microbial species identified can be extracted from independent searches and normalized relative to the total number of bacterial TSMs and spePEPs per sample, respectively.
This step preferably requires calculating, for each given acquisition window, an index which is herein called “Species Proteotyping Index” or “SPi” for each identified taxon and each fraction:
SPi =1 x TSMs + k x spePEP wherein “TSMs” is the number of TSMs in said acquisition window, “spePEP” is the number of taxon-specific peptide sequences in said acquisition window, and k is a number comprised between 1.5 and 2.5.
This index which is calculated for each identified taxon, combining the results of TSM and spePEPs of the taxon, varies depending on the retention time. A mathematical modelisation of the SPi index along the retention time allows identifying the SPi maximum or the SPi maxima (in case of multiple identification of the same taxon in various fractions). Once it is calculated, it is possible to identify the elution characteristic of the taxon which corresponds to the elution value for which the maximum value of SPi for each organism is obtained. As shown in the example below, it is in particular possible to fit nonlinear regression curves of SPi and retention time, for example with a software such as GraphPad Prism in order to obtain the SPi maximum value.
In a preferred embodiment, only organisms with SPi values > 0 over more than three consecutive acquisition time intervals are retained in the analysis.
With this SPi index, it is possible to calculate the retention time (center value) of each species and the r-squared value of the modelisation of SPi maxima versus retention time as an indication of confidence. Thereafter, it is possible to determine reliably the exact elution characteristics of each of the identified organisms, thus their respective sample origins.
Additional step: evaluation of the quantity of peptides
The method of the invention enables not only to identify but also to quantify the relative ratio of organisms present in each of the starting biological samples. For this purpose, TSM values of each of the organism identified at a given retention time can be directly compared. Furthermore, the absolute quantity of protein biomass of each organism identified at a given retention time can be measured by the specific tandem mass spectrometry signal assigned to each organism. For this purpose, specific quantities of isotopically labelled standards can, for example, be subjected to MS/MS analysis together with the peptides from the samples and a standard curve can be constructed from the ion signals generated from these standards. Using this standard curve, the relative abundance of a given ion can be converted to an absolute amount of the corresponding protein in the sample. Other alternatives relying on mass spectrometry of specific quantities of standards performed separately in exactly the same conditions can be carried out.
Therefore, the method of the invention allows both to generate MS/MS spectra characteristic of the peptides / proteins contained in the sample to be analyzed, and from these spectra to quantify the protein biomass contained in the sample to be analyzed and, via these characteristics, to quantify the relative proportion of organism(s) present in said sample to be analysed or the absolute amount of organism(s) present in said sample to be
analysed in referene to standards. It also allows to quantify the amount of particular peptides, for example peptides that are characteristic of antibiotic resistance, virulence factors, or toxins.
The measurement of peptide abundance can be performed by a method selected from the group comprising a method using spectral counts, extracted Ion Chromatograms, a quantification method based on mass spectrometry data or associated liquidchromatography data, the MS/MS total ion current and methods based on peptide fragments isolation and quantification such as selected reaction monitoring (SRM), multiple reaction monitoring (MRM) or parallel reaction monitoring (PRM).
In a particular embodiment, the method of the invention therefore also contains an additional step h’) of calculating the relative or absolute amount of the detected organisms in each biological sample.
In this embodiment, the quantities of proteins subject to proteolysis have been preferably normalised before step a) or if not, the quantity of peptides in each eluted fraction of step b) has been calculated so that the fractions mixed in the final sample contain the same equivalent quantities of peptides.
X biological samples
The number X of biological samples that can be analyzed by using the present method is not limited. It is typically superior to two (2), so that multiplexing is effectively achieved. This number X can be extended to 50, 100, 150, 180, 200, 250, 300, 500, 1000 or even more, once the mass spectrometer devices allow longer runs to be performed or are able to analyze more data in shorter time.
The innovative label-free multiplex strategy described here can thus be applied to any biological sample containing any kind of organisms, e.g., microbial isolates, or part of organisms, e.g. blood or saliva of an animal, leg or antenna of a parasite, or floor extracted from a cereal. This sample may even be of unknown nature such as a bioterrorism type sample, an extraterrestrial sample, or a new microbial isolate that has not been yet taxonomically characterized. Generally speaking, it is any sample for which it is desired to characterize by MS/MS the biological material of a protein nature that it contains or by which it is contaminated.
Advantageously, this sample can be a biological fluid; a plant fluid such as sap, nectar and root exudate; a sample in a culture medium or in a biological culture reactor such as a cell culture of higher eukaryotes, yeasts, fungi, bacteria, viruses or algae; a liquid obtained from an animal or plant tissue; an animal or plant tissue; one or more cells; a cell pellet; a sample in a food matrix; a sample in a chemical reactor; a sample from a wastewater treatment plant; a sample from a composting plant; city water, river water, pond water, lake water, sea water, swimming pool water, water from cooling towers or water from underground sources; a sample from a liquid industrial effluent; wastewater from intensive livestock farming or chemical, pharmaceutical or cosmetic industries; a sample from an air filtration or a coating; a sample from an object such as a piece of fabric, a garment, a sole, a shoe, a tool, a weapon, etc. A sample of an object such as a fragment of fabric, a garment, a sole, a shoe, a tool, a weapon, etc. may be taken from the following objects: a pharmaceutical product; a cosmetic product; a perfume; a soil sample or a mixture thereof.
The biological fluid is advantageously selected from the group consisting of blood such as whole blood or anti -coagulated whole blood, blood serum, blood plasma, lymph, saliva, spit, tears, sweat, semen, urine, stool, milk, cerebrospinal fluid, interstitial fluid, isolated bone marrow fluid, mucus or fluid from the respiratory, intestinal or genitourinary tract, cell extracts, tissue extracts and organ extracts. Thus, the biological fluid can be any fluid naturally secreted or excreted from a human or animal body or any fluid recovered, from a human or animal body, by any technique known to the skilled person such as extraction, sampling or washing. The steps of recovery and isolation of these different fluids from the human or animal body are performed prior to the implementation of the process according to the invention. The sample could be any derived products such as organoids, tumoroids, iPSC cells, or CAR-T cells.
The sample used in the context of the present invention may contain, by its nature or possibly as a result of contamination, one or more cells (identical or different). In the context of the present invention, "cell" is understood to mean both a cell of prokaryotic type and a cell of eukaryotic type. Among the eukaryotic cells, the cell may be a yeast such as a yeast of the genus Saccharomyces or Candida, a fungus, a parasite, a plant cell or an animal cell such as a mammalian or insect cell. The prokaryotic cells are bacteria which can be gram positive or gram negative, or archaea. Among these bacteria, we can mention, as examples and in a non-exhaustive way, the bacteria belonging to the spirochetes and chlamydiae phyla, the bacteria belonging to the families of the enterobacteria (such as Escherichia cold), the streptococcaceae (such as Streptococcus'), the microccaceae (such as Staphylococcus), the legionellae, the mycobacteria, the bacillaceae, the cyanobacteria and other. Among these archaea, we can mention, as
examples and in a non-exhaustive way, the archaea belonging to the phyla of Crenarchaeotes (such as Sulfolobus) and Euryarchaeotes (such as Thermococcus).
The sample used in the present invention may include one (or more) pathogen(s). In this case, prior to the implementation of the process according to the invention, the sample may be subjected to a pre-treatment to inactivate it. Any technique commonly used to inactivate a sample comprising pathogens can be used as part of this pre-treatment. Advantageously, this pre-treatment can consist of heating said sample at 99°C for a period of time between 20 min and 90 min and, in particular, of the order of 60 min (i.e. 60 min ± 10 min) and then allowing it to cool to room temperature (i.e. 22°C ± 3°C). Alternatively, this pre-treatment may consist of heating said sample to a temperature above 99°C, in particular above 100°C and in particular of the order of 110°C (i.e. 110°C± 5°C) for a period of between 5 min and 45 min and, in particular, of the order of 20 min. Inactivation of the pathogen can be also performed with other means, such as the use of chemical reagents, or a combination of methods, e.g. chemical reagents and temperature treatment.
In the following, "sample" is understood to mean both a sample that has undergone this pre-treatment and a sample that is pathogen-free and, in fact, has not undergone such inactivation.
Accordingly, the organism present in the sample to be analyzed can be e.g., a virus, a bacterium, an archaeon, a yeast, a fungus, an algae, a plant, an animal, a parasite, or any other organism, pathogenic or not, alive or not, taxonomically characterized or not.
In a particular embodiment, the sample analyzed by the methods of the invention contains only one kind of organism, e.g., one type of virus, one type of bacterium, etc.
In another particular embodiment, the sample analyzed by the methods of the invention contains more than one kind of organisms. It can e.g., contain several specie of one organism (e.g., several bacterial specie or several yeast strains), or a mixture of several organisms (e.g., bacteria and viruses, or yeast and fungi, or a virus and its animal host).
If the number X of samples to analyze is so high that the method cannot be performed on each sample separately in a convenient manner by the mass spectrometer currently available (e.g., because their possibility are limited by the duration of the runs), then it is possible to pool a number of samples together and then to apply the method of the invention to the pooled sample. For example, if there are 2000 samples to analyze, one can pool several samples together to reduce the number of samples to multiplex and analyze on the MS/MS run, and depending on the results, reanalyze the samples for which the pool gives an interesting result. This amount can be easily adjusted by the skilled person in view of the number of samples, the need to detect the organisms in each sample or in group samples, the performance of the mass spectrometer devices, etc.
FIGURE LEGENDS
Figure 1. Workflow for multiplex proteotyping of bacterial isolates using Phylopeptidomics.
Figure 2. Distribution of TSMs attributed to P. putida (A), K. aerogenes (B) and R. pickettii (C) in 1-min acquisition windows. A) Distribution of total TSMs for 1-min gradient intervals over the whole 60-min chromatographic run for P. putida, B)
Distribution of total TSMs for 1-min gradient intervals over the whole 60-min chromatographic run for K. aerogenes. C) Distribution of total TSMs in 1-min gradient intervals over the whole 60-min chromatographic run for R pickettii. Data were obtained by shotgun proteomics. W: window.
Figure 3. Number of TSMs attributed at different retention times for a single R. pickettii HPLC fraction. Number of TSMs pointing to the identification of R. pickettii in fraction 1 from R. pickettii for each 1-min acquisition period.
Figure 4. Scatter plot of Species Proteotyping index (SPi) for R. pickettii and Sagitulla stellata. Proteotyping results of R pickettii for which two fractions contributed to Ml 1 is represented in squares. Proteotyping results of Sagittula stellata for which only one fraction contributed to Ml 1 is represented in dots. Lorentzian fitting are indicated. A and B represents the SPi maximum of the two fractions of R. pickettii in Mi l and C represent the SPi maximum of the fraction of Sagittula stellata.
Figure 5. The fraction number correlates almost perfectly with the corresponding proteotyping-retention time value (SPi maxima). A second polynomial regression fitting was established using the theoretical fractions with the SPi maximum in chronological order; the equation y = 0.7018x2 +94.781x + 662.8 was used to determine the SPi identified for each fraction.
Figure 6. Scatter-plot of SPi values of each organisms in 20-sec acquisition time window. SPi values of each organisms composing the mix are represented as function of 20 sec acquisition time. The first line is associated to S. stellata, the second line to R.
pomeroyi, the third line to O. indolifex, the fourth line to S. cerevisiae, the fifth line to K. aerogenes and the last line to M. tractuosa.
Figure 7. Retention time of each fractions of the 23 mix are reproducible. A box plot was established with all retention time value of the 23 mixtures as function of each fraction (Fl, F2, F3, F4, F5, F6). For the six fractions, the median is at 554, 666, 780, 894, 1010, and 1159 sec of chromatography, respectively. The minimum and maximum values are [536-570], [639-682], [764-803], [880-908], [993-1026], and [1143-1170], respectively.
Figure 8. Scatter plot of Species Proteotyping index of mixtures with different quantities. A) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to K. aerogenes with a peptide quantity of 280 ng, the following peaks are associated to AT. tractuosa (60ng), O. indolifex (100 ng), R. pomeroyi (340 ng), S. cerevisiae (180ng) and S. stellata (80ng). B) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to S. stellata with a peptide quantity of 140 ng, the following peaks are associated to R. pomeroyi (120 ng), M. tractuosa (40 ng), S. cerevisiae (160 ng), K. aerogenes (160 ng) and O. indolifex (100 ng).
Figure 9. Scatter plot of Species Proteotyping index of mixtures containing a same organisms in several fractions. These scatter plot show the distribution of SPi values as function of 20-s window acquisition time with a same organism present twice in a mix. The peptide quantity of each fractions was normalized (100 ng). Same organisms are represented with the same color line. A) The SPi values of each organisms are plotted and fit by Lorentzian model. R. pomeroyi, K. aerogenes, S. cerevisiae, R. pomeroyi, K.
aerogenes and S. cerevisiae are associated to each peak in the left to right order. B) The SPi values of each organisms are plotted and fit by Lorentzian model. S.stellata, R. pomeroyi, O. indolifex, K. aerogenes, O. indolifex and S. cerevisiae are associated to each peak in the left to right order. C) The SPi values of each organisms are plotted and fit by Lorentzian model. S. cerevisiae, R. pomeroyi, K. aerogenes, S.stellata, K. aerogenes, and R. pomeroyi are associated to each peak in the left to right order. D) The SPi values of each organisms are plotted and fit by Lorentzian model. R. pomeroyi, K. aerogenes, S. cerevisiae, S.stellata, S. cerevisiae and K. aerogenes are associated to each peak in the left to right order.
Figure 10. Scatter plot of Species Proteotyping index of mixtures containing a same organisms in following fractions. These scatter plots show the distribution of SPi values as function of 20-s window acquisition time with a same organism present in consecutive fractions within a mix. Large wide peak is an indicator of the presence of a same organism in consecutive fractions. The peptide quantity of each fractions was normalized (100 ng). Same organisms are represented with the same color line. A) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to O. indolifex, the second line to S. cerevisiae, the third line to R pomeroyi and the last line to K. aerogenes. B) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to S. cerevisiae, the second line to K. aerogenes, the third line to Rpomeroyi and the last line to S. cerevisiae. C) The SPi values of each organisms are plotted and fit by Lorentzian model. The first line is associated to K. aerogenes, the second line to S. stellata, the third line to R. pomeroyi, the fourth line to O. indolifex and the last line to S. cerevisiae.
EXAMPLES
EXAMPLE I: use of C18 reverse phase chromatographic for peptide separation
MATERIAL & METHODS
1. Microbial cultures
For this study, 20 microbial strains were used, 13 from a commercial source and 7 bacteria isolated in the case of a previous study [2], Bacteria were grown under aerobic conditions in liquid media until the stationary phase was reached. Culture media were lysogeny broth (LB; BD Bacto) IX and diluted 0.1X in water; trypticase soy broth (TSB; Biomerieux) diluted 0.5* in water; 0.1X TSB+ trace elements, as described by Vujicic-Zagar [13], (diluted TSB); 10 g TSB+ 2 g Yeast extract + 1 g Beef extract based on DSM medium #948 (TSB & extracts); Marine Broth (MB); Peptone-Tryptone- Yeast extract-Glucose medium (PTYG); and low-nutrient R2A broth (R2A14) diluted 0.5 X in water.
Cells from 5-mL cultures were harvested by centrifugation at 4 000 g for 10 min. A second centrifugation for 2 min was systematically performed to remove residual liquid medium. Resulting pellets were weighed and stored at -20 °C until use.
2. Cell lysis and production of peptide digests
Cell pellets were resuspended in a specific volume (100 pL per 1.7 mg wet biomass) of lx lithium dodecyl sulfate sample loading buffer (Thermo Fisher Scientific) prepared without dyes and supplemented with 5% beta-mercaptoethanol (v/v). Protein extracts were prepared as described in Hayoun et al. [12], Briefly, after heating for 5 min at 99 °C, samples were sonicated for 5 min in an ultrasonic water bath (VWR ultrasonic cleaner USC 300T). Samples were then transferred to 2-mL or 0.5-mL screw cap microtubes
(Sarstedt) containing an equimolar mix of 0.1 mm silica beads, 0.1 mm glass beads, and 0.5 mm glass beads. Cells were disrupted using a Precellys Evolution bead-beater (Bertin Technologies) operated at 10 000 rpm for 10 x 30-s cycles, with 30-s pauses between each cycle. Beads and debris were removed by centrifugation at 16 000 g for 2 min. Supernatants were transferred into new microtubes and incubated at 99 °C for 5 min. Lysates were stored at -20 °C until use. For SP3 proteolysis, a stock solution of Sera-Mag carboxylate-modified magnetic beads consisting of a 50:50 mixture of hydrophilic and hydrophobic beads (Sigma) in Milli-Q water was prepared, as recommended [15], Samples were then treated according to the protocol developed by Hayoun et al. [10] in 96-well plate format. Briefly, 40 pg of magnetic beads (4 pL) were added to each cell lysate (20 pL), followed by half the total volume of formic acid, and CH3CN to a final concentration of 85%. Trapped proteins were purified by successive washes with 70% ethanol and CH3CN using a Smart2 MBS (Tecan) neodymium magnetic rack. Tryptic peptides were produced by re-suspending the paramagnetic beads in 10 pL of 50 mM NH4HCO3 supplemented with 0.01% Protease Max surfactant (Promega) and 1 pg.pL-1 trypsin gold (Promega). Samples were incubated on the plate for 15 min at 50 °C. After removal of the paramagnetic beads, the resulting 8-pL peptide solution was acidified by addition of trifluoroacetic acid (0.5% final concentration) and the total volume adjusted to 50 pL with aqueous 0.1% TFA. The concentration of tryptic digests was measured using a Pierce colorimetric Peptide Assay, as recommended (Thermo Scientific).
3. Off-line HPLC fractionation of peptide digests
Peptide digests were fractionated off-line by reverse-phase chromatography on an Agilent 1100 series HPLC system equipped with a UV detector. Peptides were separated on a
ZORBAX StableBond Cl 8 column (Agilent) with particle size 5 pm, average pore size 300 A, length 15 cm and internal diameter 4.6 mm. A total of 25 pg of peptides was loaded onto the column, and a 30-min linear gradient was developed from 2.5% to 50% B at a flow rate of 400 pL.min'1. Mobile phases A and B were an aqueous solution of 0.1% formic acid (FA) in water, and 0.1% FA in acetonitrile, respectively. Fractions were collected every 30 s for 30 min after a 5-min delay corresponding to the dead volume. An Agilent 1100 G1364C fraction collector was used to collect fractions in 96-well plates. Microplates were stored at 4 °C until use.
4. Assembly of Mi l and M21 concatenated samples A 40 pL-volume of each of the chosen fractions were combined in a single tube. Concatenates were thoroughly mixed. Samples were then dried down and resuspended in 50 pL TFA 0.1% for nanoLC-MS/MS analysis. The Ml 1 sample was a mixture of eleven 1-min fractions eluted from a 30-min gradient, as presented in Table 1. The fractions included were: Fl; F3; F5; F7; F9; Fl l; F13; F15; F17; F19; F21. The M21 sample corresponded to a mixture of twenty-one 30-s fractions - Fl; F2; F3;... : F21 - eluted from a 30-min gradient, as presented in Table 1.
Table 1: Composition of the concatenated fractions samples making up Mil and M21
5. Tandem mass spectrometry
NanoLC-MS/MS analyses were conducted using an ultimate 3000 nanoLC system (Thermo Fisher Scientific) coupled to a Q-Exactive HF tandem mass spectrometer (Thermo Fisher Scientific) as described [16], Peptides were desalted on a reverse-phase PepMap 100 C18 p-precolumn (5 pm, 100 A, 300 pm i.d. x 5 mm; Thermo Fisher Scientific). Desalted peptides were separated on a nanoscale PepMap 100 C18 nanoLC column (3 pm, 100 A, 75 pm i.d x 50 cm; Thermo Fisher Scientific) at a flow rate of 0.3 pL.min'1, applying a two-slope 60-min gradient as follows: 4-25% B from 0 to 50 min and 25-40% B from 50 to 60 min [10], Mobile phase A consisted of an aqueous solution of 0.1% (v/v) FA in water; phase B consisted of 0.1% FA in 80% acetonitrile. The mass spectrometer was operated in data-dependent mode (DDA) with an MS acquisition scan range of 350 to 1500 m/z. The 20 most abundant precursor ions were selected for fragmentation, applying a 10-s dynamic exclusion window and a 1.6-m/z isolation window. Only ions with a 2+ or 3+ charge were selected. HCD fragmentation was carried out with a normalized collision energy of 27 eV.
6. MS/MS data interpretation for proteotyping at the species taxonomical rank
Proteotyping was performed as previously described [11,17], Briefly, ion peak lists were extracted using Mascot Daemon software, version 2.6.0 (Matrix Science), with the following settings: minimum mass (400), maximum mass (5000), grouping tolerance (0), intermediate scans (0), and threshold (1000). A Python script was applied to split the molecular ion peak list file into 1-min acquisition periods. MS/MS spectra for each of the resulting files were interpreted using Mascot version 2.6.1 (Matrix Science) against the NCBInrS database [18], an in-house assembled subset of the National Center for Biotechnology Information non-redundant (NCBInr) database (downloaded the 3rd of
January 2018 as ftp://ftp.ncbi.nlm.nih.gov/blast/db/FASTA/nr.gz). Filtering was applied to the database to retain one representative taxon per species belonging to the Bacteria, Archaea, and Eukaryota superkingdoms. This database contains 50 080 649 protein sequences totaling 20054 069 172 amino acid residues. A second query was performed on a database comprising all known annotated genomes from NCBInr associated with the genera identified in the first-round search. A third query was performed on an even more reduced database comprising only the annotated genomes of the strains belonging to the species identified in the second query.
For MS/MS spectrum -to-peptide assignment, the following parameters were applied in the first search: mass tolerance of 3 ppm on parent ion, 0.02 Da on MS/MS, 2+ or 3+ as possible peptide charges, a maximum of one missed cleavage, carbamidomethylation of cysteine as fixed modification, oxidation of methionine as variable modification, trypsin as proteolytic enzyme. In the second and third searches, the same parameters were used, except for a mass tolerance of 5 ppm on parent ions and a maximum of two missed cleavages was allowed. The p-values used for peptide validation were 0.3, 0.15, and 0.05 in homology threshold mode for the first, second, and third search rounds, respectively. Peptide sequences were mapped to taxa at the species, genus, family, order, class, phylum, and superkingdom taxonomical ranks, as previously described [2,3], resulting in Taxon- to-Spectrum Matches (TSMs). Taxonomies were identified based on these TSMs and from the number of taxon-specific peptide sequences (spePEP). For each 1-min window, the TSMs and spePEPs for the microbial species identified were extracted and normalized relative to the total number of bacterial TSMs and spePEPs per sample, respectively. Only species for which TSMs and spePEPs were assigned in more than three consecutive acquisition time intervals were retained. A filter combining the two parameters (lx TSMs
+ 2x spePEPs) for each 1-min window for each species was used to calculate the retention time (center value) and the r-squared value as an indication of confidence. This “Species Proteotyping index” (SPi) filter varied depending on the retention time. GraphPad Prism version 6.0 (GraphPad software) was used to fit nonlinear regression curves of SPi and retention time.
Data availability
All mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE partner repository under dataset identifiers PXD035870 and 10.6019/PXD035870.
RESULTS
Strategy for multiplex proteotyping of microbial isolates based on hydrophobicity fractionation of peptides
The label-free multiplexing protocol described here is a new method developed to increase the throughput of tandem mass spectrometry proteotyping of microbial isolates. As shown in Figure 1, the concept involved the production of peptide mixtures from each individual microbial isolate. These peptide mixtures were then fractionated based on their hydrophobicity, and multiple fractions were subsequently concatenated to associate a specific reverse-phase chromatography retention time with each isolate. After merging the different peptide fractions, a single nanoLC-MS/MS analysis was performed to sequentially generate MS/MS spectra for each organism over a low- to high- hydrophobicity sequence. To establish the nature of each microorganism in this type of multiplexed nanoLC-MS/MS run, a highly discriminative proteotyping approach must be
employed, ensuring identification without false-positives at each hydrophobicity window. For this tandem mass spectrometry-based proteotyping, MS/MS spectra were first assigned to peptide sequences, resulting in peptide-to-spectrum-matches (PSMs). The TSMs at the different taxonomic ranks could then be extracted by mapping the peptides to all theoretical proteomes contained in the database. Taxon-specific peptides were also obtained based on the search for the lowest common ancestor. TSMs and specific peptides were then used to identify the microorganisms present in the sample, as described in Materials & Methods.
Bacterial isolates can be proteotyped from a short acquisition window, even with a low number of MS/MS spectra
The tryptic peptides obtained after proteolysis of the soluble proteomes of the three species - Pseudomonas putida, Klebsiella aerogenes, and Ralstonia pickettii - were analyzed by shotgun tandem mass spectrometry on a Q-Exactive HF instrument following a 60-min gradient. The resulting datasets comprised 37 139, 35 759, and 27 339 MS/MS spectra, respectively. Query of the comprehensive NCBInr database resulted in 21 749, 20 701, and 12 073 PSMs, respectively. Thus, the ratio of MS/MS spectra interpreted was 59%, 58%, and 44%, respectively. The number of unique peptide sequences was 9 706, 9 137, and 7 401, respectively. The total number of TSMs extracted at the species level was 21 825 for P. putida, 20 681 for K. aerogenes, and 11 856 for R. pickettii. Importantly, the general profile of the peptides eluted from the reverse-phase chromatography was quite similar for the three peptidomes.
For the present label -free proteotyping concept, it was assessed whether short, 1-min, windows of tandem mass spectrometry data were sufficient to confidently identify the expected species. The P. putida peptidome dataset was split into 60 short windows corresponding to 1 min of acquisition time. Each of these small datasets was analyzed without a priori on the proteotyping pipeline. The P. putida species was systematically identified in each of the fractions, based on significant levels of TSMs and spePEPs. As shown in Figure 2 (Panel A), a significant number of TSMs was obtained for most fractions. However, the first five fractions and the last five fractions - the most hydrophilic and hydrophobic peptides, respectively - produced lower numbers of TSMs. For the other 50 fractions, an average of 396 ± 73 TSMs was obtained. Therefore, the TSM signal was relatively stable over a very broad separation window. The number of spePEP sequences for each fraction was similarly distributed, with an average of 227 ± 44 spePEPs per fraction, excluding the five most hydrophilic and five most hydrophobic fractions. Similar results were observed with the datasets acquired on K. aerogenes (357 ± 95 TSMs and 184 ± 66 specific peptides) and R. pickettii (257 ± 73 TSMs and 13 ± 7 specific peptides), as shown in Figure 2 (Panel B & C). These results indicated that a 1- min acquisition with the Q-Exactive HF instrument should be sufficient to identify the expected species by proteotyping.
Subsequently, the peptidome for the pure R. pickettii isolate was resolved into 60 x 30-s fractions by reverse-phase chromatography with a 30-min acetonitrile gradient. One of the fractions was selected, Fl - eluting at 12 min - for analysis by nanoLC-MS/MS with a 60-min acetonitrile gradient. The resulting dataset was subdivided into 60 small subdatasets corresponding to 1-min acquisition windows, and each of these sub-datasets was interpreted for proteotyping. For 11 contiguous fractions, the interpretation undoubtedly
pointed to the presence of R. pickettii. These results confirm that a 1-min acquisition window with the Q-Exactive HF instrument provides enough information to proteotype the expected species. Figure 3 shows the distribution of the number of TSMs allowing the identification of R. pickettii plotted against the retention time. These results show that it is possible to identify an isolate using a single HPLC fraction, equivalent to a reduced proportion of the eluted peptides.
Multiplex proteotyping applied to 10 isolates analyzed in a single nanoLC-MS/MS run.
In addition to the R. pickettii peptidome, nine other microbial peptidomes were resolved individually by reverse-phase chromatography with a 30-min gradient of acetonitrile, systematically collecting 30-s fractions. A sample consisting of specific fractions of these 9 peptidomes was assembled, together with two fractions from the A. pickettii peptidome. Thus, a duplicate of one of the isolates at different chromatographic elution times was included. The mixture of peptides (Ml 1) was then injected into a nanoLC reverse-phase column coupled to a tandem mass spectrometer and subjected to a 60-min gradient and MS/MS analysis. The resulting MS/MS dataset comprised 30 881 MS/MS spectra. Once again, the data were split into 1-min acquisition windows for interpretation. We propose to name the combination of TSMs and spePEPs at a given retention time the “Species Proteotyping index” (SPi), as both parameters contribute to the proteotyping result. To increase the specificity of this index, a weighting factor of two was applied to spePEPs. Ten species were identified when all the fractions were merged, and only species with SPi values > 0 over more than three consecutive acquisition time intervals were retained. The ten species identified perfectly matched those expected for this Ml 1 sample.
To establish the link between each fraction and the corresponding organism, a proteotyping-retention time corresponding to the SPi maximum was established for each isolate. To do so, SPi values were plotted according to the 1-min windows for each of the ten species contained in the mixture. Figure 4 shows the plots obtained for two species, namely Sagitulla stellata and R. pickettii. As observed from this figure, well-defined elution peaks were obtained. The proteotyping interpretation profile for each species has a bell shape representative of the elution of the peptides from each specific fraction included in the mixture. The shape of these retention-time-resolved proteotyping results allows the center of the elution peak to be precisely defined through experimental-data curve fitting with a Lorentzian function. As expected, a single peak was observed for S. stellata fraction Fl 5 with a retention time (A stellata SPi maximum) at 37.6 min, whereas two well-separated peaks were observed for R. pickettii at 12.4 min and 22.4 min (R. pickettii SPi maxima). These results are in full agreement with the two fractions, Fl and F7, from this organism contributing to the Ml 1 assemblage. Lorentzian curve fitting was applied to each peak corresponding to the eight other fractions of Mi l to determine the respective proteotyping elution time: Sphingomonas yabuuchiae F3 at 16.3 min, Microbacterium oxydans F5 at 19.3 min, Stenotrophomonas maltophilia F9 at 26.1 min, Serratia marcescens Fl l at 29.7 min, Kineococcus radiotolerans F13 at 33.5 min, Pseudopedobacter saltans F17 at 41.5 min, Pseudomonas aeruginosa F19 at 45.3 min, and Methylobacterium extorquens F21 at 49.3 min. As expected, these proteotyping- retention times were increasing with the fraction numbers from the initial chromatography runs from which Ml 1 was assembled. Figure 5 shows the relationship between fraction numbers and SPi maximum. A second polynomial regression fitting gave y = 0.7018x2 + 94.781x + 662.8, with a correlation coefficient (r-squared) of 0.9997.
The average difference between two consecutive proteotyping-retention times was thus 221 (±21) s, allowing clear distinction between the fractions. Consequently, within the F1-F21 range, a larger number of fractions corresponding to even shorter hydrophobic windows could be introduced to allow greater multiplexing.
Label-free multiplex proteotyping is successful on fractions assembled from 20 isolates
The multiplexing capacity of the method of the invention was next studied. The M21 chimeric sample was assembled by mixing 21 distinct 30-s fractions from 20 individual peptidomes prepared from isolates. Like for Mi l, two fractions from the R. pickettii isolate were included in this assemblage as a control. The peptidome assemblage was once again analyzed by shotgun tandem mass spectrometry with a 60-min gradient, resulting in a dataset comprising 28 883 recorded MS/MS spectra. Proteotyping analysis, carried out as described for the Ml 1 mixture, revealed the presence of 20 distinct species in the dataset: Bacillus cereus, Bacillus thuringiensis, Deinococcus deserti, Deinococcus proteolyticus, K. radiotolerans, K. aerogenes, Massilia timonae, Marivirga tractuosa, M. extorquens, M. oxydans, Oceanibulbus indoliflex, P. aeruginosa, P. putida, P. saltans, R. pickettii, Ruegeria pomeroyi, S. stellata, S. marcescens, S. yabuucchiae, and S. maltophilia. Notably, several pairs of isolates belonging to the same genus were identified, and discriminated between at the species level: two Pseudomonas, two Bacillus, and two Deinococcus.
Fractio Expected SPi low SPi high Bacterial Measured Fraction n used SPi fraction fraction species SPi identifie maximu limit (s) limit (s) identified maximu d m (s) in
(s)
1 755.009 704.746 805.469 S. l l 1 maltophilia
2 856.126 805.469 906.98 D. 896 2 proteolyticus
3 958.031 906.98 1009.27 B. 939 3
9 thuringiensis
4 1060.724 1009.27 1112.36 S. 1065 4
9 6 marcescens
5 1164.205 1112.36 1216.24 M. timonae 1165 5
6 1
6 1268.474 1216.24 1320.90 K. aerogenes 1222 6
1 4
7 1373.531 1320.90 1426.35 P. putida 1345 7
4 5
8 1479.376 1426.35 1532.59 S. stellata 1474 8
5 4
9 1586.009 1532.59 1639.62 R. pickettii 1570 9
4 1
10 1693.43 1639.62 1747.43 K. 1701 10
1 6 radiotoleran s
11 1801.639 1747.43 1856.03 R. pomeroyi 1790 11
6 9
12 1910.636 1856.03 1965.43 D. deserti 1943 12
9
13 2020.421 1965.43 2075.60 O. indolifex 2042 13
9
14 2130.994 2075.60 2186.57 5. 2146 14
9 6 yabuuchiae
15 2242.355 2186.57 2298.33 M. oxydans 2243 15
6 1
16 2354.504 2298.33 2410.87 B. cereus 2342 16
1 4
17 2467.441 2410.87 2524.20 R. pickettii 2514 17
4 5
18 2581.166 2524.20 2638.32 P. saltans 2579 18
5 4
19 2695.679 2638.32 2753.23 M. 2694 19
4 1 extorquens
20 2810.98 2753.23 2868.92 P. 2780 20
1 6 aeruginosa
21 2927.069 2868.92 2985.40 M. tractuosa 2921 21
6 9
Table 2 shows the attribution of each organism to its corresponding fraction.
Table 2. Correlation between SPi maximum and fraction origin for the 20 species identified in the M21 assemblage.
Values obtained for mix M21 are shown. The first column lists the experimental fraction used to create the mix. The 'Expected SPi max’ is a theoretical value calculated using the correlation curve. The Tow’ and ‘high’ fraction limits correspond to the range for each fraction. The fraction identified must fall within the range. The ‘Measured SPi maximum’ corresponds to the value obtained from the nonlinear regression curve, corresponding to the maximum of the SPi. The ‘fraction identified’ column corresponds to the theoretical calculation of the fraction number from the SPi maximum measured.
In the first four columns, the theoretical time is calculated as an approximation of the expected retention time. These time values were calculated using the second polynomial equation curve for M21. The low fraction and high fraction limits indicate the theoretical limits of the range for each fraction. The Lorentzian nonlinear regression of the SPi values for each microorganism was used to determine the retention time for each species identified. The 21 SPi maximum for the different species agreed perfectly with the chronological order, with values between the theoretical low and high limits. The second polynomial regression (y = 0.394x2 + 99.935x + 654.68; r2 = 0.9988) was used to determine the respective origins of each of the identified species, i.e., to which initial fraction they belong. As expected, the retention time agreed perfectly with the attribution of the fraction. The difference between consecutive proteoty ping-retention times was 108 (±32) s on average. This result is in full agreement with the time difference obtained between two consecutive proteotyping-retention times for Ml 1. As for Ml 1, the presence of two distinct fractions of R. pickettii was confirmed by the two corresponding SPi peaks.
DISCUSSION
The aim of this study was to develop a label-free multiplexing strategy for analysis of samples by tandem mass spectrometry. Indeed, tandem mass spectrometry-based proteotyping has an important role to play in the massive identification of isolates, especially environmental microorganisms [11] through deeper characterization of microbiomes [19], or medically-relevant but poorly characterized pathogens [20], Improving the throughput of this methodology requires robust sample preparation methods that can be applied to any sample [12], and miniaturization of the experimental load upstream of the mass spectrometry step [10], Additional improvements to throughput, such as multiplexing samples to reduce the per-sample costs of mass spectrometry are of considerable interest. For single-cell proteomics analyses, multiplex isobaric labeling has been successfully used to increase throughput [21], however, the cost of the chemical reagents may hinder its universal adoption. No similar approach to multiplexing has yet been reported for proteotyping.
The label-free multiplexing concept herein developed relies on prefractionation, where the peptide digest from each isolate to be characterized is first resolved by reverse-phase chromatography. This off-line peptide separation should ideally be performed with the same separation mode and in similar conditions to the on-line separation applied during the final nanoLC-MS/MS analysis. Here, HPLC at normal flow was used with a reversephase column with distinct characteristics to the column used with the nanoflow LC- MS/MS system. Although there were notable differences between the two chromatography runs, the elution profile of the fractions was consistent across both systems. To collect peptide fractions, a HPLC system equipped with an automatic collector was used, allowing fraction collection in 96-well plates. Numerous samples
could be fractionated with the automated procedure, but operator handling of the 96-well plate was required after fractionation. In the future, this sample preparation protocol could be fully automated to produce and store a single specific fraction for each isolate, thus minimizing lab ware-related costs. Because reverse-phase chromatography with HPLC is a highly robust and reproducible technique, this sample preparation step resulted in a robust procedure. Indeed, all of the microbial digests analyzed in this study presented similar chromatographic profiles during prefractionation, and each peptide fraction was clearly distinguishable along the acetonitrile gradient during the second chromatography run. Despite the differences in chromatographic conditions, overlap between adjacent fractions was quite low during the second chromatography run.
The label -free multiplexing concept was validated in the present study using two mixtures of peptide fractions, Ml 1 and M21 - comprising 11 and 21 fractions, respectively. These mixtures were subjected to a 60-min gradient analysis on a Q-Exactive HF tandem mass spectrometer incorporating a high-field Orbitrap analyzer. Microorganisms could be identified at the species level from acquisition data recorded over just 1 min thanks to the high number of informative MS/MS spectra acquired with this system. The methodology could perform equally well whatever the tandem mass spectrometer used, provided the system generates similarly informative MS/MS spectra. Even higher performance can be expected with more recent tandem mass spectrometers that can deliver more MS/MS spectra per unit of time. For example, next-generation instruments such as the Exploris 480 can deliver denser datasets for an equal run time, thus potentially allowing a larger number of samples to be multiplexed, or the final chromatography to be performed at higher speed. Here, the equivalent of 21 hydrophobicity -resolved microorganisms were identified in a single 60-min tandem mass spectrometry acquisition, resulting in a rate of
identification of one microorganism every 3 min of mass spectrometry. While whole-cell MALDI-TOF mass spectrometry proteotyping of isolates was reported to be quick to perform [22], the superior performances of tandem mass spectrometry proteotyping in identifying microorganism not yet referenced in a database [23] and grown in any type of media are of high interest. Here, the label -free multiplexing approach provides interesting performances and decreases mass spectrometry costs per sample.
In terms of proteotyping performances, the high potential of tandem mass spectrometry has been extensively demonstrated [11, 24, 25], The results presented here confirm the discriminative power of this methodology even in a complex situation where a high number of isolates are mixed together. Microorganisms from the same genus were easily distinguished at the species level, as shown for three genera: Pseudomonas with Pseudomonas aeruginosa and Pseudomonas putida; Deinocococcus with Deinococcus deserti and Deinococcus proteolyticus; and Bacillus with Bacillus cereus and Bacillus thuringiensis. In the latter case, the two bacterial species are generally difficult to phylogenetically discriminate as their 16S rRNA share more than 99% sequence identity.
In this proof-of-concept study, it was not tested if differences in peptide concentration that could arise if the quantities of microbial proteins subjected to proteolysis were not normalized (or at least estimated) before subjecting the resulting peptides to reversephase fractionation had a significant effect. Because UV-monitoring of the quantity of peptides in each eluted fraction can be performed automatically, it should be relatively simple to correct the volume of each fraction to be included in the final mixture. Also, the proteotyping pipeline relying on TSMs and peptides specific to the identified species,
along with the analysis of the elution profile for each species identified, as presented here, will not be affected by small variations in terms of peptide quantities.
The presence of exactly the same species in neighboring samples can lead to more complex analysis. For example, if F15, F16, and F17 originated from samples containing exactly the same R. picketti species, the Lorentzian fitting procedure might be difficult to perform, as an unexpectedly broad peak would be obtained. The possibility of such an occurrence could be included in the data-treatment procedure to ascertain whether this is the case. If a broad peak is obtained while the initial material has been correctly normalized before the tandem mass spectrometry analysis, it is recommended assigning the identified organism to the different fractions covered by the eluted peak. Moreover, dubious identification can be easily verified by analyzing a new fraction of the initial material. Typically, all dubious identifications can be multiplexed and reanalyzed in a single nanoLC-MS/MS runs.
EXAMPLE II: use of reverse phase spin columns for peptide separation MATERIAL & METHODS
1. Biological material
The five bacteria used in this study (Klebsiella aerogenes, Oceanibulbus indolifex, Marivirga Iracluosa, Sagitulla stellata and Ruegeria pomeroyi) were from a commercial source and cultivated as mentioned above. S. cerevisiae was from commercial baker’s yeast bought in supermarket that was solubilized as follows: 108 mg in 25 mL of PBS at pH 7.4 (lx) (GIBCO). Pellets cells were obtained after two centrifugations for the
removing of residual liquid medium as described above. Resulting pellets were weighed and stored at -20 °C until use.
2. Preparation of peptide digests from isolates
Peptides were prepared as previously detailed. Briefly, a specific volume of lx lithium dodecyl sulfate sample loading buffer (Thermo Fisher Scientific) prepared without dyes and supplemented with 5% beta-mercaptoethanol (v/v) was added to each cell pellet (100 pL per 1.7 mg wet biomass). After their disruption, protein lysates were stored at -20 °C until use. Peptide digests were obtained by SP3-based proteolysis performed in a 96 wellplate format as described by Hayoun et al (2020). Lysate (20 pL) was mixed with 40 pg of Sera Mag beads (4pL), formic acid (12 pL) and CH3CN to a final concentration of 85%. Successive washes with 70 % ethanol and CH3CN were made to purify trapped proteins using a Smart2 MBS (Tecan) neodymium magnetic rack. Finally, beads were resuspended in 10 pL of 50 mM NH4HCO3 supplemented with 0.01% Protease Max surfactant (Promega) and 1 pg.pL'1 trypsin gold (Promega), and incubated for 15 min at 50°C. After removal of the paramagnetic beads, the total volume was adjusted to 50 pL with aqueous 0.1% TFA and the resulting peptide solution was acidified by addition of trifluoroacetic acid (0.5% final concentration). The concentration of peptides in the samples was measured using a Pierce colorimetric Peptide Assay, as recommended (Thermo Scientific).
3. Spin-column reverse-phase fractionation of peptide digests
Peptide digests were fractionated by Cl 8 micro spin columns (Harvard Apparatus) with particle size of 10 pm and average pore size of 300 A as recommended by the supplier. First the column was washed twice with a solution containing 80 % of CH3CN and twice
with a 0.5 % trifluoroacetic acid (TFA) - H2O solution. A quantity of 40 pg of peptides was loaded onto the column and centrifugated during 1 min at 1000 x g. The eluate was load again onto the column and centrifugated. Then, a step gradient from 1% to 32% CH3CN in 0.1% formic acid was applied by incrementing 1% at each step. A fraction was collected at each step after a centrifugation step for 1 min at 1000 x g, giving a total number of 32 fractions. The quantity of peptides from each fraction was established by Nanodrop (Thermo Fisher) measurement at 214 nm. Fractions with 9%, 12%, 15%, 18%, 21% and 25% CH3CN + 0.1%formic acid were used for the mixture composition.
4. Assemblage of different mix of 6 fractions Identical quantities of peptides (100 ng) from six fractions of different hydrophobicity were pooled to obtain different mixtures (M601-M623) representing combinations of the six microorganisms. Additional mixtures (M624 and M625) were assembled with different peptide quantities. Finally, seven mixtures (M626 to M632) were assembled with less than six organisms, by repeating the same organism for multiple hydrophobic fractions. Mixtures and their exact composition are listed in the Table 3.
Table 3. List of assemblages of five fractions and species for each of their five fractions (from the least hydrophobic to the most hydrophobic).
5. Tandem mass spectrometry
Assembled peptide mixtures (M601-M632) were analyzed by nanoLC-MS/MS with an ultimate 3000 nanoLC system (Thermo Fisher Scientific) coupled to a Q-Exactive HF tandem mass spectrometer as described ([16]). Briefly, peptides were desalted on a reverse-phase PepMap 100 C18 p -precolumn. Then, they were separated on a nanoscale PepMap 100 Cl 8 nanoLC column with a flow rate of 0.3 pL.min'1 following a two-slope 20-min gradient with 4-25% B from 0 to 17 min and 25-32% B from 17 to 20 min. Mobile phase A consisted of an aqueous solution of 0.1% (v/v) formic acid in water; phase B consisted of 0.1% formic acid in 100% CH3CN. The data-dependent mode (DDA) was conducted with an MS acquisition range of 350 to 1500 m/z. The 20 most abundant precursor ions were selected for fragmentation, applying a 10-s dynamic exclusion window and a 1 .6-/77 z isolation window and 8.3e5 for the intensity threshold.
6. MS/MS data interpretation for multiplex proteotyping at the species taxonomical rank
Tandem mass spectrometry proteotyping was performed as previously described. Briefly, ion peak lists were extracted using Mascot Daemon software, version 2.6.0 (Matrix Science) generating a MGF file per sample. Then, the mgf file was split into 20- sec
acquisition time window by an in-house python script. MS/MS spectra for each of the resulting files were interpreted using Mascot version 2.6.1 (Matrix Science) against the NCBInrS database ([20]). A cascade search for the taxonomical analysis was applied as described above, for each 20 sec splitted files. Peptide sequences were mapped to taxa at the species, genus, family, order, class, phylum, and superkingdom taxonomical ranks, as previously described ([2]) ([3]), resulting in Taxon-to-Spectrum Matches (TSMs). TSMs and taxon-specific peptide sequences (spePEP) were used for the taxonomic identification. The deconvolution of TSMs and spePEPs from an assemblage of fractions was performed as previously described from these 20 sec windows. Only species identified with more than three consecutive acquisition time intervals were selected. Then, the Species Proteotyping index (SPi) was calculated by combining the two parameters (IxTSMs + 2 x SpePEP) and used for the calculation of the retention time. A non-linear regression curves of SPi as function of retention time was used for the determination of the retention time using GraphPad Prism version 6.0 (GraphPad software). Average SPi peak width was calculated by half-height section x half-width section of the modeled peaks.
RESULTS
Strategy for multiplexing microbial isolates using C18 spin columns
It is herein proposed a strategy to produce different fractions of peptides from microbial isolates using Cl 8 reverse phase spin columns eluted with solvents with different hydrophobic characteristics. This strategy is exemplified with a set of six isolates that will be analyzed all at once by tandem mass spectrometry. For each isolate, proteins are extracted and proteolyzed. The peptides from each digest are fractionated using their
hydrophobicity characteristic with Cl 8 spin column with simple application of solvents and centrifugation. Six fractions are made starting from 1 % to 32 % of acetonitrile supplemented with 0.1 % formic acid with stepwise increase of 1%. For each of the six isolates a specific fraction is selected, taking care that each fraction is different for the six isolates in order that a specific reverse phase chromatography retention time is associated with each isolate. Then, the six fractions of peptides are mixed together resulting in six pools of peptides with different hydrophobicity characteristics that will eluted sequentially from a reverse phase chromatography. The mixture of peptides is then analyzed by a single run of nanoLC-MS/MS. The interpretation of the MS/MS spectra to identify the organisms present along the gradient per windows of 20 sec is similar to our previous study. Briefly, the file containing the MS/MS spectra recorded along their retention time is sliced into 20 sec portions. Taxonomical analysis of each of the subfiles by phylopeptidomics identifies the organisms present. Then, a specific index named Species Proteotyping index (SPi), is calculated for each of the organisms identified in the whole dataset for each window of 20 sec. The SPi versus time data can be fitted to a Lorentzian model to obtain the maximum of the peak, called SPi max, corresponding to the retention time of each isolate. As indicated in Figure 6, six well-defined elution peaks are clearly distinguished for an experimental mixture that has been analyzed with a nanoLC-MS/MS run operated with a 20 min gradient. In this analytical run six organisms have been identified at the species level: Sagittula slellala. Ruegeria pomeroyi. Oceanibulbus indolifex. Saccharomyces cerevisiae. Klebsiella aerogenes and Marivirga tractuosa. Each organism is associated to a characteristic peak of elution and shows a SPi-max clearly defined and well separated from the others. Thus, a specific retention time can be assigned to each of the organisms. The order of assignation of the species to
each isolate is in the expected order. These results shows that the proof of concept of multiplexing isolates for tandem mass spectrometry using a C18 spin column is validated. Last, a correlation curve showing the SPi-max index confirmed the assignation order.
The label-free multiplexing method is robust and reproducible
To test the robustness and possible limits of the multiplexing method, the six peptide fractions that were collected systematically for the six isolates were used to create a diversity of combinations of assemblages. A total of 23 different mixtures were created with a peptide quantity for each fraction normalized at 100 ng. Each of these 23 multiplexed samples were analyzed with a 20 min nanoLC-MS/MS gradient. For all of the 23 mixtures, the 6 expected species were systematically identified and their order along the chromatography was found matching perfectly with the expected order. When analyzing the data obtained for these 23 assemblages, it was found that the SPi-max intensities associated to each fraction was relatively constant for the five first fractions, but decreased for the more hydrophobic fraction. Indeed, the integrated signal of the modeled SPi peak is in average for the 23 assemblages 2120 ± 264 TSMs x sec for the first fraction, 2187 ± 438 TSMs x sec for the second fraction, 2687 ± 559 TSMs x sec for the third fraction, 2595 ± 563 TSMs x sec for the fourth fraction, 2241 ± 716 TSMs x sec for the fifth fraction, but only 1132 ± 538 TSMs x sec for the sixth fraction. Overall, the five first fractions have an average SPi-max of 2366 ± 257 TSMs x sec. This lower level of signal for the sixth fraction is not interfering in the correctness of the identification of the species, nor the correctness of the SPi-max measurement. The low taxonomical value of the most hydrophobic fraction is due to i) lower level of MS signals due to poor ionization of hydrophobic peptides, and ii) low number of assigned MS/MS spectra
recorded because of poor fragmentation. Figure 7 shows a boxplot of the measured retention times for the six peaks and the 23 mixtures. For the six fractions, the median SPi-max is at 554, 666, 780, 894, 1010, and 1159 sec of chromatography, respectively. The range of values are [536-570], [639-682], [764-803], [880-908], [993-1026], and [1143-1170], respectively. Thus, a clear separation of SPi-max retention times for each of the six fractions and a perfect linearity of retention times as a function of the increasing fraction number are observed in Figure 7. Interestingly, the lowest and the highest values of retention for each SPi peak are very close and within less than 43 sec at maximum, showing the high reproducibility of retention times for any of the six possible fractions and whatever the microorganism origin of the fractions. In average, the difference between the maximum and the minimum values is 34 sec. Based on these observations, a reference table of retention times associated to each of the six fractions could be establish
(Table 4)
Fractions Retention time Retention time Retention time high reference value low range value range value
6
6 1158 1124 1192
Table 4 : Reference values of retention time associated to each fraction
An average of all the retention time of the 23 mixtures obtained for specific fraction was made to have a reference value. The difference between the minimum and the maximum values are used to determine the range of time of retention time for each fraction with the RT low range and the RT high range.
With the exact same experimental analytical set-up, these values can be used to directly associate a SPi-max retention time to one of the six possible fractions, thus a specific
isolate. Interestingly, it was noted that the width of each elution peak is similar regardless the organism found for the five first fractions with a width of 78 ± 8 sec in average. For the most hydrophobic fraction, the width is 45 ± 10 s as the intensity of this peak is half of the others.
The multiplex proteotyping method can manage different quantities of peptides
An important question for a routine application of the methodology is whether all the six fractions should be equal in terms of peptide quantities or if differing quantities can be afford. To test the possible limits, two mixtures of six fractions were assembled with different peptide quantities.
Mix M624 comprises six peptide fractions: 280 ng of peptides from K. aerogenes (first fraction), 60 ng of M. tractuosa (second fraction), 100 ng of 0. indolifex (third fraction), 340 ng of R. pomeroyi (fourth fraction), 180 ng of S. cerevisiae (fifth fraction), and 80 ng of S. stellata (sixth fraction). Its analysis by a 20 min gradient nanoLC-MS/MS and interpretation is indicated Figure 8 (Panel A). Six organisms could be detected and their corresponding SPi values were plotted against the elution time. Six peaks were clearly distinguished. As expected because of the differing quantities of peptides per fraction, differences in term of SPi signals are observed between organisms. For example, the SPi- max intensity associated to K. aerogenes (280 ng of peptides) is the double of that associated to M. tractuosa (60 ng). Regarding the integration of the whole SPi peak signal, the signal of the first species is 3.16 times higher than that of the second, a ratio corresponding exactly to the one expected from the injected quantities. The same was observed for R. pomeroyi with 3.17 times higher than the SPi value of other organisms excepted for K. aerogenes. Mix M625 is the assemblage of six peptide fractions: 140 ng
of peptides from S. stellata (first fraction), 120 ng of R. pomeroyi (second fraction), 40 ng of M. tractuosa (third fraction), 160 ng of S. cerevisiae (fourth fraction), 160 ng of K. aerogenes (fifth fraction), and 100 ng of O. indolifex (sixth fraction). Figure 8 Panel B shows the results of its analysis with the same previous condition. Once again, the six organisms were correctly identified and their elution peaks were clearly distinguished. The order of their respective retention times is in perfect agreement with the composition of the mixture and selected fractions. As indicated in the figure, here the SPi signal is relatively more similar between the peaks but a slight drop of SPi values were noted for M. tractuosa (third fraction) and O. indolifex (sixth fraction) as expected as their quantities are lower than the others with 40 and 100 ng, respectively.
The retention times associated to each fraction for the two mixtures are 562 ± 7 sec; 673 ± 7 sec; 777 ± 0 sec; 884 ± 4 sec; 1016 ± 5 sec; and 1163 ± 1 sec for the least to the most hydrophobic fractions. These values are perfectly matching the elution windows defined for the mixtures made with fractions with equal amounts of peptides which are indicated in Table 4. Thus, it was observed in these two examples, that the difference of quantities of peptides per fraction does not affect the modeled retention time of the identified organism associated to each fraction.
The method is robust even if a same organism is present in several fractions in the mix
Another important question for a routine application of the methodology is whether the methodology is robust and delivers the expected identification when in the series of six samples to multiplex, several are the same species. Seven mixtures of six fractions were assembled to represent several scenarios. First, it was analyzed whether the presence of
the same organism in several distant fractions is disturbing the identification procedure. The analyses of the four first mixtures are presented in Figure 9. M629 comprises only 3 organisms: R. pomeroyi, S. cerevisiae and K. aerogenes. Each of the three organisms contributed to two fractions. As shown in Figure 9 (Panel A), two SPi peaks are clearly distinguished for each of these three organisms. The peaks at 556 and 885 sec corresponds to R. pomeroyi. At 788 and 1155 sec the peaks indicate the presence of S. cerevisiae and the peaks at 678 and 998 sec are from ", aerogenes. As previously observed, the last peak (F6) associated to S. cerevisiae shows a decreasing of SPi intensity. Indeed, SPi intensity for the last peak (F6) is at 119 and the average of SPi intensity for last fraction is 202 ± 69. M630 contains peptides corresponding to four organisms: K. aerogenes, O. indolifex, R. pomeroyi and S. stellata (Figure 9, Panel B). Two organisms have two distinct peaks: O. indolifex at 790 and 1017 sec, and R. pomeroyi at 659 and 1162 sec. The remaining two organisms are identified by a single peak observed at 549 sec for S. stellata and 901 sec for K. aerogenes. Fractions from four organisms were assembled for Mix M632: K. aerogenes, S. cerevisiae, R. pomeroyi and S. stellata. As shown in Figure 9 (Panel C), two organisms are identified with two peaks: R. pomeroyi at 656 and 1164 sec, and K. aerogenes at 795 and 1002 sec. For the remaining two organisms, a single peak was observed at 543 sec for S. cerevisiae and 901 sec for S. stellata. Mix M632 corresponds to peptides from four organisms: R pomeroyi, K. aerogenes, S. cerevisiae, and S. stellata. Figure 9 (Panel D) shows that two peaks can be distinguished for two organisms: S. cerevisiae at 786 and 998 sec and K. aerogenes at 679 and 1159 sec. For the remaining two organisms, a single peak was observed at 557 sec for R. pomeroyi and 897 sec for S. stellata. For the four mixtures, the identified organisms and the order of the retention time of the corresponding fractions are in perfect agreement with the composition of this mix.
For these four assemblages, the average retention time for each fraction is: 551 ± 6 sec, 668 ± 12 sec, 789 ± 4 sec, 896 ± 7 sec, 1004 ± 9 sec and 1160 ± 4 sec for the least to the most hydrophobic fractions. These retention times are in perfect agreement with the reference values mentioned in Table 4 with a CV <1%, showing again the robustness of the method.
Another more complex scenario was tested: three mixtures containing a same organism with two or three adjacent fractions were assembled and analyzed. The M626 assemblage comprises four organisms: 0. indolifex, S. cerevisiae, R. pomeroyi and K. aerogenes. Figure 10 (panel A) shows the SPi values of each selected organisms plotted as a function of time, with a point every 20 sec acquisition time window. Four peaks are observed, the first peak being wider than the other peaks. Indeed, its width is 156 sec while the average of the other peaks is 78 ± 8 sec in average. The same observation can be done for the mix M627 represented on Figure 10 (panel B), where three organisms were identified. A single peak was associated to K. aerogenes, two peaks were attributed to S. cerevisiae, and a very wide peak was assigned to R. pomeroyi with a width of 128 sec. A diminution of the SPi intensity with an amplitude of 82 TSMs for the last peak (F6) associated to S. cerevisiae is observed. The average in term of SPi intensity for the last fraction is 202 ± 69 TSMs. It may be due to the previous larger peak associated to R. pomeroyi. M628 assemblage is made of peptide fractions from five organisms: K. aerogenes, S. stellata, R. pomeroyi, O. indolifex and S. cerevisiae are detected at 566 sec, 677 sec, 775 sec, 893 sec and 1014 sec, respectively, as indicated in Figure 10 (panel C). For all the peaks, the width is in agreement with a single fraction. Only five peaks are observed but the more hydrophobic peak corresponding to S. cerevisiae can be assigned to the fifth and sixth fractions taking into account the width and the start and end values of the peak. In these
last worst-case scenarios, it was highlighted that the width of the peak gives directly the information whether the same organism is present in a single fraction only or in several consecutive fractions in the assemblage. In conclusion, the six retention times expected and the peak width if less than six SPi-max are measured allow the final assignation of the organisms to each of the six initial samples.
DISCUSSION
Isolating and identifying quickly microorganisms is pivotal for diagnostic purposes, searching for uncharacterized branches of the Tree of Life, or screening novel catalysts of interest in biotechnology. Tandem mass spectrometry-based proteotyping is an impressively accurate methodology, but improving its throughput is key to popularizing this approach. The aim of this second example was to make cost-effective an innovative label-free multiplex proteotyping protocol that we previously developed. While the first description of the methodology (example I) was relying on fractionation of the peptidome obtained from each isolate by reverse phase chromatography using HPLC, here it is proposed a simplified alternative based on Cl 8 spin columns. The robustness of the methodology was also evaluated by exploring different worst-case scenarios.
A selected peptide fraction from each isolate is differentiated from the fractions arising from the other isolates by its hydrophobic characteristics. Because HPLC requires equipment and specific expertise, it is proposed here to fraction the peptidomes with Cl 8 spin columns that would require only a centrifuge. In principle, this approach can be fully automatized and is cost-effective. Here, six fractions differing in hydrophobicity are used for creating a mixture representative of six isolates and analysed into a single 20 min nanoLC-MS/MS run. It was noted that each of the resulting peaks of elution of species-
specific peptides are clearly identified. Without any optimisation and automation, the fractionation step took about 1 h for treating in parallel six isolates. The current yield of identification is around 3 min of tandem mass spectrometry per isolate. This results in a rather competitive methodology in terms of costs.
The best practice to establish a novel analytical method is to test its robustness whatever the conditions. Here, the robustness of the methodology was tested by using 23 sets of six isolates with the same organisms but with different combinations. Fractions used for the artificial combinations were normalized with a peptide quantity of 100 ng per added fraction, resulting into mixtures of 600 ng of peptides injected per nanoLC-MS/MS. The peptides of each of the six microorganisms present in a mixture can delineate a specific taxon-retention time on the final reverse-phase chromatography. The dispersion of each of these six retention times observed for this set of 23 mixtures is minimal and shows a very good reproducibility. If similar conditions of nanoLC-MS/MS are used, reference values for the six expected species-retention times can be defined. Whether difference in peptide quantities in the six fractions could influence the results was also tested in this work. Introduction of fractions with different quantities in the range of 60-280 ng did not introduce any noticeable bias. Nevertheless, the normalization with a peptide quantity of 100 ng in each fraction is recommended, especially for the most hydrophobic fraction. Indeed the total number of SPi for the sixth fraction is always lower compared to the results from the five first fractions.
Finally, another potential difficulty was considered: the presence of the same organism in multiple isolates in the same multiplex set of six samples. Indeed, many culturomics project have to deal with dereplication of isolates to avoid waste of time in characterizing the same biological object several times ([34]). The same appears true for clinical
diagnostic where several samples commonly contain the same type of pathogens. Two types of mixtures were created: four sets with at least an organism repeated in two different fractions and three with the same organism repeated but in neighboring fractions. Once again, this worst-case scenario did not affect the identification success and the correct assignation to the expected initial samples. In this case, the width of peaks assigned to a given species, i.e. its retention time range, and its average retention time give information on the presence of the same organisms in neighboring fractions. As an example, it was noted that in the case where three neighboring fractions contain the same organism, the SPi-max value for this organism corresponds exactly to the fraction in the middle, but the width of the peak is enlarged. Also, if the number of organisms is inferior to the number of introduced fractions, i.e. six here, a simple ratio of the width of peaks and the average of the expected width can confirm the presence of the same organism in neighboring fractions. Thus, these three parameters should be taken into account to determine the number of fractions containing the same organism. For mixtures composed with the same organism in different but not successive fractions two peaks of elution for the organism can be easily distinguished with width in full agreement with the average value. In such case, two SPi-max are obtained.
Interestingly, the method makes no use of pre-recorded mass spectra databases such as with whole-cell MALDI-TOF mass spectrometry ([1]), but is rather based on the whole known diversity through the current database of annotated full genome sequences. It is thus applicable to any of the organisms already genome sequenced and their relatives, as well as uncharacterized branches of the tree of life because partial conserved information can be always retrieved. Regarding the range of application of the methodology, it was tested whether bacteria or yeasts can be easily identified. Specifically, S. cerevisiae was
introduced in several scenarios to prove that the method applies equally well on Bacteria and Eukaryota. S. cerevisiae was always identified. The extraction of peptides from S. cerevisiae or from the bacteria is identical as based on an optimized protocol previously developed for complex samples including both types of material ([12]). However, it was observed less signal in terms of abundance (TSMs) and taxon-specificity (SpePEPs) compared to the signals for bacteria while the quantity of peptides was identical. This comes from the dynamic range of the proteomes that differ between Eukaryota and Bacteria on the one hand, and the number of genomes for this specific branch of the tree of life that may lower the Saccharomyces cere\'isiae- c\ \c peptides on the other hand.
In conclusion, the methodology presented here enables robust and cost-effective multiplexing of samples for microbial identification by tandem mass spectrometry proteotyping. The results show how simple the methodology is because the fractionation step before the analysis can be performed with Cl 8 spin columns. Based on this affordable technology, the fractionation of peptidomes can be fully automatized to reduce operations and time to result. The current protocol allows the identification of 6 organisms in a single 20-min nanoLC-MS/MS analytical run, but there is room for improvement as the latest generation of tandem mass spectrometers have several fold more capacities than the mass spectrometer used here.
This protocol paves the way for proteotyping of isolates by tandem mass spectrometry in clinical settings for the diagnosis of difficult organisms that are not yet in whole-cell MALDI-TOF databases or the culturomics pipeline dedicated to the discovery of new microorganisms.
BIBLIOGRAPHIC REFERENCES
(1) Suarez, S. Ribosomal Proteins as Biomarkers for Bacterial Identification by Mass Spectrometry in the Clinical Microbiology Laboratory. J. Microbiol. Methods 2013, 7.
(2) Hayoun, K.; Pible, O.; Petit, P.; Allain, F.; Jouffret, V.; Culotta, K.; Rivasseau, C.; Armengaud, J.; Alpha-Bazin, B. Proteotyping Environmental Microorganisms by Phylopeptidomics: Case Study Screening Water from a Radioactive Material Storage Pool. Microorganisms 2020, 8 (10), 1525.
(3) Pible, O.; Allain, F.; Jouffret, V.; Culotta, K.; Miotello, G.; Armengaud, J. Estimating Relative Biomasses of Organisms in Microbiota Using “Phylopeptidomics.” Microbiome 2020, 8 (1), 30.
(4) Lagier, J.-C.; Armougom, F.; Million, M.; Hugon, P.; Pagnier, I.; Robert, C.; Bittar, F.; Foumous, G.; Gimenez, G.; Maraninchi, M.; Trape, J.-F.; Koonin, E. V.; La Scola, B.; Raoult, D. Microbial Culturomics: Paradigm Shift in the Human Gut Microbiome Study. Clin. Microbiol. Infect. 2012, 18 (12), 1185-1193.
(5) Mesuere, B.; Devreese, B.; Debyser, G.; Aerts, M.; Vandamme, P.; Dawyndt, P. Unipept: Tryptic Peptide-Based Biodiversity Analysis of Metaproteome Samples. J. Proteome Res. 2012, 11 (12), 5773-5780.
(6) Boulund, F.; Karlsson, R.; Gonzales-Siles, L.; Johnning, A.; Karami, N.; AL- Bayati, O.; Ahren, C.; Moore, E. R. B.; Kristiansson, E. Typing and Characterization of Bacteria Using Bottom-up Tandem Mass Spectrometry Proteomics. Mol. Cell. Proteomics 2017 , 76 (6), 1052-1063.
(7) Mooradian, A. D.; van der Post, S.; Naegle, K. M.; Held, J. M. ProteoClade: A Taxonomic Toolkit for Multi-Species and Metaproteomic Analysis. PLOS Comput. Biol. 2020, 76 (3), el007741.
(8) Kuhring, M.; Doellinger, J.; Nitsche, A.; Muth, T.; Renard, B. Y. Taxlt: An Iterative Computational Pipeline for Untargeted Strain-Level Identification Using MS/MS Spectra from Pathogenic Single-Organism Samples. J. Proteome Res. 2020, 19 (6), 2501-2510.
(9) Petit, P. C. M.; Pible, O.; Eesbeeck, V. V.; Alban, C.; Steinmetz, G.; Mysara, M.; Monsieurs, P.; Armengaud, J.; Rivasseau, C. Direct Meta-Analyses Reveal Unexpected Microbial Life in the Highly Radioactive Water of an Operating Nuclear Reactor Core. Microorganisms 2020, 8 (12), 1857.
(10) Hayoun, K.; Gaillard, J.-C.; Pible, O.; Alpha-Bazin, B.; Armengaud, J. High- Throughput Proteotyping of Bacterial Isolates by Double Barrel Chromatography- Tandem Mass Spectrometry Based on Microplate Paramagnetic Beads and Phylopeptidomics. J. Proteomics 2020, 226, 103887.
(11) Lozano, C.; Kielbasa, M.; Gaillard, J.-C.; Miotello, G.; Pible, O.; Armengaud, J. Identification and Characterization of Marine Microorganisms by Tandem Mass Spectrometry Proteotyping. Microorganisms 2022, 10 (4), 719.
(12) Hayoun, K.; Gouveia, D.; Grenga, L.; Pible, O.; Armengaud, J.; Alpha-Bazin, B. Evaluation of Sample Preparation Methods for Fast Proteotyping of Microorganisms by Tandem Mass Spectrometry. Front. Microbiol. 2019, 10, 1985.
(13) Vujicic-Zagar, A.; Dulermo, R.; Le Gorrec, M.; Vannier, F.; Servant, P.; Sommer,
S.; de Groot, A.; Serre, L. Crystal Structure of the IrrE Protein, a Central Regulator of DNA Damage Repair in Deinococcaceae. J. Mol. Biol. 2009, 386 (3), 704-716.
(14) Reasoner, D. J.; Geldreich, E. E. A New Medium for the Enumeration and Subculture of Bacteria from Potable Water. Appl. Environ. Microbiol. 1985, 49 (1), 1-7.
(15) Hughes, C. S.; Foehr, S.; Garfield, D. A.; Furlong, E. E.; Steinmetz, L. M.; Krijgsveld, J. Ultrasensitive Proteome Analysis Using Paramagnetic Bead Technology. Mol. Syst. Biol. 2014, 10 (10), 757.
(16) Trapp, J.; Almunia, C.; Gaillard, J.-C.; Pible, O.; Chaumot, A.; Geffard, O.; Armengaud, J. Proteogenomic Insights into the Core-Proteome of Female Reproductive Tissues from Crustacean Amphipods. J. Proteomics 2016, 135, 51-61.
(17) Hirtz, C.; Mannaa, A. M.; Moulis, E.; Pible, O.; O’Flynn, R.; Armengaud, J.; Jouffret, V.; Lemaistre, C.; Dominici, G.; Martinez, A. Y .; Dunyach-Remy, C.; Tiers, L.; Lavigne, J.-P.; Tramini, P.; Goldsmith, M.; Lehmann, S.; Deville de Periere, D.; Vialaret, J. Deciphering Black Extrinsic Tooth Stain Composition in Children Using Metaproteomics. ACS Omega 2022, 7 (10), 8258-8267.
(18) Grenga, L.; Pible, O.; Miotello, G.; Culotta, K.; Ruat, S.; Roncato, M.; Gas, F.; Bellanger, L.; Claret, P.; Dunyach-Remy, C.; Laureillard, D.; Sotto, A.; Lavigne, J.; Armengaud, J. Taxonomical and Functional Changes in COVID -19 Faecal Microbiome Could Be Related to SARS-CoV -2 Faecal Load. Environ. Microbiol. 2022, 1462- 2920.16028.
(19) Van Den Bossche, T.; Arntzen, M. 0.; Becher, D.; Benndorf, D.; Eijsink, V. G. H.; Henry, C.; Jagtap, P. D.; Jehmlich, N.; Juste, C.; Kunath, B. J.; Mesuere, B.; Muth,
T.; Pope, P. B.; Seifert, J.; Tanca, A.; Uzzau, S.; Wilmes, P.; Hettich, R. L.; Armengaud, J. The Metaproteomics Initiative: A Coordinated Approach for Propelling the Functional Characterization of Microbiomes. Microbiome 2021, 9 (1), 243.
(20) Grenga, L.; Pible, O.; Armengaud, J. Pathogen Proteotyping: A Rapidly Developing Application of Mass Spectrometry to Address Clinical Concerns. Clin. Mass Spectrom. 2019, 14, 9-17.
(21) Dou, M.; Clair, G.; Tsai, C.-F.; Xu, K.; Chrisler, W. B.; Sontag, R. L.; Zhao, R.; Moore, R. J.; Liu, T.; Pasa-Tolic, L.; Smith, R. D.; Shi, T.; Adkins, J. N.; Qian, W.-J.;
Kelly, R. T.; Ansong, C.; Zhu, Y. High-Throughput Single Cell Proteomics Enabled by Multiplex Isobaric Labeling in a Nanodroplet Sample Preparation Platform. Anal. Chem. 2019, 91 (20), 13119-13127.
(22) Bizzini, A. Matrix- Assisted Laser Desorption Ionization Time-of-Flight Mass Spectrometry, a Revolution in Clinical Microbial Identification. 6.
(23) Armengaud, J. Metaproteomics to Understand How Microbiota Function: The Crystal Ball Predicts a Promising Future. Environ. Microbiol. 2022, 1462-2920.16238.
(24) Heyer, R.; Benndorf, D.; Kohrs, F.; De Vrieze, J.; Boon, N.; Hoffmann, M.; Rapp, E.; Schluter, A.; Sczyrba, A.; Reichl, U. Proteotyping of Biogas Plant Microbiomes Separates Biogas Plants According to Process Temperature and Reactor Type. BiotechnoL Biofuels 2016, 9 (1), 155.
(25) Karlsson, R.; Gonzales-Siles, L.; Gomila, M.; Busquets, A.; Salva-Serra, F.; Jaen- Luchoro, D.; Jakobsson, H. E.; Karlsson, A.; Boulund, F.; Kristiansson, E.; Moore, E. R. B. Proteotyping Bacteria: Characterization, Differentiation and Identification of Pneumococcus and Other Species within the Mitis Group of the Genus Streptococcus by Tandem Mass Spectrometry Proteomics. PLOS ONE 2018, 13 (12), e0208804.
(27) Manadas, B., Mendes, V.M., English, J., Dunn, M.J., 2010. Peptide fractionation in proteomics approaches. Expert Rev. Proteomics 7, 655-663.
(28) Deng, W ., Sha, J., Plath, K., Wohlschlegel, J. A., 2021. Carboxylate-Modified Magnetic Bead (CMMB)-Based Isopropanol Gradient Peptide Fractionation (CIF) Enables Rapid and Robust Off-Line Peptide Mixture Fractionation in Bottom-Up Proteomics. Mol. Cell. Proteomics 20, 100039.
(29) Herraiz, T.; Casal, V. Evaluation of Solid-Phase Extraction Procedures in Peptide Analysis. J. Chromatogr. A 1995, 708 (2), 209-221
(30) Rappsilber, J., Mann, M., Ishihama, Y., 2007. Protocol for micro-purification, enrichment, pre-fractionation and storage of peptides for proteomics using StageTips. Nat. Protoc. 2, 1896-1906.
(31) Liu X, Rossio V, Paulo JA. Spin column-based peptide fractionation alternatives for streamlined tandem mass tag (SL-TMT) sample processing. J Proteomics. 2023, 276: 104839
(32) Bache, N., Geyer, P.E., Bekker-Jensen, D.B., Hoerning, O., Falkenby, L., Treit, P.V., Doll, S., Paron, I., Muller, J.B., Meier, F., Olsen, J.V., Vorm, O., Mann, M., 2018. A Novel LC System Embeds Analytes in Pre-formed Gradients for Rapid, Ultra-robust Proteomics. Mol. Cell. Proteomics 17, 2284-2296.
(33) Hosp, F., Scheltema, R.A., Eberl, H.C., Kulak, N.A., Keilhauer, E.C., Mayr, K., Mann, M., 2015. A Double-Barrel Liquid Chromatography-Tandem Mass Spectrometry
(LC-MS/MS) System to Quantify 96 Interactomes per Day*. Mol. Cell. Proteomics 14, 2030-2041.
(34) Kapinusova, G., Jani, K., Smrhova, T., Pajer, P., Jarosova, I., Suman, J., Strejcek, M., Uhlik, O., 2022. Culturomics of Bacteria from Radon- Saturated Water of the World’s Oldest Radium Mine. Microbiol. Spectr. 10, e01995-22.
Claims
1. Method for identifying and/or quantifying the organisms present in each of X biological samples, wherein X is an integer superior to 2, said method comprising the following successive steps : a) Providing one peptide sample PSi from each of said X biological samples, wherein i is an integer comprised between 1 and X, b) Separating said peptide samples PSi into at least N fractions Fj according to their hydrophobicity level, each fraction Fj being characterized by its hydrophobicity level Hj and/or another elution characteristic Ej, wherein N is an integer superior or equal to X, and wherein j is an integer comprised between 1 and N, c) Pooling at least one fraction Fj from each peptide sample PSi,, each fraction Fj being characterized by its hydrophobicity level Hj and/or another elution characteristic Ej so as to generate a final sample wherein all the fractions Fj added in the final sample have a different hydrophobicity level Hj and/or elution characteristic Ej,. d) Injecting the final sample of step c) into a chromatographic system coupled to a tandem mass spectrometer for resolving the peptides by their hydrophobicity, establishing preferably the charge z and the mass spectra of the eluted peptides, e) Analysing the MS/MS spectra obtained in step d) by comparing them with a database of known peptide sequence data and matching each peptide sequence of said MS/MS spectrum to one or more organisms present in said database, f) Identification of the organisms present in each of the X biological samples, by classifying the organisms identified in step e) to each initial biological sample, by using the specific elution characteristics at which the organisms were identified.
2. The method of claim 1, wherein the separation in step b) is performed by hydrophobicity filtering.
3. The method of claim 1 or 2, wherein the separation in step b) is performed by reverse-phase chromatography, preferably by reverse-phase HPLC or by reverse phase spin column.
4. The method of any of claims 1 to 3, wherein the separation in step b) is performed by reverse-phase HPLC with an elution time of about 30 seconds.
5. The method of any of claims 1 to 4, wherein, in the pooling step c), the final sample contains:
• the fraction Fi of the peptide sample PSi having an hydrophobicity level Hi and an elution characteristic Ei,
• the fraction F2 of the peptide sample PS2 having an hydrophobicity level H2 and an elution characteristic E2,
• the fraction F3 of the peptide sample PS 3 having an hydrophobicity level H3 and an elution characteristic E3,
• (etc...)
• the fraction Fj of the peptide sample PSi having an hydrophobicity level Hj and an elution characteristic Ej, wherein Hi H2 H3 Hj and wherein Ei E2 E3 ^ ... ^ Ej.
6. The method of any of claims 1 to 5, wherein more than one fraction Fj of the same peptide sample PSi is added in the final sample of step c), provided that said fractions Fj have different hydrophobicity levels Hj and/or different elution characteristics Ej.
7. The method of any of claims 1 to 6, wherein the step d) is performed by using a HPLC system coupled to tandem mass spectrometer, preferably a nanoLC device coupled to a MS/MS tandem spectrometer.
8. The method of any of claims 1 to 7, wherein the step d) is performed by subjecting the final sample of step c) to a 60-min HPLC gradient coupled to a tandem MS/MS analysis.
9. The method of any of claims 1 to 8, wherein the resulting MS/MS spectra dataset in step d) is divided into interpretable sub-datasets corresponding to short windows of elution times, for example of less than 2 minutes, typically of about 1 minute or of about 30 seconds.
10. The method of claim 9, comprising calculating, for each given acquisition window, the index:
SPi = I x TSMs + k x spePEP for each identified taxon and each fraction, wherein “TSMs” is the number of TSMs in said acquisition window, “spePEP” is the number of taxon-specific peptide sequences in said acquisition window, and k is a number comprised between 1.5 and 2.5, in order to calculate the elution characteristic corresponding to the maximum Spi for each organism.
11. The method of any of claims 1 to 10, wherein a polynomial regression model is used to determine the exact elution characteristic of each of the identified organism, thus their respective sample origins.
12. The method of claim 10, wherein only organisms with SPi values > 0 over more than three consecutive acquisition time intervals are retained.
13. The method of any of claims 1 to 12, wherein said organism is a virus, a bacterium, an archaeon, a yeast, a fungus, an algae, a plant, an animal, a parasite, or any other organism, pathogenic or not, alive or not, taxonomically characterized or not.
14. The method of any of claims 1 to 13, comprising an additional step h) of calculating the relative or absolute amount of the detected organisms in each biological sample.
15. The method of any of claims 1 to 14, wherein the quantities of proteins subject to proteolysis have been normalised before step a) or if not, wherein the quantity of peptides in each eluted fraction of step b) is calculated so that the fractions mixed in the final sample contain the same equivalent quantities of peptides.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23305557 | 2023-04-13 | ||
| PCT/EP2024/060182 WO2024213796A1 (en) | 2023-04-13 | 2024-04-15 | Label-free multiplex proteotyping of microbial isolates |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4695611A1 true EP4695611A1 (en) | 2026-02-18 |
Family
ID=86329220
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24717246.3A Pending EP4695611A1 (en) | 2023-04-13 | 2024-04-15 | Label-free multiplex proteotyping of microbial isolates |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4695611A1 (en) |
| WO (1) | WO2024213796A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6558946B1 (en) | 2000-08-29 | 2003-05-06 | The United States Of America As Represented By The Secretary Of The Army | Automated sample processing for identification of microorganisms and proteins |
| ES2683032T3 (en) | 2010-12-17 | 2018-09-24 | Biomerieux, Inc | Methods of isolation, accumulation, characterization and / or identification of microorganisms using a sample filtration and transfer device, and said device |
| FR3006055B1 (en) | 2013-05-24 | 2020-08-14 | Commissariat Energie Atomique | PROCESS FOR CHARACTERIZING BY TANDEM MASS SPECTROMETRY A BIOLOGICAL SAMPLE |
| EP2835751A1 (en) | 2013-08-06 | 2015-02-11 | Commissariat A L'energie Atomique Et Aux Energies Alternatives | Method of deconvolution of mixed molecular information in a complex sample to identify organism(s) |
-
2024
- 2024-04-15 EP EP24717246.3A patent/EP4695611A1/en active Pending
- 2024-04-15 WO PCT/EP2024/060182 patent/WO2024213796A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024213796A1 (en) | 2024-10-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP6668415B2 (en) | Apparatus and method for microbiological analysis | |
| Roux-Dalvai et al. | Fast and accurate bacterial species identification in urine specimens using LC-MS/MS mass spectrometry and machine learning | |
| Abiraami et al. | Soil metaproteomics as a tool for monitoring functional microbial communities: promises and challenges | |
| JP5754742B2 (en) | Method for characterizing at least one microorganism by mass spectrometry | |
| Picotti et al. | The implications of proteolytic background for shotgun proteomics | |
| Keilhauer et al. | Accurate protein complex retrieval by affinity enrichment mass spectrometry (AE-MS) rather than affinity purification mass spectrometry (AP-MS) | |
| Lay Jr | MALDI-TOF mass spectrometry and bacterial taxonomy | |
| Arsène-Ploetze et al. | Proteomic tools to decipher microbial community structure and functioning | |
| Everley et al. | Liquid chromatography/mass spectrometry characterization of Escherichia coli and Shigella species | |
| Mappa | Alpha-Bazin, B.; Pible, O.; Armengaud, J. Mix24X, a lab-assembled reference to evaluate interpretation procedures for tandem mass spectrometry proteotyping of complex samples | |
| Lay Jr et al. | Rapid identification of bacteria based on spectral patterns using MALDI-TOFMS | |
| Rajoria et al. | Elucidation of protein biomarkers for verification of selected biological warfare agents using tandem mass spectrometry | |
| Mott et al. | Comparison of MALDI-TOF/MS and LC-QTOF/MS methods for the identification of enteric bacteria | |
| Velichko et al. | Classification and identification tasks in microbiology: mass spectrometric methods coming to the aid | |
| EP4695611A1 (en) | Label-free multiplex proteotyping of microbial isolates | |
| US9995751B2 (en) | Method for detecting at least one mechanism of resistance to glycopeptides by mass spectrometry | |
| Al-Shahib et al. | Coherent pipeline for biomarker discovery using mass spectrometry and bioinformatics | |
| Dupas et al. | Importance of taking Single Amino Acid Variant and accessory proteome variability into account in Data Independent Acquisition Proteomics: illustrated with Legionella pneumophila analysis | |
| Ravi Kumar et al. | ChipFilter: Microfluidic Based Comprehensive Sample Preparation Methodology for Metaproteomics | |
| Chabas et al. | A Simplified Label-Free Method for Proteotyping Sets of Six Isolates in a Single Liquid Chromatography-High-Resolution Tandem Mass Spectrometry Analysis | |
| RK et al. | ChipFilter: Microfluidic-Based Comprehensive Sample Preparation Methodology for Microbial Consortia. | |
| Drissner et al. | Rapid profiling of human pathogenic bacteria and antibiotic resistance employing specific tryptic peptides as biomarkers | |
| US20170275667A1 (en) | Method for quantifying at least one microorganism group via mass spectrometry | |
| McFarland et al. | Bacterial Identification at the Serovar Level by Top-Down Mass Spectrometry | |
| Everley | Characterization of Microorganisms of Interest to Homeland Security and Public Health Utilizing Liquid Chromatography/Mass Spectrometry |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251107 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |