WO2025199014A1 - Mass spectrometry-based top-down methods of characterizing proteins in protein corona - Google Patents
Mass spectrometry-based top-down methods of characterizing proteins in protein coronaInfo
- Publication number
- WO2025199014A1 WO2025199014A1 PCT/US2025/020200 US2025020200W WO2025199014A1 WO 2025199014 A1 WO2025199014 A1 WO 2025199014A1 US 2025020200 W US2025020200 W US 2025020200W WO 2025199014 A1 WO2025199014 A1 WO 2025199014A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- protein
- corona
- sample
- proteins
- disease
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
- G01N33/6803—General methods of protein analysis not limited to specific proteins or families of proteins
- G01N33/6848—Methods of protein analysis involving mass spectrometry
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N27/00—Investigating or analysing materials by the use of electric, electrochemical, or magnetic means
- G01N27/26—Investigating or analysing materials by the use of electric, electrochemical, or magnetic means by investigating electrochemical variables; by using electrolysis or electrophoresis
- G01N27/416—Systems
- G01N27/447—Systems using electrophoresis
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N30/00—Investigating or analysing materials by separation into components using adsorption, absorption or similar phenomena or using ion-exchange, e.g. chromatography or field flow fractionation
- G01N30/02—Column chromatography
- G01N30/88—Integrated analysis systems specially adapted therefor, not covered by a single one of the groups G01N30/04 - G01N30/86
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/68—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
- G01N33/6803—General methods of protein analysis not limited to specific proteins or families of proteins
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N30/00—Investigating or analysing materials by separation into components using adsorption, absorption or similar phenomena or using ion-exchange, e.g. chromatography or field flow fractionation
- G01N30/02—Column chromatography
- G01N30/88—Integrated analysis systems specially adapted therefor, not covered by a single one of the groups G01N30/04 - G01N30/86
- G01N2030/8809—Integrated analysis systems specially adapted therefor, not covered by a single one of the groups G01N30/04 - G01N30/86 analysis specially adapted for the sample
- G01N2030/8813—Integrated analysis systems specially adapted therefor, not covered by a single one of the groups G01N30/04 - G01N30/86 analysis specially adapted for the sample biological materials
- G01N2030/8831—Integrated analysis systems specially adapted therefor, not covered by a single one of the groups G01N30/04 - G01N30/86 analysis specially adapted for the sample biological materials involving peptides or proteins
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N30/00—Investigating or analysing materials by separation into components using adsorption, absorption or similar phenomena or using ion-exchange, e.g. chromatography or field flow fractionation
- G01N30/02—Column chromatography
- G01N30/62—Detectors specially adapted therefor
- G01N30/72—Mass spectrometers
Definitions
- This disclosure generally relates to a method of characterizing proteins in a biological sample by adding protein-binding agents, such as nanoparticles (NPs), to the sample to generate protein coronas, separating the protein-binding agents from the protein coronas, and characterizing the proteins from the protein coronas using Capillary Electrophoresis (CE) and Mass Spectrometry (MS).
- protein-binding agents such as nanoparticles (NPs)
- CE Capillary Electrophoresis
- MS Mass Spectrometry
- Nanomedicine applies nanotechnology concept to medicine, i.e., employing biocompatible nanoparticles (NPs) for controlled and/or targeted delivery of therapeutic (bio)molecules to desired tissues/organs, imaging, and disease diagnosis.
- NPs nanoparticles
- the overall efficacy of nanomedicine is strongly impacted by protein/biomolecular corona, i.e., the composition and decoration of various types of biomolecules (e.g., mostly proteins) that bind to the surface of NPs after they are exposed to biological fluids (e.g., human plasma). It has been well documented that the composition and decoration of participated proteins in protein corona determines the biological fate and pharmacokinetics of NPs.
- protein corona has also been recognized as a useful analytical technique to discover new protein biomarkers of diseases because it can reduce the complexity of biological fluids (e.g., plasma), which enables easier detection and identification of biomarkers.
- biological fluids e.g., plasma
- MS-based bottom-up proteomics has been widely recognized as an efficient way for the characterization of protein corona, providing the identification of gene products in the protein corona.
- MS-based BUP fails to identify exact forms of protein molecules (i.e., proteo forms) in the protein corona because of the “peptide- to-protein” inference problem.
- proteoforms from the same gene due to sequence variations and post-translational modifications (PTMs) could have divergent biological functions and proteoforms are important for modulating disease progression.
- PTMs post-translational modifications
- the present disclosure provides a method for characterizing proteins (including one or more of the protein’ s proteoforms) in a biological sample by adding one or more protein-binding agents to the biological sample, where the protein-binding agents have a surface capable of binding protein, and allowing a protein corona to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one or more protein binding agents.
- the method also includes separating the complex from the biological sample to generate a complex sample including the complex, separating the one or more proteinbinding agents from the protein corona in the complex sample to generate a protein corona sample, and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
- CE Capillary Electrophoresis
- MS/MS Tandem Mass Spectrometry
- separating the complex from the biological sample may include one or more rounds of washing and centrifuging to remove a supernatant including the biological sample and any proteins not bound to or that have a low affinity for the one or more proteinbinding agents.
- the method may include separating one or more proteinbinding agents from the protein corona in the complex sample by eluting the protein corona from the one or more protein-binding agents.
- the elution buffer may contain a detergent, such as SDS, and/or an ionic liquid with 4-12 carbons in the hydrocarbon chain.
- separating the one or more protein-binding agents from the protein corona may include dissolving the one or more protein-binding agents.
- the method includes buffer-exchanging the buffer in the protein corona sample to a mass spectrometry (MS)-compatible buffer prior to characterizing the proteins in the protein corona sample.
- the CE may be Capillary Zone Electrophoresis (CZE) or Capillary Isoelectric Focusing (cIEF).
- the protein-binding agents may be nanoparticles.
- the present disclosure provides a method for detecting one or more biomarkers or a pattern of one or more biomarkers associated with a disease following the method of characterizing proteins described herein, where the biological sample is obtained from the subject.
- the present disclosure also provides a method of diagnosing a disease in a subject, by following the characterization method described herein.
- the disease may be cancer or a neurological disease.
- compositions including the protein corona sample produced by the method described herein, which may include the protein corona and substantially no protein-binding agent.
- FIG. 1 is an example schematic of the MS-based top-down proteomics (TDP) workflow for protein corona using polystyrene nanoparticles (PSNPs), a human plasma sample, and capillary zone electrophoresis (CZE)-tandem MS (MS/MS) (see, EXAMPLE 1).
- TDP top-down proteomics
- FIG. 2 includes transmission electron microscopy (TEM) images of bare NPs, NPs after protein binding, and NPs after protein elution with 0.4% SDS (see, EXAMPLE 1).
- TEM transmission electron microscopy
- FIG. 3 includes examples of total ion current (TIC) electropherograms of eluted protein corona after CZE-MS/MS analyses in “high-high” mode.
- Three protein corona samples (samples 1-3) were prepared in parallel and analyzed by CZE-MS/MS. Each sample was measured in triplicate (see, EXAMPLE 1).
- FIG. 4 shows proteoform intensity correlations between any two samples. The data are from “high-high” mode. Log-log plots are shown in the figure (see, EXAMPLE 1).
- FIG. 5 includes heatmaps of detected proteoform intensity across three samples. Proteoform intensity was log2 transformed and used to create the heat map using GraphPad Prism (see, EXAMPLE 1).
- FIG. 6 shows extracted ion electropherograms (EIE) of three large proteins (1-3). The m/z ion with the highest intensity for each protein was used for the peak extraction with a mass tolerance of 200 ppm (see, EXAMPLE 1).
- FIG. 7 shows the averaged mass spectrum of each protein across the peak and the corresponding deconvoluted masses of various proteoforms.
- the mass deconvolution was performed using the UniDec (Universal Deconvolution) software with default settings (see, EXAMPLE 1).
- FIG. 8 is a bar graph representing numbers of proteoform identifications, proteoform- spectrum matches (PrSMs), and proteoform family identifications from TDP (see, EXAMPLE 1).
- FIG. 9 is a bar graph representing numbers of peptide identifications, peptide- spectrum matches (PSMs), and protein group identifications from BUP (see, EXAMPLE 1).
- FIG. 10 includes four example proteoforms of SAA1 identified by TDP with sequences and fragmentation patterns and the five proteins included in the protein group SAA1 identified by BUP (see, EXAMPLE 1).
- FIG. 11 is a bar graph representing numbers of proteoform and proteoform families identified from the protein corona by using different workflows.
- the protein corona sample was from treating the protein corona-coated PSNPs by 1% SDS for 3 h at 60 °C (see, EXAMPLE 1).
- FIG. 12 is a Venn diagram of identified proteoforms from the protein corona by three different methods (CZE-MS/MS, CZE-FAIMS-MS/MS, and RPLC-MS/MS).
- the protein corona sample was from treating the protein corona-coated PSNPs by 1% SDS for 3 h at 60 °C (see, EXAMPLE 1).
- FIG. 13 is an example schematic workflow of cIEF-MS/MS-based TDP for NP protein corona.
- Polystyrene NPs PSNPs
- the figure was created using BioRenderTM and used here with permission (see, EXAMPLE 2).
- FIG. 14 includes base peak electropherograms of duplicate cIEF-MS/MS runs (see, EXAMPLE 2).
- FIG. 15 is a Venn diagram of proteoform overlaps between duplicate measurements (see, EXAMPLE 2).
- FIG. 16 is a line graph showing proteoform intensity correlation between the duplicate runs. Log2 (proteoform intensity) was used, and the proteoforms having proteoform feature intensities in both runs were used (see, EXAMPLE 2).
- FIG. 17 shows sequence and fragmentation pattern of one APOA1 proteoform (proteoform 1), having one N-terminal acetylation and diphosphorylation (see, EXAMPLE 2).
- FIG. 18 is a bar graph representing mass distribution of proteoforms identified in the two replicate runs (see, EXAMPLE 2).
- FIG. 19A shows a base peak electropherogram of protein corona by cIEF-MS/MS using an Orbitrap ExplorisTM 480 mass spectrometer in low-high mode. Deconvoluted masses of detected large proteoforms of proteins 1 (FIG. 19C and FIG. 19D) and 4 (FIG. 19B) are shown. UniDec software was used for mass deconvolution with default settings (see, EXAMPLE 2).
- FIG. 20 shows a brief example workflow of preparing the protein corona sample for TDP after incubating the PSNPs with a human plasma sample to form the protein corona on the surface of PSNPs (see, EXAMPLE 3).
- FIG. 21 is a graph representing DLS analysis of bare PSNPs (uncoated) and protein corona-coated.NPs (protein corona) (see, EXAMPLE 3).
- FIG. 22 includes an example schematic design of the high-throughput cIEF-MS/MS for protein corona analysis (see, EXAMPLE 3).
- FIG. 23 shows total ion current (TIC) electropherograms of protein corona proteoforms by CIEF-MS/MS in “High-High (HH)” and “Low-High (LH)” modes. Six selected electropherograms from runs #4, #8, #14, #16, #20, and #23 in HH and LH modes are shown (see, EXAMPLE 3).
- FIG. 24 includes intensity correlations of overlapped proteoforms between any two cIEF-MS/MS runs. Six runs were randomly selected for this analysis. Proteoform intensities were log2-transformed for the plot, and Pearson’s correlation coefficient (r) values were labeled (see, EXAMPLE 3).
- FIG. 25 is a graph representing mass distribution of the identified proteoforms from 25 cIEF-MS/MS runs (High-High) (see, EXAMPLE 3).
- FIG. 26 includes sequences and fragmentation patterns of four distinct proteoforms of Apolipoprotein A-I (APOA1) identified using cIEF-MS/MS-based TDP in "high-high” mode (see, EXAMPLE 3).
- APOA1 Apolipoprotein A-I
- MS mass spectrometry
- BUP bottom-up proteomics
- this disclosure describes a robust and reproducible MSbased top-down proteomics (TDP) technique for characterizing proteins (including one or more proteoforms) in the protein corona by adding protein-binding agents (e.g., NPs) to a biological sample to form a complex including the protein-binding agent and protein corona, separating the complex from the biological sample, separating the protein-binding agents from the protein corona to generate a protein corona sample, and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
- CE Capillary Electrophoresis
- MS/MS Tandem Mass Spectrometry
- the present TDP approach has successfully identified about 900 proteoforms in the protein corona of polystyrene NPs, ranging from 2-70 kDa, revealing proteoforms of 48 protein biomarkers with combinations of post-translational modifications, signal peptide cleavages, and/or truncations — details that bottom-up proteomics (BUP) could not fully discern.
- BUP bottom-up proteomics
- Example embodiments are provided so that this disclosure will be thorough and will fully convey the scope to those who are skilled in the art. Numerous specific details are set forth such as examples of specific compositions, components, devices, and methods, to provide a thorough understanding of embodiments of the present disclosure. It will be apparent to those skilled in the art that specific details need not be employed, that example embodiments may be embodied in many different forms and that neither should be construed to limit the scope of the disclosure. In some example embodiments, well-known processes, well-known device structures, and well-known technologies are not described in detail.
- compositions, materials, components, elements, features, integers, operations, and/or process steps are also specifically includes embodiments consisting of, or consisting essentially of, such recited compositions, materials, components, elements, features, integers, operations, and/or process steps.
- the alternative embodiment excludes any additional compositions, materials, components, elements, features, integers, operations, and/or process steps, while in the case of “consisting essentially of,” any additional compositions, materials, components, elements, features, integers, operations, and/or process steps that materially affect the basic and novel characteristics are excluded from such an embodiment, but any compositions, materials, components, elements, features, integers, operations, and/or process steps that do not materially affect the basic and novel characteristics can be included in the embodiment.
- suitable nanoscale and/or microscale materials include, but are not limited to, organic materials, non-organic materials, or combinations thereof.
- the materials are micelles, liposomes, iron oxide, graphene, silica, protein-based materials, polystyrene, silver, and gold materials, such as colloidal gold, quantum dots, palladium, platinum, titanium, and combinations thereof.
- nanoparticles are liposomes.
- One skilled in the art is able to select and prepare suitable nanoscale and/or microscale material(s).
- a “small molecule” as used herein is a molecule having a molecular weight of less than 5 kilodaltons (kDa).
- the small molecule can be synthetic or natural. In other words, the small molecule may be synthetically generated or it can be naturally produced.
- the small molecules described in the methods herein may also be referred to as protein-recruitment agents because they encourage recruitment of different proteins (and other types of biomolecules) to the surface of materials, such as nanoparticles.
- the small molecule may also be referred to as a “high-abundance protein-binding agent” or a “high- abundance protein-interacting agent.”
- High-abundance protein refers to one of the seven most abundant proteins (albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and fibrinogen) in human plasma. Such proteins collectively represent 85% of the total protein mass in human plasma.
- Low-abundance protein refers to any proteins present in human plasma that are not included in the seven high- abundance proteins.
- a “biomolecule” as used herein is a molecule produced by a living organism and essential to one or more biological processes.
- a biomolecule may also be synthetically generated to mimic the function of a naturally generated biomolecule.
- suitable biomolecules include, but are not limited to, macromolecules (such as proteins, carbohydrates, lipids, nucleic acids, etc.) as well as smaller molecules (such as vitamins, hormones, etc.).
- a biomolecule may also be referred to herein as a biological material.
- a small molecule as used herein can be considered a “small biomolecule” if it fits the definition of a small molecule (a molecule having a molecular weight of less than 5 kilodaltons (kDa)) and the definition of a biomolecule (a molecule produced by a living organism and essential to one or more biological processes or a molecule synthetically generated to mimic the function of a naturally generated bio molecule).
- a small molecule a molecule having a molecular weight of less than 5 kilodaltons (kDa)
- a biomolecule a molecule produced by a living organism and essential to one or more biological processes or a molecule synthetically generated to mimic the function of a naturally generated bio molecule.
- a “metabolite” as used herein is an intermediate or end product of metabolism. It is intended to include all natural, synthetic, and biological small molecules, such as amino acids, alcohols, polyols, alkaloids, organic acids, sugars (e.g., glucose) as well as nucleotides (e.g., inosine-5'-monophosphate and guanosine-5'-monophosphate).
- Lipids are a group of organic compounds including all natural, synthetic, and biological fatty compounds, such as glycerolipids (e.g., triacylglycerols also known as triglycerides, TG or TAG; diacylglycerols also known as diglycerides, DG or DAG, such as 1,2- diacylglycerols and 1,3-diacylglycerols), glycerophospholipids (e.g., phosphatidylcholine, phosphatidylethanolamine, L-a-phosphatidylinositol), sterol lipids, very-low-density lipoprotein (VLDL), low density lipoprotein (LDL), and high-density lipoprotein (HDL), fatty acids, prenol lipids, and sphingolipids.
- glycerolipids e.g., triacylglycerols also known as triglycerides,
- Nutrient as used herein is a substance used by an organism to survive, grow, and reproduce. In some cases a metabolite can also be considered a nutrient, such as amino acids, fatty acids, vitamins (e.g., vitamin B complex), minerals and choline.
- Plant-derived molecule refers to any molecule derived from a plant. Suitable plant-derived molecules include small molecules such as auxin, gibberellic acid, alkaloids, phenylpropanoids; plant-made biologies, such as anti-cancer biologies; and phytopharmaceutical drugs, such as those with properties against human health problems such as allergy, inflammation, etc.
- endogenous refers to a substance, e.g. a nucleic acid, protein, enzyme, small molecule, etc., that is produced from within a host organism (e.g., a human) and/or that is naturally occurring or naturally found inside a host organism.
- a host organism e.g., a human
- an endogenous substance refers to a substance produced by and found inside a host organism.
- an endogenous substance may be encoded by the genome of the host organism.
- the substance may be encoded by an autonomously replicating plasmid carried by the host organism.
- an endogenous substance is a substance that was present in a host organism’s biological sample when the biological sample was originally isolated from nature, i.e., the substance is native to the organism.
- an “endogenously produced” substance may be expressed/generated by a host organism’s own machinery.
- the host organism has not been genetically engineered to produce the substance.
- a host organism may endogenously produce a native or non-native substance.
- the method for detecting proteins in a biological sample may include a step for depleting one or more proteins in a biological sample.
- depleting one or more proteins in a biological sample may include running the biological sample through a depletion column or through a spin column with resin.
- running the biological sample through a depletion column or spin column with resin may reduce the complexity of biological samples for analysis via antibody-based techniques or proteomics techniques.
- the complexity of biological samples may be reduced for top-down proteomics analysis.
- depletion columns or spin columns may be used to reduce the complexity of biological samples, such as serum, plasma, etc., which contain high concentrations of albumin and immunoglobulins.
- depletion columns or spin columns may be used to remove highly abundant proteins, such as albumin and IgG, from biological samples.
- a suitable depletion or spin column may be a High SelectTM Depletion Spin Column (Thermo ScientificTM).
- protein depletion methods may be used for applications in drug delivery and/or imaging.
- protein depletion methods, such as those including depletion or spin columns may include particles, such as lipid-based NPs, with a small molecule, such as phosphatidylcholine (PtdChos), on their surface.
- PtdChos phosphatidylcholine
- proteins such as albumin
- proteins may be attracted and/or bound to the surface of the small molecule-coated NPs, which may enhance the small molecule-coated NP’s blood circulation time and/or allow them to be removed quickly by the immune system.
- Biomolecules that may be detected by the disclosed method may include small molecules (such as lipids, fatty acids, glycolipids, sterols, monosaccharides, vitamins, hormones, neurotransmitters, metabolites, etc.), monomers (such as amino acids, monosaccharides, isoprene, nucleotides, etc.), oligomers (such as oligopeptides, oligosaccharides, terpenes, oligonucleotides, etc.), and polymers (such as polypeptides, proteins and/or their proteoforms, polysaccharides, polyterpenes, polynucleotides, nucleic acids, etc.).
- small molecules such as lipids, fatty acids, glycolipids, sterols, monosaccharides, vitamins, hormones, neurotransmitters, metabolites, etc.
- monomers such as amino acids, monosaccharides, isoprene, nucleotides, etc.
- the method described herein may be used to detect proteins (including one or more proteoforms).
- proteoforms of proteins may be different forms of a protein produced from the genome with a variety of biological variations, such as sequence variations, splice isoforms, post-translational modifications, etc., which may alter the primary sequence and composition at the whole-protein level.
- proteoforms may carry different biological functions.
- the method described herein may be used to detect a number of unique biomolecules in a biological sample.
- the method may be used to detect a number of different proteins and/or proteoforms in a biological sample.
- the method described herein may be used to detect one type of protein and/or the method may be used to detect more than one type of protein.
- the method may be used to detect an amount of one type of protein present in a biological sample and/or present in a protein corona.
- the method may be used to detect an amount of more than one type of protein present in a biological sample and/or present in a protein corona.
- the method may be used to detect how many distinct proteins are present in a biological sample and/or are present in a protein corona.
- the method may include adding one or more small molecules along with the one or more protein-binding agents to the biological sample.
- one or more small molecule covers a “small molecule combination” (i.e., a combination of one or more small molecules).
- the method includes adding one or more small molecule or small molecule combination to the biological sample.
- the method may include adding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. small molecules to the biological sample.
- the method may include adding a combination of one or more small molecules to the biological sample.
- the small molecule combination may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. different small molecule types.
- the method may include adding one or more small molecule combination to the biological sample. In some embodiments, the method may include adding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. small molecule combinations to the biological sample. In some embodiments, the number of small molecules or small molecule combinations added to the biological sample may be any number to reach a desired concentration of small molecules or small molecule combinations in the biological sample.
- the small molecule or small molecule combination may be added to the biological sample at a concentration of about Ipg/ml to about Ig/ml.
- the small molecule or small molecule combination may be added to the biological sample at a concentration of about 1 pg/ml, about 2 pg/ml, about 3 pg/ml, about 4 pg/ml, about 5 pg/ml, about 6 pg/ml, about 7 pg/ml, about 8 pg/ml, about 9 pg/ml, about 10 pg/ml, about 20 pg/ml, about 30 pg/ml, about 40 pg/ml, about 50 pg/ml, about 60 pg/ml, about 70 pg/ml, about 80 pg/ml, about 90 pg/ml, about 100 pg/ml, about 200 pg/ml, about
- about 10 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, about 100 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, about 1,000 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, one or more small molecule or small molecule combination may be added to the biological sample at diverse concentrations. In some embodiments, the diverse concentrations may be from about 1 pg/ml to about Ig/ml.
- a small molecule may have a molecular weight of less than about 5 kDa, about 4.9 kDa, about 4.8 kDa, about 4.7 kDa, about 4.6 kDa, about 4.5 kDa, about
- a small molecule may be any molecule that weighs less than about 5 kDa.
- a small molecule as used herein may be considered a “small biomolecule” if it is a molecule having a molecular weight of less than 5 kilodaltons (kDa) and is produced by a living organism and essential to one or more biological processes (or is synthetically generated to mimic the function of a naturally generated biomolecule).
- the small molecule may be synthetically generated (i.e., artificial).
- the small molecule may be naturally produced.
- the small molecule may be endogenous to the host organism from which the biological sample is taken. In other embodiments, the small molecule may be exogenous to the host organism from which the biological sample is taken.
- the small molecule may be native or nonnative to the host organism.
- the small molecule may be native or non-native to the host organism from which the biological sample is taken, e.g. a lipid, protein, or nucleic acid, and exogenously added to the biological sample.
- the small molecule may encourage recruitment of different proteins (and other types of biomolecules) to the surface of protein-binding agents, such as nanoparticles.
- the small molecule may be capable of altering a protein corona.
- the small molecule may be capable of altering a protein corona that has formed on the surface of a protein-binding agent, such as a nanoparticle.
- the small molecule may be capable of depleting high- abundance plasma proteins, such as albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and fibrinogen, in a biological sample.
- the small molecule may be referred to as a “high-abundance protein depleting agent.”
- the small molecule may be capable of depleting at least one, at least two, at least three, at least four, at least five, at least six, or all seven of the high-abundance plasma proteins.
- the small molecule may be capable of depleting albumin.
- a small molecule may physically or chemically interact with one or more proteins in a biological sample.
- the small molecule may be referred to as a “protein-interacting agent.”
- the small molecule may bind proteins in the biological sample.
- the small molecule may bind albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and/or fibrinogen.
- the small molecule may bind one or more abundant protein types.
- binding the abundant proteins may deplete the number of abundant proteins in a biological sample.
- a small molecule may change the conformation of one or more proteins in the biological sample.
- the small molecule may be a metabolite and/or a derivative thereof.
- a metabolite may be any natural, synthetic, or biological small molecule, such as an amino acid, alcohols, polyols, alkaloids, organic acids, sugars (e.g., glucose) as well as nucleotides (e.g., inosine-5'-monophosphate and guanosine-5'-monophosphate).
- the small molecule may be a lipid and/or a derivative thereof.
- lipids may be a group of organic compounds including all natural, synthetic, and biological fatty compounds, such as glycerolipids (e.g., triacylglycerols also known as triglycerides, TG or TAG; diacylglycerols also known as diglycerides, DG or DAG, such as 1,2- diacylglycerols and 1,3-diacylglycerols), glycerophospholipids (e.g., triacylglycerols also known as triglycerides, TG or TAG; diacylglycerols also known as diglycerides, DG or DAG, such as 1,2- diacylglycerols and 1,3-diacylglycerols), glycerophospholipids (e.g.
- phosphatidylcholine phosphatidylethanolamine, L-a-phosphatidylinositol
- sterol lipids very-low-density lipoprotein (VLDL), low density lipoprotein (LDL), and high-density lipoprotein (HDL)
- VLDL very-low-density lipoprotein
- LDL low density lipoprotein
- HDL high-density lipoprotein
- fatty acids prenol lipids, and sphingolipids.
- the small molecule may be a nutrient and/or a derivative thereof.
- a metabolite can also be considered a nutrient, such as amino acids, fatty acids, vitamins (e.g. vitamin B complex), minerals and choline.
- the small molecule may be a plant-derived molecule and/or a derivative thereof, such as a small molecule (e.g., auxin, gibberellic acid, alkaloid, phenylpropanoid), a plant-made biologic (e.g., anti-cancer biologies), or phytopharmaceutical drugs (e.g., those with properties against human health problems, such as allergy, inflammation, etc.).
- a small molecule e.g., auxin, gibberellic acid, alkaloid, phenylpropanoid
- a plant-made biologic e.g., anti-cancer biologies
- phytopharmaceutical drugs e.g., those with properties against human health problems, such as allergy, inflammation, etc.
- the small molecule may be a metabolite and/or a derivative thereof, lipid and/or a derivative thereof, nutrient and/or a derivative thereof, plant-derived molecule and/or a derivative thereof, or a combination thereof.
- the small molecules may be a combination of phosphatidylethanolamine and/or a derivative thereof, L-a-phosphatidylinositol and/or a derivative thereof, inosine 5’- monophosphate and/or a derivative thereof, and vitamin B complex and/or a derivative thereof.
- the one or more protein-binding agents may be added to the biological sample at a concentration of about Ipg/ml to about Ig/ml.
- the one or more protein-binding agents may be added to the biological sample at a concentration of about 1 pg/ml, about 2 pg/ml, about 3 pg/ml, about 4 pg/ml, about 5 pg/ml, about 6 pg/ml, about 7 pg/ml, about 8 pg/ml, about 9 pg/ml, about 10 pg/ml, about 20 pg/ml, about 30 pg/ml, about 40 pg/ml, about 50 pg/ml, about 60 pg/ml, about 70 pg/ml, about 80 pg/ml, about 90 pg/ml, about 100 pg/ml, about 200 pg/ml, about 300
- the protein-binding agent may be any material of which a single unit is sized (in at least one dimension) at about 1 nm and about 100,000 nm.
- the protein-binding agent may be about 1 nm, about 50 nm, about 100 nm, about 500 nm, about 1,000 nm, about 5,000nm about 10,000 nm, about 50,000 nm, about 100,000 nm, about 50 nm to about 100,000 nm, about 100 nm to about 100,000 nm, about 500 nm to about 100,000 nm, about 1,000 nm to about 100,000 nm, about 5,000 nm to about 100,000 nm, about 10,000 to about 100,000 nm, about 50,000 to about 100,000 nm, about 1 nm to about 50,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 5,000
- the protein-binding agent may be made from an organic material, a non-organic material, or a combination thereof.
- the material may be micelles, liposomes, iron oxide, graphene, silica, protein-based materials, polystyrene, silver, and gold materials, quantum dots, palladium, platinum, titanium, and a combination thereof.
- the protein-binding agent may have a poly dispersity index (PDI) of about 0.01 to about 10.
- PDI poly dispersity index
- the protein-binding agent may have a PDI of about 0.01, about 0.02, about 0.03, about 0.04, about 0.05, about 0.06, about 0.07, about 0.08, about 0.09, about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.02 to about 0.6, about 0.03 to about 0.5, about 0.04 to about 0.4, 0.05 to about 0.3, 0.06 to about 0.2, 0.07 to about 0.1, 0.08 to about 0.09, about 0.01 to about 0.6, about 0.01 to about 0.5, about 0.01 to about 0.4, about 0.01 to about 0.3, about 0.01 to about 0.2, about 0.01 to about 0.1, about 0.01 to about 0.09, about 0.01 to about 0.08, about 0.01 to about 0.07, about 0.01 to about 0.06, about 0.01 to about
- the PDI of the protein-binding agent may be about 0.01 to about 0.7. In a specific embodiment, the protein-binding agent PDI may be about 0.7. In another embodiment, the protein-binding agent PDI may be about 0.3. In another embodiment, the proteinbinding agent PDI may be about 0.2.
- the soft corona may include lower-affinity proteins that are reversibly bound to the protein-binding agent’s surface and/or to other proteins in the protein corona.
- proteins in the soft corona may be exchanged or detached over time.
- larger proteins with lower affinities may aggregate to the protein-binding agent’s surface first, and over time, smaller proteins with higher affinities may replace them (i.e., “hardening” the corona).
- the type and/or number of proteins available to bind a protein-binding agent may be altered by one or more small molecules in the biological sample.
- phosphatidylcholine may bind high-abundance proteins, such as albumin, thus altering the protein corona composition to include less high-abundance proteins, such as albumin.
- the sample containing the one or more protein-binding agents may be incubated to allow a protein corona to form on the surface of the one or more protein-binding agents.
- the sample may be incubated with the one or more protein-binding agent.
- the biological sample may be incubated for at least 10 seconds to about 24 hours.
- the biological sample may be incubated for at least about 10 seconds, at least about 15 seconds, at least about 20 seconds, at least about 25 seconds, at least about 30 seconds, at least about 40 seconds, at least about 50 seconds, at least about 60 seconds, at least about 90 seconds, at least about 2 minutes, at least about 3 minutes, at least about 4 minutes, at least about 5 minutes, at least about 6 minutes, at least about 7 minutes, at least about 8 minutes, at least about 9 minutes, at least about 10 minutes, at least about 15 minutes, at least about 20 minutes, at least about 25 minutes, at least about 30 minutes, at least about 45 minutes, at least about 50 minutes, at least about 60 minutes, at least about 90 minutes, at least about 2 hours, at least about 3 hours, at least about 4 hours, at least about 5 hours, at least about 6 hours, at least about 7 hours, at least about 8 hours, at least about 9 hours, at least about 10 hours, at least about 12 hours, at least about 14 hours, at least about 15 hours, at least about 16 hours, at least about 17 hours, at least about
- the incubation temperature can be determined by one skilled in the art, and includes temperatures of about 4° C to about 40° C, about 4° C to about 20° C, about 10° C to about 15° C, about 10° C to about 40° C, about 4° C, about 5° C, about 6° C, about 7° C, about 8° C, about 9° C, about 10° C, about 11° C, about 12° C, about 13° C, about 14° C, about 15° C, about 16° C, about 17° C, about 18° C, about 19° C, about 20° C, about 21° C, about 22° C, about 25° C, about 30° C, about 35°, about 37° C, etc.
- the method may be performed at room temperature (e.g., about 37° C.; e.g., about 35° C to about 40° C).
- the biological sample may be diluted.
- the biological sample may be diluted using a suitable buffer, such as phosphate buffer saline (PBS), Tris-HCl buffer, ammonium bicarbonate buffer, HEPES (4-(2-hydroxyethyl)-l- piperazineethanesulfonic acid) buffer, MOPS (3-(N-morpholino)propanesulfonic acid) buffer, PIPES (1,4-Piperazinediethanesulfonic acid) buffer, or EPPS (4-(2-Hydroxyethyl)-l- piperazinepropanesulfonic acid) buffer.
- PBS phosphate buffer saline
- Tris-HCl buffer Tris-HCl buffer
- ammonium bicarbonate buffer such as phosphate buffer saline (PBS), Tris-HCl buffer, ammonium bicarbonate buffer, HEPES (4-(2-hydroxyethyl)-l- piperazineethanesulfonic acid) buffer,
- the biological sample may be diluted to a final concentration of about 25% to about 85%, about 30% to about 80%, about 35% to about 75%, about 40% to about 10%, about 45% to about 65%, about 50% to about 60%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, or about 85%.
- the biological sample may be diluted to a final concentration of about 55% using PBS.
- separating the complex including the protein corona and the one or more protein binding agents from the biological sample may include one or more rounds of washing and centrifuging to remove the supernatant including the biological sample and any proteins not bound or that have low affinity for the one or more protein -binding agents.
- at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 washing and/or centrifugation steps may be performed.
- six washes may be performed. Proteins that are removed from the surface of protein-binding agents after centrifugation with at least 5000 g may be considered “low affinity proteins,” i.e., proteins which have low affinity for binding to the surface of proteinbinding agents.
- the complex may be separated from the remainder of the biological sample, for example, by centrifugation (such as gradient centrifugation), size exclusion chromatography, magnetic separation, field-flow fractionation, etc.
- the complex may be washed and resuspended.
- the proteins from the complex may be reduced, alkylated, and/or digested.
- proteins in the hard protein corona, soft protein corona, or a combination of the hard and soft protein corona may be detected.
- proteins present in the protein corona may be detected after a biological sample has been incubated with one or more protein-binding agent, and optionally one or more small molecule.
- the protein corona composition may change over time. For example, in some embodiments, a protein corona composition after 10 minutes of incubation may be different than a protein corona composition after 20 minutes of incubation.
- molecules other than proteins may be detected that have been bound to the protein-binding agent.
- separating the one or more protein-binding agents from the protein corona in the complex sample may include eluting the protein corona from the one or more protein-binding agents.
- eluting the protein corona from the one or more protein-binding agents may include adding an elution buffer containing a detergent and/or an ionic liquid to the complex sample.
- An ionic liquid with detergent-like capabilities may be used, for example an ionic liquid with 4-12 carbons in its hydrocarbon chain.
- a suitable detergent and/or ionic liquid can easily be selected by one skilled in the art, and includes, but is not limited to, sodium dodecyl sulphate (SDS), sodium deoxycholate (SDC), NP-40, 1 -butyl- 3 -methyl imidazolium tetrafluoroborate (BMIM BF4), l-dodecyl-3-methylimidazolium chloride (C12Im- Cl), or a combination thereof.
- detergent concentration i.e., sodium dodecyl sulfate, SDS
- elution temperature, time, or a combination thereof may influence protein recovery.
- the amount of detergent and/or ionic liquid in the elution buffer can be determined by one skilled in the art, and includes percentages of about 0.1%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, or about 5%.
- the elution buffer may contain about 0.1% to about 5%, about 0.5% to about 4.5%, about 1% to about 4%, about 1.5% to about 3.5%, about 2% to about 3%, about 0.1% to about 4.5%, about 0.1% to about 4%, about 0.1% to about 3.5% about 0.1% to about 3%, about 0.1% to about 2.5%, about 0.1% to about 2%, about 0.1% to about 1.5%, about 0.1% to about 1%, about 0.1% to about 0.5%, about 0.5% to about 5%, about 1% to about 5%, about 1.5% to about 5%, about 2% to about 5%, about 2.5% to about 5%, about 3% to about 5%, about 3.5% to about 5%, about 4% to about 5%, or about 4.5% to about 5% detergent and/or ionic liquid.
- the elution buffer may contain about 0.1% to about 5% detergent and/or ionic liquid.
- eluting the protein corona from the one or more protein-binding agents may include adding the elution buffer including detergent and/or an ionic liquid to the complex sample and incubating the complex sample.
- the incubation time can be determined by one skilled in the art, and includes times of about 0.5 hours, about 1 hour, about 1.5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, or about 5 hours.
- the complex sample may be incubated for about 0.5 to about 5 hours, about 1 hour to about 4.5 hours, about 1.5 hours to about 4 hours, about 2 hours to about 3.5 hours, about 2.5 hours to about 3 hours, about 0.5 hours to about 4.5 hours, about 0.5 hours to about 4 hours, about 0.5 hours to about 3.5 hours, about 0.5 hours to about 3 hours, about 0.5 hours to about 2.5 hours, about 0.5 hours to about 2 hours, about 0.5 hours to about 1.5 hours, about 0.5 hours to about 1 hour, about 1 hour to about 5 hours, about 1.5 hours to about 5 hours, about 2 hours to about 5 hours, about 2.5 hours to about 5 hours, about 3 hours to about 5 hours, about 3.5 hours to about 5 hours, about 4 hours to about 5 hours, or about 4.5 hours to about 5 hours.
- the complex sample may be incubated with the elution buffer for about 0.5 to about 5 hours.
- the incubation temperature can be determined by one skilled in the art, and includes temperatures of about 20°C, about 25°C, about 30°C, about 35°C, about 40°C, about 45°C, about 50°C, about 55°C, about 60°C, about 65°C, about 70°C, about 75°C, about 80°C, about 85°C, about 90°C, 95°C, or about 100°C.
- the complex sample may be incubated at about 20°C to about 100°C, about 25°C to about 95°C, about 30°C to about 90°C, about 35°C to about 85°C, about 40°C to about 80°C, about 45°C to about 75°C, about 50°C to about 70°C, about 55°C to about 65°C, about 20°C to about 95°C, about 20°C to about 90°C, about 20°C to about 85°C, about 20°C to about 80°C, about 20°C to about 75°C, about 20°C to about 70°C, about 20°C to about 65°C, about 20°C to about 60°C, about 20°C to about 55°C, about 20°C to about 50°C, about 20°C to about 45°C, about 20°C to about 40°C, about 20°C to about 35°C, about 20°C to about 30°C, about 20°C to about 25°C, about 25°C to about 100°C, about
- eluting the protein corona from the one or more proteinbinding agents may include adding the elution buffer, containing about 0.1% to about 5% detergent and/or ionic liquid, to the complex sample and incubating the complex sample for about 0.5 to about 5 hours at about 20°C to about 100°C.
- separating the one or more protein-binding agents from the protein corona may include dissolving the one or more protein-binding agents.
- different protein-binding agents may be dissolved by different solutions.
- the dissolving solution may be selected based off the type of protein-binding agent used.
- one or more gold protein-binding agents may be dissolved by a potassium-iodide/iodine etching solution.
- one or more polystyrene protein-binding agents may be dissolved by ethyl acetate.
- one or more iron oxide nanoparticles may be dissolved by ethylenediaminetetraacetic acid (EDTA).
- the dissolving solutions may digest the protein-binding agents while keeping the protein corona shell intact.
- compositions including the complex (including the protein corona and the one more protein-binding agents) and the dissolving solution are provided herein.
- a composition including the complex (including the protein corona and the one more protein-binding agents) and EDTA is also provided herein.
- compositions including one or more proteins obtained from one or more complex (including one or more proteins and a protein-binding agent) where the protein-binding agent has been removed from the composition.
- the composition may contain substantially no protein-binding agent.
- the protein-binding agent may be removed from the composition by adding a detergent and/or an ionic liquid as described herein.
- the protein-binding agent may be dissolved in the composition.
- the dissolved protein-binding agent, such as nanoparticles may stay in the composition during CE.
- the dissolved protein-binding agent, such as nanoparticles may be removed from the composition before CE.
- the protein corona sample may be buffer-exchanged to a mass spectrometry (MS)-compatible buffer prior to characterizing the proteins in the protein corona sample.
- MS mass spectrometry
- the protein samples may contain a high-concentration of detergent (i.e., SDS) and/or ionic liquid in the elution buffer, which can be incompatible with mass spectrometry.
- the detergent and/or ionic liquid is removed from the composition.
- removing the detergent and/or ionic liquid maintains high protein recovery during the cleanup.
- the buffer exchange approach is used for sample cleanup with a molecular weight cutoff membrane.
- the buffer-exchange may include about a 3-kDa, 5-kDa, 10- kDa, 15-kDa, 20-kDa, 25-kDa, or 30-kDa molecular weight cutoff filter.
- rounds of washing and/or centrifuging may be included in the buffer exchange sample clean up.
- at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 washing/centrifugation steps may be performed.
- six washes may be performed.
- urea may be used as the washing buffer.
- about 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7.0, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8.0, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9.0, 9.1, 9.2, 9.3 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, or 10M washing buffer may be used.
- about 8M washing buffer may be used.
- about 8M urea may be used as the washing buffer.
- the CE may be Capillary Zone Electrophoresis (CZE) or Capillary Isoelectric Focusing (cIEF).
- CE may be used to detect and/or characterize proteins from protein coronas.
- CE may be used to determine protein isoforms and/or study post-translational protein modifications.
- CZE and cIEF are described in SUN, L. et al. (2014) Capillary zone electrophoresis for analysis of complex proteomes using an electrokinetically pumped, sheath flow nanospray interface. Proteomics. 14(0): 622-628 and XU, T. and SUN, L. (2021) A Mini Review on Capillary Isoelectric Focusing- Mass Spectrometry for Top-Down Proteomics. Front Chem. 9:651757, which are both incorporated by reference in their entirety.
- cIEF separates molecules based on their isoelectric point (pl), the pH at which a molecule carries no net charge.
- pl isoelectric point
- a pH gradient may be established in the capillary, and molecules may migrate to the point where their net charge is zero.
- CZE separates molecules based on the size and/or charge.
- molecules may migrate through a capillary filled with an electrolyte solution under the influence of an electric field, with smaller and more highly charged molecules moving faster.
- about 100 ng, about 150 ng, about 200 ng, about 250 ng, about 300 ng, about 350 ng, about 400 ng, about 450 ng, about 500 ng, about 550 ng, about 600 ng, about 650 ng, about 700 ng, about 750 ng, about 800 ng, about 850 ng, about 900 ng, about 950 ng, or about 1,000 ng of proteins and in the protein corona sample may be loaded into a capillary for CE.
- about 100 ng to about 1,000 ng, about 200 ng to about 900 ng, about 300 ng to about 800 ng, about 400 ng to about 700 ng, about 500 ng to about 600 ng, about 100 ng to about 900 ng, about 100 ng to about 800 ng, about 100 ng to about 700 ng, about 100 ng to about 600 ng, about 100 ng to about 500 ng, about 100 ng to about 400 ng, about 100 ng to about 300 ng, about 100 ng to about 200 ng, about 200 ng to about 1,000 ng, about 300 ng to about 1,000 ng, about 400 ng to about 1,000 ng, about 500 ng to about 1,000 ng, about 600 ng to about 1,000 ng, about 700 ng to about 1,000 ng, about 800 ng to about 1,000 ng, or about 900 ng to about 1,000 ng of proteins in the protein corona sample may be loaded into a capillary for CE.
- about 100 to about 1,000 ng of proteins in the protein corona sample may be
- the method described herein may detect more proteins (including one or more proteoforms) compared to using a biological sample without the one or more protein-binding agent.
- the present method may result in an increase in the number of detected proteins compared to using a biological sample without the one or more protein-binding agent.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- the disclosed method may increase the detection of low- abundance proteins.
- adding one or more small molecules to the biological sample with one or more protein-binding agents may reduce the detection of one or more high-abundance proteins by at least about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 99%.
- the present disclosure can provide a method of increasing the number of proteins detected in a biological sample compared to other known methods.
- protein coronas may be used to detect one or more protein types and/or the amount of one or more protein types present in a biological sample.
- multiple protein-binding agent types may be used to detect more than one protein and, multiple protein-binding agent types may be used.
- different protein-binding agents may attract different protein types.
- using more than one type of protein-binding agent may increase the number of detected proteins in a biological sample.
- each protein-binding agent type may have distinct physicochemical properties.
- the protein corona formed around the protein-binding agents may be different for different protein-binding agents.
- the present disclosure also provides a method for detecting biomarkers in a biological sample.
- the present disclosure provides a method for detecting one or more biomarkers or a pattern of one or more biomarkers associated with a disease, including adding one or more protein-binding agents, having a surface capable of binding proteins, to at least two biological samples, where each biological sample from different subjects diagnosed with the disease.
- the method includes allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents.
- the method also includes separating the complex from the biological sample to generate a complex sample, separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample, and detecting the one or more biomarkers or the pattern of one or more biomarkers in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
- CE Capillary Electrophoresis
- MS/MS Tandem Mass Spectrometry
- the one or more biomarkers or pattern of one or more biomarkers may be associated with a health spectrum condition, such as a disease or disorder.
- the method may include adding one or more small molecules, as described herein, to at least two biological samples, where each biological sample may be from different subjects diagnosed with the disease.
- the method may also include adding one or more proteinbinding agent (which may all be the same type or a combination of different types), as described herein, having a surface capable of binding proteins to the biological samples.
- a biomarker pattern may include multiple types of biomarkers.
- a biomarker pattern may include a certain amount of one or more biomarkers in a biological sample.
- individual organisms or biological samples may have different biomarker patterns.
- a biomarker pattern, similar to a biomarker may be used to detect and/or diagnose a disease or disorder in a subject.
- a biomarker may be a measurable indicator of some biological state or condition.
- bio markers may be upregulated or downregulated according to different disease types or disease stages.
- the biomarker may be a molecular, physiologic, histologic, and/or radiographic biomarker. Additionally or alternatively, the biomarker may be predictive, prognostic, or diagnostic. In some embodiments, predictive biomarkers may help optimize ideal treatments. Examples of predictive biomarkers may include HER2/neu in breast cancer or EGFR1 mutations in non-small cell lung cancer.
- diagnostic biomarkers can be a traceable substance that is introduced into an organism as a means to examine aspects of health.
- a diagnostic biomarker may be used as a substance whose detection indicates a particular disease state.
- the presence of an antibody may indicate an infection.
- a diagnostic biomarker may be prostate-specific antigen (PSA), which may be used as a proxy of prostate size with rapid changes potentially indicating cancer.
- PSA prostate-specific antigen
- multiple protein-binding agent types may be used.
- different protein-binding agents may attract different biomarker types.
- using more than one type of protein-binding agent may increase the number of detected biomarkers in a biological sample.
- the health spectrum condition may be any health state, including complete well-being, minor health issues, chronic conditions, and severe illnesses.
- the health spectrum refers to overall well-being and the factors that influence it, whether they lead to optimal health or contribute to illness.
- a health spectrum may include a condition a subject may have which is not considered a disease or disorder yet for medical treatment.
- a health spectrum condition may include a predisposition for a disease or disorder.
- the presence of biomarker(s) may indicate a subject is at risk of developing a health spectrum condition and/or a disease/disorder.
- the biomarker(s) for a disease/disorder may be the same for a disease or disorder predisposition. Additionally or alternatively, the biomarker(s) for a disease/disorder may be different for the same disease or disorder.
- the health spectrum condition may be a disease or a disorder. Some health spectrum condition examples include, but are not limited to, a predisposition to, or risk of developing, or diagnosed obesity, heart disease, liver disease, kidney disease, depression, cancer, etc.
- the disease may be any condition that adversely affects the structure or function of all or part of an organism and is not immediately due to any external injury.
- the disease the biomarker may be associated with may be a neoplastic disease, a cardiovascular disease, a metabolic disease, an infectious disease, an inflammatory disease, a congenital disease, a hereditary disease, a degenerative disease, a neurological disease, or a combination thereof.
- the disease may be a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof.
- a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof.
- the disease may be a neurological disease selected from Alzheimer’s disease, brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke, frontotemporal dementia, demyelinating disease, multiple sclerosis, Devic's disease, central pontine myelinolysis, progressive multifocal leukoencephalopathy, leukodystrophies, Guillain-Barre syndrome, progressing inflammatory neuropathy, Charcot-Marie-Tooth disease, chronic inflammatory demyelinating polyneuropathy, anti-MAG peripheral neuropathy, and a combination thereof.
- Alzheimer’s disease brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke
- Suitable biomarkers for Alzheimer’s disease include, but are not limited to, for example, the amyloid beta (AP)42/40 ratio, phosphorylated tau (p-tau), serum neurofilament light chain (NfL), and glial fibrillary acidic protein (GFAP).
- AP amyloid beta
- p-tau phosphorylated tau
- NfL serum neurofilament light chain
- GFAP glial fibrillary acidic protein
- Suitable cancer biomarkers include, but are not limited to, for example, AHSG (a2- HS-Glycoprotein), AKR7A2 (Aflatoxin Bl aldehyde reductase), AKT3 (PKB y), ASGR1 (ASGPR1), BDNF, BMP1 (BMP-1), BMPER, C9, CA6 (Carbonic anhydrase VI), CAPG (CapG), Carcino-embryonic antigen, CDH1 (Cadherin-1), CHRDL1 (Chordin-Like 1), CKB-CKM-(CK- MB), CLIC1 (chloride intracellular channel 1), CM Al (Chymase), CNTN1 (Contactin- 1), COL18A1 (Endostatin), CRP, CTSL2 (Cathepsin V), DDC (dopa decarboxylase), EGFR (ERBB1), FGA-FGB-FGG (D-dimer),
- biomarkers for breast cancer include, but are not limited to, Circulating Tumor Cells (EpCAM, CD45, cytokeratins 8, 18+, 19+), ER/PR, HER- 2/neu, CA15-3, CA27.29, and the like.
- Biomarkers for colorectal cancer include, but are not limited to, for example, EGFR, KRAS, UGT1A1, Fibrin/ fibrinogen degradation product (DR- 70), Human hemoglobin (fecal occult blood), and the like.
- Biomarkers associated with leukemia/lymphoma include, but are not limited to, e.g., CD20 antigen, CD30, FIP1L1- PDGFRalpha, PDGFR, Philadelphia Chromosome (BCR/ABL), PML/RAR alpha, TPMT, UGT1A1, and the like.
- Biomarkers associated with lung cancer include but are not limited to, e.g., ALK, EGFR, KRAS and the like.
- Biomarkers associated with ovarian cancer include but are not limited to, e.g., ROMA (HE4+CA-125), OVA1 (multiple proteins), HE4, CA-125, and the like.
- Biomarkers associated with hepatocellular cancer include but are not limited to AFP-L3%, and the like.
- Biomarkers associated with gastrointestinal stromal tumors include but are not limited to c-Kit, and the like.
- Biomarkers associated with pancreatic cancer include but are not limited to CAI 9-9, and the like.
- Biomarkers are known in the art, and can be found in, for example, Bigbee W, Herberman R B. Tumor markers and immunodiagnosis. In: Bast R C Jr., Kufe D W, Pollock R E, et al., editors. Cancer Medicine. 6th ed. Hamilton, Ontario, Canada: BC Decker Inc., 2003; Andriole G, Crawford E, Grubb R. et al.
- Biomarkers associated with a cardiovascular disease may include, but are not limited to, lipid profile, glucose, and hormone level and physiological biomarkers based on measurement of levels of important biomolecules such as serum ferritin, triglyceride to HDLp (high density lipoproteins) ratio, lipophorin-cholesterol ratio, lipid-lipophorin ratio, LDL cholesterol level, HDLp and apolipoprotein levels, lipophorins and LTPs ratio, sphingolipids, Omega-3 Index, and ST2 level, among others.
- Suitable biomarkers for cardiovascular disease can be found in the art, for example, but not limited to, in van Holten et al.
- Biomarkers associated with a neurological disease may include, but are not limited to, e.g., Api-42, t-tau andp-tau 181, a-synuclein, among others. See, e.g., Chintamaneni and Bhaskar “Biomarkers in Alzheimer's Disease: A Review” ISRN Pharmacol. 2012. 2012: 984786. Published online 2012 Jun. 28, incorporated by reference in its entirety.
- a method of diagnosing a disease or identifying another health spectrum condition, such as a predisposition for a disease in a subject includes adding one or more protein-binding agents, having a surface capable of binding proteins, to a biological sample from the subject and allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents.
- the method also includes separating the complex from the biological sample to generate a complex sample, separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample, and detecting one or more biomarkers or a pattern of one or more biomarkers associated with the disease in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
- CE Capillary Electrophoresis
- MS/MS Tandem Mass Spectrometry
- the “health spectrum” covers a range of health states, including complete well-being, minor health issues, chronic conditions, and severe illnesses.
- some health spectrum condition examples include both mental and physical conditions at a predisease or pre-disorder level all the way to severe illness, such as, but are not limited to, obesity, heart disease, liver disease, kidney disease, depression, cancer, etc.
- the disease may be a neoplastic disease, a cardiovascular disease, a metabolic disease, an infectious disease, an inflammatory disease, a congenital disease, a hereditary disease, a degenerative disease, a neurological disease, or a combination thereof.
- the disease may be a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof.
- the disease may be a neurological disease selected from Alzheimer’s disease, brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke, frontotemporal dementia, demyelinating disease, multiple sclerosis, Devic's disease, central pontine myelinolysis, progressive multifocal leukoencephalopathy, leukodystrophies, Guillain-Barre syndrome, progressing inflammatory neuropathy, Charcot-Marie-Tooth disease, chronic inflammatory demyelinating polyneuropathy, anti-MAG peripheral neuropathy, and a combination thereof.
- Alzheimer’s disease brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke
- the disclosed method may be used to provide a predisposition (e.g. a chance or likelihood) of developing the disease.
- the disclosed method may be used to detect early onset of a disease (i.e., early disease diagnosis).
- apolipoproteins may be detected, which are indicators of cardiovascular and neurodegenerative disorders.
- the present method may detect protein categories that are important in disease onset and progression.
- biomarkers(s) for a disease/disorder may be the same and/or different for a predisposition of a disease/disorder.
- the same biomarkers that are detected for diagnosing a disease/disorder may be detected to identify a predisposition for that disease/disorder.
- the biomarker concentration or amount may be lower in the sample for identifying a predisposition for a disease compared to diagnosing the disease.
- diagnosing a disease, or identifying another health spectrum condition, such as a predisposition for a disease or health spectrum condition may help prevent or reduce the risk of developing conditions such as obesity, heart disease, liver disease, kidney disease, depression, cancer, etc.
- subjects diagnosed with a disease, a health spectrum condition, or identified as having a predisposition for a disease or health spectrum condition may take measures to prevent or slow the progression of the disease or condition.
- EXAMPLE 1 CZE-MS characterization of proteins in protein corona
- Ammonium bicarbonate (ABC), 3 -(trimethoxy silyl) propyl methacrylate, dithiothreitol (DTT), iodoacetamide (IAA), and Amicon® Ultra (0.5 mL, 10 kDa cutoff size) centrifugal filter units were ordered from Sigma- Aldrich® (St. Louis, MO).
- LC/MS grade water, acetonitrile (ACN), HPLC-grade acetic acid (AA), and fused silica capillaries (50 mm i.d., 360 mm o.d., Polymicro Technologies) were purchased from Fisher ScientificTM (Pittsburgh, PA).
- Acrylamide was obtained from Acros OrganicsTM (Fair Lawn, NJ). Healthy human plasma protein was purchased from Alternative ResearchTM (www.innov-research.com) and diluted to 55% using phosphate buffer solution (PBS, lx). Plain polystyrene nanoparticles (PSNPs, ⁇ 100 nm) were obtained from Polysciences® (www.polysciences.com).
- the protein-NP complexes underwent two cold PBS washes under identical conditions. Subsequently, two-thirds of the resulting protein-NP complexes were collected for the TDP experiment, while the remaining protein-NP complexes were used for the BUP experiment.
- Sample Preparation for TDP The protein-NP complexes were treated with four different conditions to elute the protein corona from the surface of the NPs.
- the protein corona-coated PSNPs were incubated in a 0.4% (w/v) SDS solution for 1.5 h at 60°C with constant agitation.
- the supernatant, containing the protein corona in a 0.4% SDS solution was separated from the PSNPs through centrifugation at 19,000g for 20 min at 4°C.
- the resulting supernatant underwent an additional centrifugation step under the same condition to ensure complete removal of the PSNPs.
- the final protein corona sample was cleaned through a buffer exchange step.
- An Amicon® Ultra Centrifugal Filter with a molecular weight cutoff (MWCO) of 10 kDa was employed for the buffer exchange, effectively eliminating SDS from the protein samples.
- 1% SDS with a 1.5 h incubation, 1% SDS with a 3 h incubation, and 2% SDS with a 3 h incubation were employed using the same procedure as the 0.4% SDS and 1.5 h incubation with one minor change.
- the SDS and PSNP solution was centrifuged at 14,000g for 20 min at room temperature to remove the PSNPs followed by a second centrifugation at 19,000g at 4°C for 20 min. The resulting supernatant was then transferred to a different tube and underwent an additional centrifugation step at 19,000g for 20 min at 4°C to ensure complete removal of the PSNPs.
- the buffer exchange protocol started with the initial wetting of the filter using 20 pL of 100 mM ammonium bicarbonate (pH 8.0) followed by centrifugation at 14,000g for 10 min. Subsequently, 200 pg of proteins was added to the filter, and centrifugation was carried out for 20 min at 14,000g. A total of 200 pL of 8 M urea in 100 mM ammonium bicarbonate solution was added followed by centrifugation at 14,000g for 20 min. This step was repeated twice under the same conditions to ensure the complete removal of SDS and other small interferences. To eliminate urea from the purified protein, the filter underwent three additional rounds of buffer exchange. Precisely, 100 mM ammonium bicarbonate was added to the filter, bringing the final volume to 200 pL. All steps were executed with centrifugation at 4°C, ensuring the thorough removal of urea from the protein corona.
- the protein samples were diluted four times using 100 mM ammonia bicarbonate followed by trypsin (1.5 pg, bovine pancreas TPCK-treated) digestion at 37°C overnight. The digestion was finally terminated by adding formic acid (0.6% (v/v) final concentration).
- the samples were desalted with Sep-Pak® Cl 8 Cartridge (Waters, Milford, MA) according to manufacturer’s protocol. The eluates were lyophilized in a vacuum concentrator and then redissolved in 70 pL of 100 mM ABC buffer (pH 8.0).
- CZE-MS/MS Linear polyacrylamide (LPA)-coated fused silica capillaries (50 pm i.d., 360 pm o.d.) were prepared according to CHEN, D. et al. (2017) Capillary zone electrophoresis-mass spectrometry with microliter-scale loading capacity, 140 min separation window and high peak capacity for bottom-up proteomics. Analyst. 142:2118-2127 and ZHU, G. et al. (2016) Thermally-initiatedfree radical polymerization for reproducible production of stable linear polyacrylamide coated capillaries, and their application to proteomic analysis using capillary zone electrophoresis— mass spectrometry. Taianta. 146:839-843.
- the CZE-MS/MS system configuration involved the integration of a CESI 8000 Plus CE system (Beckman Coulter) with an Orbitrap ExplorisTM 480 mass spectrometer (Thermo Fisher ScientificTM), employing an in-house-built electrokinetically pumped sheath-flow CE-MS nanospray interface.
- the interface featured a glass spray emitter with an orifice size of 30-35 pm, filled with sheath buffer composed of 0.2% (v/v) formic acid and 10% (v/v) methanol.
- the spray voltage was about 2 kV.
- the length of the LPA-coated CZE capillary was 100 cm.
- the capillary’s inlet was securely affixed within the cartridge of the CE system, while its outlet was inserted into the emitter of the interface.
- the capillary outlet-emitter orifice distance was maintained at approximately 0.5 mm.
- TDP For TDP, a 5 psi pressure was applied to load ⁇ 240 ng of corona proteins (2.4 mg/mL, injection volume of 100 nL) into the capillary and then the inlet of the capillary was inserted into the background electrolyte (BGE, 5% (v/v) acetic acid) for CZE separation with a separation voltage of 30 kV.
- BGE background electrolyte
- 115 nL of each corona peptide sample was loaded for CZE-MS/MS. Following this, the capillary’s inlet was immersed into the BGE (5% (v/v) acetic acid), initiating the CZE separation process under a separation voltage of 30 kV.
- the maximum ion injection times were set as 50 and 100 ms, respectively.
- the precursor isolation width was 2 m/z.
- the dynamic exclusion was applied with a duration of 15 s, and the exclusion of isotopes was enabled.
- precursor ions in full MS spectra were isolated with a 2 m/z window and subjected to fragmentation through higher-energy collisional dissociation (HCD) with a normalized collision energy (NCE) of 25%. Only precursor ions with an intensity exceeding 1 x 10 4 and a charge state ranging from 5 to 60 nm were selected for fragmentation.
- Product ions were detected with a resolution of 120,000 (at 200 m/z), utilizing three microscans and maintaining a normalized AGC target value of 100% for both conditions. Dynamic exclusion was enabled with a duration of 30 s and a mass tolerance of 10 ppm (parts per million). Additionally, the “Exclude isotopes” function was activated.
- RPLC-MS/MS for TDP of Protein Corona The RPEC separation was performed using an EASY-nEC 1200 system from Fisher ScientificTM. A 1 pF aliquot of the protein corona sample (0.3 mg/mE) was loaded onto a home-packed C4 capillary column (75 pm i.d. x 360 pm o.d., 20 cm in length, 3 pm particles, 300 A, Bio-C4, Sepax) and separated at a flow rate of 400 nE/min. A gradient composed of mobile phase A (2% ACN in water containing 0.1% FA) and mobile phase B (80% ACN with 0.1% FA) was used for separation.
- the gradient profile consisted of an 80 min program: 0-60 min, 20-100% B; 60-80 min, 100% B.
- the EC system required an additional 30 min for column equilibration and sample loading between runs, resulting in approximately 110 min per EC-MS run.
- the sample was run in triplicate.
- NP Characterization Dynamic light scattering (DES) and zeta potential analyses were performed to measure the size distribution and surface charge of the NPs before and after protein corona formation using a Zetasizer® Nano Series DES instrument (Malvern PanalyticalTM). A Helium Neon laser with a wavelength of 632 nm was used for the size distribution measurement at room temperature. Protein corona profiles at the surface of the NPs were studied by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) analysis. Transmission electron microscopy (TEM) was carried out using a JEM-2200FS (JEOL® Ltd.) operated at 200 kV.
- TEM Transmission electron microscopy
- the instrument was equipped with an in-column energy filter and an Oxford InstrumentsTM X-ray energy-dispersive spectroscopy (EDS) system.
- EDS Oxford InstrumentsTM X-ray energy-dispersive spectroscopy
- a total of 20 pL of the bare PSNPs was deposited onto a copper grid and used for imaging.
- 20 pL of the sample was negatively stained using 20 pL of uranyl acetate (1%) and finally washed with deionized (DI) water, deposited onto a copper grid, and used for imaging on the same day.
- DI deionized
- proteoform identification and quantification were performed using the TopPIC (top-down mass spectrometry-based proteoform identification and characterization) pipeline.
- TopPIC top-down mass spectrometry-based proteoform identification and characterization
- RAW files were converted to mzML files using the Msconvert tool.
- the spectral deconvolution that converted precursor and fragment isotope clusters into the monoisotopic masses and proteoform features was then performed using TopFD (top-down mass spectrometry feature detection, version 1.6.3).
- TopFD top-down mass spectrometry feature detection, version 1.6.3
- the resulting mass spectra and proteoform feature information were stored in msalign and text files, respectively.
- the database search was performed using TopPIC (version 1.6.3) against a home-built protein database ( ⁇ 1,000 protein sequences), in which the proteins identified in the BUP data in this work and in literature were included.
- the maximum number of unexpected mass shifts was one.
- the mass error tolerance for precursors and fragments was 50 ppm.
- To estimate FDRs of proteoform identifications the target-decoy approach was used, and proteoform identifications were filtered by a 1 and 5% FDR at the PrSM level and proteoform level, respectively.
- the lists of identified proteoforms (TDP) and protein groups (BUP) from all runs are shown in the Supporting Information of SADEGHI, S.A.
- FIG. 1 illustrates a detailed example TDP workflow.
- PSNPs polystyrene nanoparticles
- CZE capillary zone electrophoresis
- FIG. 1 illustrates a detailed example TDP workflow.
- PSNPs were incubated with a human plasma sample to form protein corona.
- the protein corona was then eluted from the PSNP surface using an elution buffer containing sodium dodecyl sulfate (SDS).
- SDS sodium dodecyl sulfate
- the eluted protein corona sample was buffer-exchanged to an MS-compatible buffer (100 mM ammonium bicarbonate, pH 8), followed by dynamic pH junction-based CZE-MS/MS analysis using an Orbitrap ExplorisTM 480 mass spectrometer (Thermo ScientificTM).
- MS-compatible buffer 100 mM ammonium bicarbonate, pH 8
- dynamic pH junction-based CZE-MS/MS analysis using an Orbitrap ExplorisTM 480 mass spectrometer (Thermo ScientificTM).
- Both “high-high” and “low-high” MS modes were used to measure proteoforms in the protein corona.
- proteoform parent ions (MS) were detected using a 480,000 resolution (at m/z 200) and fragment ions (MS/MS) were measured using a 120,000 resolution (at m/z 200).
- the “high-high” mode was mainly used to identify proteoforms smaller than 30 kDa.
- proteoform parent ions (MS) were detected using a 7,500 resolution (at m/z 200) and fragment ions (MS/MS) were still measured using a 120,000 resolution (at m/z 200).
- Protein corona was formed on the surface of PSNPs after incubating the PSNPs and human plasma sample.
- the particle size increased from 78.9 ⁇ 0 nm to 105.3 ⁇ 3.8 nm after the formation of protein corona based on the dynamic light scattering (DLS) measurement.
- the particle size distribution was reasonably narrow according to the standard deviation (SD) and polydispersity index (PDI) data.
- SD standard deviation
- PDI polydispersity index
- the protein corona was then eluted from the surface of PSNPs using a 0.4% SDS buffer for 1.5 hours at 60°C initially.
- the eluted protein corona sample was measured by SDS-PAGE and CZE-MS/MS (FIG. 3, FIG. 4, and FIG. 5). Three samples were prepared in parallel starting from the PSNP and human plasma incubation for evaluating the overall reproducibility of the TDP technique.
- the SDS-PAGE data shows consistent proteoform profiles across the three samples with strong bands at around 25 kDa and between 50-75 kDa.
- the CZE-MS/MS data (“high-high” mode) also show consistent total ion current (TIC) electropherograms, the number of proteoform identifications (98+13), and the number of proteoform- spectrum matches (PrSMs, 650+26) across the three samples (FIG. 3).
- proteoform intensity was obtained using the TopDiff software (LUBECKYJ, R. A. et al., (2019) Large-Scale Qualitative and Quantitative Top-Down Proteomics Using Capillary Zone Electrophoresis- Electrospray Ionization-Tandem Mass Spectrometry with Nanograms of Proteome Samples. J. Am. Soc. Mass Spectrom. 30:1435-1445) from the “high-high” mode data. The shared proteoforms among any two samples were utilized for the analysis.
- the intensities of the shared proteoforms from technical triplicate measurements were averaged and used to generate the plots.
- the proteoform intensity heatmap in FIG. 5 further illustrates the reproducibility of this TDP technique in measuring proteoforms (3-70 kDa) across three protein corona samples prepared in parallel.
- the CZE-MS/MS in “high-high” mode only enabled the identification of proteoforms around 10 kDa or smaller.
- the proteoforms larger than 28 kDa in FIG. 5 are from the “low-high” mode.
- the theoretical mass of HSA is 66,438 Da considering the 17 disulfide bonds (native form).
- the most abundant proteoform (66,560 Da) represents a +122-Da mass shift compared to the native form, presumably due to a combination of one phosphorylation (+80 Da) and one acetylation (+42 Da) or cysteinylation (+119 Da).
- the other two proteoforms of HSA most likely have additional glycosylation PTMs according to literature data and the information in the UniProt knowledgebase (www.uniprot.org/uniprotkb/P02768/entry).
- the 266-Da mass difference could be due to lipidation, for example, adding a stearic acid (octadecanoic acid, 284 Da) molecule through reaction between the carboxyl group of lipid and amine group of protein via removing a H2O molecule.
- the five different proteoforms could represent APO Al with 0-4 stearic acid modification sites.
- One previous TDP study provided strong evidence that APOA1 proteoforms are closely related to the indices of cardiometabolic health.
- APOA1 is a prognostic marker in renal and liver cancers according to the human protein atlas (www.proteinatlas.org/ENSGOOOOOH8137-APOAl).
- the data here suggests that coupling PSNP-based protein corona and MS-based TDP could be a valuable strategy for discovering novel APOA1 proteoform biomarkers of diseases (i.e., cancers and cardiometabolic diseases) using human plasma samples.
- the protein For the 51-kDa protein, two proteoforms were detected with 51,200 Da and 51,860 Da. Based on the mass and the BUP data of the protein corona sample, the protein might be Fibrinogen beta chain (50,763 Da in the mature form without PTMs) or Clusterin (50,062 Da in the mature form without PTMs). The MS/MS spectra of those two detected proteoforms do not match well with theoretical b- or y-type of fragment ions from Fibrinogen beta chain or Clusterin. More additional studies may be conducted in the future to provide more information about the identities of those two proteoforms.
- the charge state distributions may have slight changes between the two conditions due to the significant conformational differences of proteoforms.
- the masses of proteoforms between the two conditions are consistent with only ⁇ 1 Da difference, most likely due to potential errors from MS measurement using a Q-TOF mass spectrometer in this experiment.
- proteoform changes e.g., artificial modifications
- the possibility of proteoform changes (e.g., artificial modifications) in the protein corona is low during sample preparation. More systematic investigations can be performed on this topic using a complex system in the future to make additional conclusions.
- the protein corona samples were analyzed by both BUP and TDP using CZE-MS/MS.
- BUP the protein corona on PSNPs was prepared through the on-bead digestion procedure in BLUME, J. E. et al. Rapid, deep and precise profiling of the plasma proteome with multi- nanoparticle protein corona. Nat. Commun. 2020, 11, 3662 and ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities . Nat. Commun, 13:6610.
- TDP the protein corona was eluted from PSNPs and cleaned prior to CZE-MS/MS analysis, as shown in FIG. 1. Three protein corona samples prepared in parallel starting from the same human plasma sample were analyzed.
- FIGs. 8-12 summarize the protein corona data from TDP (FIG. 8) and BUP (FIG. 9).
- Single-shot TDP analysis of the protein corona sample by CZE-MS/MS consistently identified nearly 600 PrSMs, 100 proteoforms, and 20 proteoform families across the three protein corona samples (FIG. 8).
- the proteoform family represents a set of proteoforms from the same gene. In total, 263 proteoforms and 50 proteoform families were identified from the three corona samples.
- Single-shot BUP analysis of the protein corona sample by CZE-MS/MS identified around 8,000 peptide- spectrum matches (PSMs), 3,000 peptides, and 200 protein groups with high reproducibility (FIG. 9).
- a protein group can contain multiple proteins sharing the same set of identified peptides and those proteins usually have similar protein sequences.
- BUP analyses of the three protein corona samples using CZE-MS/MS identified, in total, 280 protein groups and 5,606 peptid
- TDP produced a much lower proteome coverage than BUP in terms of the number of identified genes (50 vs. -280) due to its lower sensitivity, resulting from much wider charge state distributions of intact proteoforms compared to peptides after electrospray ionization.
- BSA bovine serum albumin
- TDP offered more advanced measurements of proteins in a proteoform- specific manner.
- 18 different proteoforms of the gene Serum amyloid A-l (SAA I) were identified by CZE-MS/MS-based TDP (“high-high” mode) from the protein corona samples.
- the SAA1 protein is a prognostic marker of renal cancer according to the Human Protein Atlas (www.proteinatlas.org/ENSG00000173432-SAAl).
- Four of the identified SAA1 proteoforms are shown in FIG. 10. They have varied length of sequences due to variations in signal peptide cleavage.
- proteoform 1 has the whole protein sequence without signal peptide cleavage and proteoforms 2-4 have signal peptide cleavage at slightly different positions.
- various PTMs occurred on those proteoforms.
- proteoform 1 carries a N-terminal acetylation and one +463-Da mass shift close to the N-terminus;
- proteoform 2 has a roughly 14-Da mass shift, corresponding to a methylation PTM;
- proteoform 3 contains a 16-Da mass shift, corresponding to an oxidation modification; proteoform 4 does not have any PTMs.
- the location of those PTMs on the proteoforms is at the underlined regions.
- proteoforms were identified with high confidence evidenced by the E- Value and the number of matched fragment ions and were reasonably well characterized according to the five-level classification system described in SMITH, E. M. et al. (2019) A five-level classification system for proteoform identifications. Nat. Methods, 16:939-940.
- proteoform 4 is level- 1 identification
- proteoforms 2 and 3 belong to level-2a identifications
- proteoform 1 is a level-3 proteoform.
- the TDP approach also enabled the determination of relative abundance of proteoforms from the same gene. For example, proteoform 1 had a much lower abundance than the other three proteoforms, suggested by their intensities.
- the BUP measurement identified a protein group SAA1, containing five proteins, which have similar protein sequences.
- the main issue of the BUP data is that the protein identification is ambiguous, which means which protein(s) in the protein group exist in the protein corona sample is unresolved due to the “peptide-to-protein” inference problem.
- proteoforms and 73 proteoform families were identified with on average about 10 proteoforms per family.
- the proteoform overlap between CZE and RPLC as well as CZE-FAIMS and RPLC is low (FIG. 12) suggesting the nice complementarity between CZE- MS/MS and RPLC-MS/MS for proteoform identifications.
- CZE-FAIMS-MS/MS only less than 40% of the proteoforms from CZE-MS/MS were covered by CZE-FAIMS-MS/MS, indicating some potential proteoform loss in the FAIMS interface.
- the identified proteoforms were reasonably well characterized.
- proteoforms have no PTMs and no additional unexplained mass shifts and they are well characterized, belonging to the level 1 identifications.
- About 130 proteoforms were identified with specific PTMs (i.e., methylation, acetylation, phosphorylation, and oxidation) with or without an accurate localization, belonging to either level 1 or 2a identifications.
- PTMs i.e., methylation, acetylation, phosphorylation, and oxidation
- TDP identified specific proteoforms of 48 protein biomarkers (Table 1). The number of proteoforms ranged from 1 to 131 for those protein biomarkers.
- the BUP data of those protein biomarkers showed that each protein group can have a range of 1-16 proteins with an average of nearly 4 proteins per protein group.
- ten protein biomarkers were identified by TDP but not by BUP, most likely due to potential sample loss during the sample preparation of BUP.
- the data agrees with Xu, T. et al. (2020) Automated. Capillary Isoelectric Focusing-Tandem Mass Spectrometry for Qualitative and Quantitative Top-Down Proteomics. Anal. Chem, 92:15890-15898, on comparing BUP and TDP for quantitative proteomics of zebrafish brains. The result here indicates the valuable and more advanced information that TDP can offer on protein biomarkers of human plasma.
- Table 1 Summary of disease-related protein biomarkers identified by TDP and BUP.
- the protein biomarkers were determined according to the information in the Human Protein Atlas (www.proteinatlas.org/) except proteins labelled by *, which were determined based on literature data. N/A represents not identified.
- EXAMPLE 2 cIEF-MS characterization of proteins in protein corona cIEF-MS-based TDP workflow for the characterization of protein corona
- PSNPs were used due to experience in changing the parameters involved in the formation of a pure protein corona, ensuring highly accurate and reproducible MS results.
- Full details on PSNP parameter changing and characterization for protein corona formation are available in ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities. Nat. Commun, 13:6610; SADEGHI, S.A. et al., (2024) Mass Spectrometry-Based Top-Down Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona. ACS Nano. 18:26024- 26036; ASHKARRAN, A. A. et al.
- PSNPs were incubated with healthy human plasma to form protein coronas. After washing with PBS, the protein corona was eluted from PSNPs using a 0.4% (w/v) SDS solution, followed by buffer exchange to a 100 mM NH4HCO3 buffer for cIEF-MS/MS.
- ampholyte concentration for cIEF-MS/MS was determined. Higher ampholyte concentration can achieve better separation resolution but also can lead to unavoidable ionization suppression of proteoforms. Three concentrations of ampholytes, 1.5%, 1%, and 0.5%, were studied using a standard protein mixture containing cytochrome c (cyt c, pl 10.8), myoglobulin (Mb, pl 6.9) and carbonic anhydrase (CAs, pl 5.4). Automated cIEF-MS was carried out using the sandwich injection approach, the electrokinetically pumped sheath flow CE-MS interface, and an AGILENT® 6545XT Q-TOF mass spectrometer.
- CIEF with a higher concentration of ampholyte could reach a better separation resolution.
- CIEF with a higher ampholyte concentration tends to need a longer analysis time due to the higher buffering capacity of ampholytes, requiring a longer time for titration.
- cIEF-MS with 0.5% ampholytes was employed for the analysis of protein coronas.
- proteoforms based on MS/MS were coupled to an Orbitrap ExplorisTM 480 mass spectrometer (Thermo ScientificTM).
- One protein corona sample was analyzed in technical duplicates by a high-high mode, employing high mass resolution for both MSI and MS2.
- proteoforms and 31 proteins were identified.
- the two runs shared 43 proteoforms, representing nearly 70% of the number of identified proteoforms in one run (FIG. 15).
- proteoform 1 is 28091.238 Da and has one N-terminal acetylation and one 157.947 Da mass shift (FIG. 17).
- the S and T amino acid residues in this specific amino acid sequence can be phosphorylated.
- the deconvoluted MS/MS spectrum of the proteoform shows clear signals of ions corresponding to losses of H2O and H3PO4. Therefore, the 157.947 Da mass shift should correspond to two phosphorylation events.
- Proteoform 1 belongs to the level 2A identification.
- Proteoform 2 is 22519.954 Da and has N- terminal truncation and a 144.354 Da mass shift between position 195 and 232. Multiple acetylation (i.e., K) and phosphorylation (i.e., S or T) could happen in this region. The 144.354 Da may be from the combination of phosphorylation, acetylation, and other PTMs.
- Proteoform 3 is 18431.319 Da and has N-terminal truncation and one 264.751 Da mass shift.
- Proteoforms 2 and 3 are level 3 identifications.
- the mass errors of matched fragment ions of the three APO Al proteoforms are smaller than 10 ppm, and for most fragment ions, especially proteoforms 2 and 3, the mass error is close to 0.
- the high mass accuracy of matched fragment ions ensures the high confidence of identifications.
- This cIEF-MS/MS-based TDP could measure diverse proteoforms of the same gene (i.e., APOA1) in the protein corona. This technique could provide a relative abundance of proteoforms from the same gene. For example, proteoform 1 of gene APOA1 has a substantially higher abundance than others, evidenced by its much higher intensity (2E10 vs. ⁇ 5E6).
- APOA1 is a prognostic marker of cancer (www.proteinatlas.org/). 12 proteoforms of APO Al were identified. Overall, over 70 proteoforms of 16 cancer-related genes were identified.
- cIEF-MS/MS-based TDP provides an advanced view of the diverse proteoforms in the protein corona, including variations such as truncations and PTMs, as well as their combinations.
- This proteoform-centric TDP approach has the potential to offer more detailed and accurate information about protein corona composition compared to the traditional peptide-centric BUP. This enhanced accuracy is fundamental for developing and improving safer and more efficient nanomedicines.
- the data also implies that TDP profiling of protein corona could be useful for discovering novel proteoform biomarkers of diseases, e.g., cancers.
- proteoforms identified in this study using the high-high mode are smaller than 10 kDa (FIG. 18).
- the other 20% of the proteoforms are in the mass range of 11-30 kDa.
- TDP To improve the measurement quality of large proteoforms, a low-high approach was employed, utilizing low- resolution MSI and high-resolution MS2. Twenty-four proteoforms were detected close to or larger than 28 kDa from 4 proteins (FIG. 19).
- proteoforms were detected from protein 1 (466 kDa) and 2 proteoforms from protein 4 (443 kDa) (FIG. 19).
- protein 1 should be human serum albumin (HSA).
- CZE-MS/MS detected three HSA proteoforms and, here, cIEF-MS/MS observed nine HSA proteoforms in a mass range of 66436-67 625 Da, and the 66 820 Da proteoform is the most abundant one.
- the theoretical mass of HSA with 17 disulfide bonds (native form) is 66 438 Da.
- the smallest HSA proteoform detected here should be the native form.
- HSA can be modified by various PTMs, e.g., phosphorylation and glycosylation.
- the HSA proteoforms detected here must be due to the combinations of PTMs and/or sequence variations.
- cIEF-MS/MS detected two proteoforms of protein 4 (about 43 kDa), not observed in the CZE-MS/MS study in EXAMPLE 1.
- cIEF separated it into two peaks (2 and 20), and each peak has two proteoforms.
- the EXAMPLE 1 CZE-MS/MS study only detected the two highly abundant proteoforms of protein 2 (51 200 and 51 860 Da) in one peak.
- HPLC-grade acetic acid (AA), MS-grade water, methanol (MeOH), formic acid (FA), Amicon® Ultra (0.5 mF, 10 kDa cut-off size) centrifugal filter units, and fused silica capillaries (50 pm i.d./360 pm o.d., Polymicro Technologies) were purchased from Fisher ScientificTM (Pittsburgh, PA).
- Acrylamide was purchased from Acros OrganicsTM (Fair Lawn, NJ).
- a healthy human plasma sample was purchased from innovative ResearchTM (www.innov- research.com) and diluted to 55% using phosphate buffer solution (PBS, IX).
- Polystyrene NPs (PSNPs, -100 nm) were obtained from Polysciences® (www.polysciences.com).
- Sample preparation and characterization Briefly, PSNPs were mixed with 55% human plasma. This mixture was stirred constantly for one hour at a temperature of 37 °C to allow the formation of a protein corona. After an hour, the protein-NP complexes were separated by centrifugation at 14,000 xg for 20 minutes to remove unbound proteins. The resulting pellet was then washed twice with cold PBS.
- DLS Dynamic Light Scattering
- the proteins were extracted from the NP surface by incubating the pellet in a 0.4% SDS solution with agitation for 1.5 hours at 60°C, and the extracted protein corona-containing supernatant was separated by centrifugation. An Amicon® Ultra centrifugal filter with a lOkDa molecular weight cutoff was used to exchange the buffer and remove the SDS. Finally, the protein corona sample in 100 mM ammonium bicarbonate (NH4HCO3) was measured using a BCA assay to determine the protein concentration, and it was adjusted to 1.5 mg/mL for MS analysis.
- NH4HCO3 ammonium bicarbonate
- cIEF-MS/MS analysis An automated cIEF-MS/MS system was built by combining a CESI 8000 Plus CE system (Beckman Coulter) with an Orbitrap ExplorisTM 480 mass spectrometer (Thermo Fisher ScientificTM) using an in-house electrokinetically pumped sheathflow CE-MS nanospray interface.
- the cIEF separation was carried out using an 80 cm long linear polyacrylamide (LPA)-coated capillary (50 pm i.d./360 pm o.d.).
- LPA linear polyacrylamide
- One end of the separation capillary was etched using hydrofluoric acid to reduce its outer diameter to approximately 100 pm.
- the interface featured a glass spray emitter with an orifice size of 30-35 pm, filled with sheath buffer composed of 0.2% (v/v) formic acid and 10% (v/v) methanol.
- the spray voltage was set to 2 kV, and the capillary outlet to emitter orifice distance was maintained at approximately 0.5 mm.
- the distance between the emitter orifice and MS inlet was about 2 mm.
- the automated cIEF-MS system was based on the "sandwich" injection approach.
- the injection sequence involved three steps: first, a 6 cm catholyte plug was injected at 10 psi for 8 seconds containing 0.3% NH 4 OH, followed by a 20 cm mixture of sample and ampholyte plug containing 0.6% ampholytes (3-10, 5-8, and 8-10.5, GE HEALTHCARETM), injected at 10 psi for 27 seconds.
- Approximately 600 ng of corona proteins (1.5 mg/mL, injection volume of 400 nL) were loaded into the capillary, and finally, a 50 cm anolyte plug was injected at 10 psi for 67 seconds containing 5% acetic acid. This combination provided efficient focusing and mobilization of the protein corona samples under a separation voltage of 30 kV.
- the Orbitrap ExplorisTM 480 mass spectrometer was used to analyze the proteoforms separated by cIEF in data-dependent acquisition (DDA) mode.
- DDA data-dependent acquisition
- Two approaches were used for data acquisition to detect both small and large (>30 kDa) proteoforms.
- a “high-resolution MSI and high-resolution MS/MS” mode i.e., “High-High” mode was employed.
- the detailed parameters for the “High-High” mode include MSI resolution 480,000 at m/z 200 with a single microscan across a m/z range of 700-3000.
- Maximum ion injection time was set to 50 ms for MS and 100 ms for MS/MS.
- Normalized AGC target 300% Ions with an intensity of over 1E4 and charge states varying from 5 to 60 were isolated with a 2 m/z window, followed by fragmentation through higher-energy collision dissociation (HCD) at 25% normalized collision energy (NCE). Dynamic exclusion was enabled with a duration of 30 seconds and a mass tolerance of 10 ppm, and isotope exclusion was activated. The fragment ions were detected with a resolution of 120,000 at m/z 200 and normalized AGC 100%. For large proteoforms (>30 kDa), a “low-resolution MSI and high-resolution MS/MS” mode, i.e., “Low- High” mode, was employed. MSI resolution of 7,500 at m/z 200 was used. The microscan setting is 3. The other parameters are the same as the “High-High” mode.
- TopPIC version 1.7.0
- BUP bottom-up proteomics
- TopPIC was configured to accommodate a single unexpected mass shift per proteoform with a maximum shift of 500 Da and maintained a mass error tolerance of 50 ppm for both precursor and fragment ions.
- a target-decoy approach was used to estimate and control the false discovery rate (FDR), setting it at 1% at the proteoform- spectrum match (PrSM) level and 5% at the proteoform level.
- FDR false discovery rate
- PrSM proteoform- spectrum match
- the quantification aggregated the intensities of each proteoform's peaks across all scans and charge states.
- the raw mass spectrometry data files were processed using XcaliburTM Qual Browser (Thermo Fisher ScientificTM) to extract proteoform intensity values and migration time information.
- Base peak chromatograms and extracted ion chromatograms were generated to visualize the separation profiles.
- the electropherograms are graphically refined using Adobe Illustrator® for figure preparation.
- a high-throughput automated cIEF-MS/MS technique was developed that took 30 minutes or less per run for TDP of NP protein coronas (FIGs. 20-22).
- the protein corona sample was prepared according to Sun, L. et al. (2013) Ultrasensitive and Fast Bottom-up Analysis of Femtogram Amounts of Complex Proteome Digests. Angewandte Chemie - International Edition, 52(51): 13661-13664. Briefly, proteoforms in the protein corona of PSNPs were eluted using a 0.4% SDS buffer and cleaned up by buffer exchange, followed by cIEF-MS/MS (FIG. 20).
- the protein coronas were uniformly formed on the surface of PSNPs evidenced by the dark shells of the particles (FIG. 21).
- the advanced cIEF-MS/MS technique for high-throughput TDP analysis of protein coronas was carried out by employing a short separation capillary for cIEF-MS with a commercial CE system (FIG. 22).
- An 80-cm long LPA-coated capillary was used, and the effective capillary length for cIEF separation was shorter than 30 cm because a “sandwich” injection approach was used.
- the “sandwich” method includes injecting a plug of catholyte (0.3% NH4OH, pH ⁇ 11), a plug of the sample with ampholyte in 100 mM NH4HCO3, and a long plug of anolyte (5% acetic acid, pH 2.4).
- a 6 cm plug of catholyte, a 20 cm sample plug containing 0.6% ampholytes (pl 3-10, 5-8, and 8-10.5 with ratios 1:1:1), and a 50 cm plug of anolyte using a standard protein mixture were used. Because of the short effective capillary length for cIEF ( ⁇ 30 cm), the analysis could be carried out in a high-throughput fashion. Also, because the total capillary length was 80 cm, a regular commercial CE system could be used, allowing the technique to be adopted easily by other researchers.
- the protein corona sample of PSNPs was analyzed using the high-throughput cIEF- MS/MS technique for 50 runs (FIG. 23). Each cIEF-MS run took less than 30 minutes, producing a 2-6-fold improvement in analysis throughput compared to the previous cIEF- MS/MS-based TDP studies. Twenty-five runs were performed in "high-high” mode and twenty- five runs in "low-high” mode to evaluate the technique for both small and large proteoform measurements.
- the cIEF-MS/MS technique produced reproducible separation, detection, and identification of proteoforms.
- the electropherograms in FIG. 23 show consistent separation profiles of proteoforms in both “high-high” and “low-high” modes.
- RSD relative standard deviation
- the NL 8.2+0.8E09, corresponding to an RSD of about 10%.
- PrSM proteoform- spectrum match
- In the “low-high” runs three large proteins (28 kDa, 51 kDa, and 66 kDa) with multiple proteoforms per protein were consistently detected. Those three proteins correspond to human serum albumin (HSA, 66 kDa), Apolipoprotein A-I (APOA1, 28 kDa), and an unknown protein (51 kDa). The data agreed reasonably with the results in EXAMPLES 1 and 2.
- FIG. 25 shows the mass distribution of identified proteoforms from all the “high-high” runs.
- the mass of identified proteoforms ranged from ⁇ 2 kDa to ⁇ 30 kDa, and the majority of them were ⁇ 10 kDa or smaller. If the large proteoforms detected in “low-high” mode are included, the mass range of identified proteoforms will be extended to 2-66 kDa.
- Table 2 Summary of migration time of seven selected proteoforms from seven proteins across 25 “High-High” runs.
- the TDP analysis of protein corona identified 53 genes, and the number of detected proteoforms per gene ranged from 1 to 102 (Table 3). 33 out of the 53 genes are biomarkers, and they span various protein families and functional classes, including but not limited to apolipoproteins, complement proteins, immunoglobulins, and cytoskeletal proteins. Many of these proteins are associated with diverse diseases and pathological states, underscoring their potential utility as diagnostic or prognostic biomarkers. Particularly noteworthy is the prominence of the apolipoprotein family within the dataset, which includes APOA1, APOA2, APOA4, APOB, AP0C2, AP0C3, APOE, and APOF.
- proteoforms were identified for most of the apolipoprotein family members. For example, 102 proteoforms of the APO Al gene were identified, and four of them are shown in FIG. 26. Those proteoforms carry variations due to signal peptide cleavages, truncations, and PTMs. Proteoform 1 has an N-terminal cleavage of the first 26 amino acid residues, most likely corresponding to the signal peptide cleavage.
- Proteoform 1 also contains a mass shift of +288.535 Da in the highlighted region. Based on the PTM information in the dbPTM database (awi.cuhk.edu.cn/dbPTM/), three lysine residues in the highlighted region can be acetylated, corresponding to a +126 Da mass shift. The +288.535 Da may correspond to the combination of acetylation and other PTMs.
- Proteoform 2 the first 70 amino acid residues were truncated, and it carries a mass shift of +340.875 Da.
- Proteoform 3 shows a truncation of the first 127 amino acid residues at the N-terminus.
- proteoform 4 also contains an unknown mass shift of +59.054 Da in the highlighted region. The exact nature of this modification requires further investigation.
- Proteoform 4 exhibits an N-terminal removal of the first 24 amino acid residues due to the signal peptide cleavage and a C-terminal truncation. This proteoform also contains an unknown mass shift of +143.988 Da in the highlighted region.
- the dataset identifies biomarkers pertinent to inflammatory and autoimmune diseases, including complement proteins (C3, C9), serum amyloid A proteins (SAA1), and serpins (SERPINA1, SERPINC1).
- the catalog also highlights biomarkers associated with neurodegenerative conditions such as Alzheimer's disease (APOE, CEU) and Parkinson's disease (ABCB9), as well as proteins involved in cancer progression and metastasis, such as ACTB, KRT1, and RAB15.
- Table 3 Summary of the identified genes and corresponding number of proteoforms from the cIEF-MS/MS-based TDP analysis of nanoparticle protein coronas.
- EXAMPLE 4 Dissolving nanoparticles for capillary electrophoresis (CE) characterization of proteins in protein corona
- TDP top- down proteomics
- BUP bottom-up proteomics
- a significant challenge in applying TDP to analyze protein corona profiles is the difficulty in detaching intact proteins from the surface of nanoparticles without altering their structure or losing important post-translational modifications.
- a novel strategy that involves the dissolution of nanoparticles after the formation of the protein corona, thereby preserving the integrity of the entire protein corona for subsequent analysis. This approach allows researchers to examine the full array of proteoforms and proteins within the corona without the interference or structural complications caused by the presence of nanoparticles.
- nanoparticle removal is tailored to the specific type of nanoparticles being used. For instance, gold nanoparticles can be dissolved using a potassium-iodide/iodine etching solution, which effectively removes the nanoparticles while leaving the protein corona intact. Similarly, for polystyrene nanoparticles, ethyl acetate can be used as a solvent to dissolve the nanoparticles without affecting the associated proteins. This strategic dissolution of nanoparticles enables a more comprehensive and accurate analysis of the protein corona, facilitating a deeper proteomics analysis of protein corona and also understanding of nanoparticle interactions within biological systems and enhancing the overall diagnostic and therapeutic potential of nanomedicines.
- concentrations of potassium iodide/iodine, ethyl acetate, and ethylenediaminetetraacetic acid (EDTA) for digesting the nanoparticle core while preserving the protein corona can vary based on the type and amount of nanoparticles and the nature of the protein corona. Below are example ranges of concentration:
- Ethylenediaminetetraacetic Acid EDTA
- EDTA 0.5 mM to 5 mM
- a compatible buffer system with EDTA e.g., Tris-HCl or PBS
- EDTA e.g., Tris-HCl or PBS
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Engineering & Computer Science (AREA)
- Chemical & Material Sciences (AREA)
- Physics & Mathematics (AREA)
- Immunology (AREA)
- Hematology (AREA)
- Urology & Nephrology (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Analytical Chemistry (AREA)
- Pathology (AREA)
- Biochemistry (AREA)
- Biomedical Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Food Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biotechnology (AREA)
- Cell Biology (AREA)
- Microbiology (AREA)
- Biophysics (AREA)
- Medicinal Chemistry (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Electrochemistry (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Other Investigation Or Analysis Of Materials By Electrical Means (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Abstract
A method for characterizing proteins in a biological sample is provided herein that includes adding one or more protein-binding agents, having a surface capable of binding proteins, to the biological sample and allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample comprising the complex; separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample; and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS). Also provided herein are methods for detecting biomarkers and diagnosing a disease in a subject following the described characterization method.
Description
MASS SPECTROMETRY-BASED TOP-DOWN METHODS OF CHARACTERIZING PROTEINS IN PROTEIN CORONA
CROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63/568,036 filed on 21 March 2024 and U.S. Provisional Application No. 63/686,961 filed on 26 August 2024. The entire content of each patent application recited above is hereby incorporated by reference.
GOVERNMENT SUPPORT STATEMENT
[0002] This invention was made with government support under R01CA247863 awarded by the National Cancer Institute (NCI), DK131417 awarded by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK), R01GM125991 and R01GM118470 awarded by the National Institute of General Medical Sciences (NIGMS), and DBI1846913 awarded by the National Science Foundation. The government has certain rights in the invention.
REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0003] This application contains references to nucleic acid sequences and/or amino acid sequences which have been submitted concurrently herewith as the sequence listing .xml file entitled “6550-000499-WO-POA_4_Feb_2025_ST26. xml”, file size 16,623 Bytes (B), created on 4 February 2025. The aforementioned sequence listing is hereby incorporated by reference in its entirety.
FIELD
[0004] This disclosure generally relates to a method of characterizing proteins in a biological sample by adding protein-binding agents, such as nanoparticles (NPs), to the sample to generate protein coronas, separating the protein-binding agents from the protein coronas, and characterizing the proteins from the protein coronas using Capillary Electrophoresis (CE) and Mass Spectrometry (MS).
BACKGROUND
[0005] This section provides background information related to the present disclosure which is not necessarily prior art.
[0006] Nanomedicine applies nanotechnology concept to medicine, i.e., employing biocompatible nanoparticles (NPs) for controlled and/or targeted delivery of therapeutic (bio)molecules to desired tissues/organs, imaging, and disease diagnosis. The overall efficacy of nanomedicine is strongly impacted by protein/biomolecular corona, i.e., the composition and decoration of various types of biomolecules (e.g., mostly proteins) that bind to the surface of NPs after they are exposed to biological fluids (e.g., human plasma). It has been well documented that the composition and decoration of participated proteins in protein corona determines the biological fate and pharmacokinetics of NPs. Therefore, achieving comprehensive and accurate characterization of the composition of protein corona is central to advance nanomedicines’ safety together with their therapeutic and diagnostic efficacy. In addition, protein corona has also been recognized as a useful analytical technique to discover new protein biomarkers of diseases because it can reduce the complexity of biological fluids (e.g., plasma), which enables easier detection and identification of biomarkers.
[0007] Mass spectrometry (MS)-based bottom-up proteomics (BUP) has been widely recognized as an efficient way for the characterization of protein corona, providing the identification of gene products in the protein corona. However, MS-based BUP fails to identify exact forms of protein molecules (i.e., proteo forms) in the protein corona because of the “peptide- to-protein” inference problem. Proteoforms from the same gene due to sequence variations and post-translational modifications (PTMs) could have divergent biological functions and proteoforms are important for modulating disease progression. For example, strong evidence has been documented that modifications of a protein [i.e., human serum albumin (HSA)] changing its physicochemical properties and producing different proteoforms can substantially influence its binding affinity to NPs and the thickness of protein corona on the NPs, leading to significant changes in NP-cell interactions. Therefore, accurate measurement of proteins and proteoforms in protein corona is needed to provide a more accurate picture of protein corona, better the understanding of how protein corona directs the interactions between NPs and cells, and offer new opportunities for novel proteoform biomarker discovery.
SUMMARY
[0008] This section provides a general summary of the disclosure and is not a comprehensive disclosure of its full scope or all of its features.
[0009] In certain aspects, the present disclosure provides a method for characterizing proteins (including one or more of the protein’ s proteoforms) in a biological sample by adding one or more protein-binding agents to the biological sample, where the protein-binding agents have a surface
capable of binding protein, and allowing a protein corona to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one or more protein binding agents. The method also includes separating the complex from the biological sample to generate a complex sample including the complex, separating the one or more proteinbinding agents from the protein corona in the complex sample to generate a protein corona sample, and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
[0010] In some aspects, separating the complex from the biological sample may include one or more rounds of washing and centrifuging to remove a supernatant including the biological sample and any proteins not bound to or that have a low affinity for the one or more proteinbinding agents. In some embodiments, the method may include separating one or more proteinbinding agents from the protein corona in the complex sample by eluting the protein corona from the one or more protein-binding agents. In some embodiments, the elution buffer may contain a detergent, such as SDS, and/or an ionic liquid with 4-12 carbons in the hydrocarbon chain. In some embodiments, separating the one or more protein-binding agents from the protein corona may include dissolving the one or more protein-binding agents.
[0011] Further, in some aspects, the method includes buffer-exchanging the buffer in the protein corona sample to a mass spectrometry (MS)-compatible buffer prior to characterizing the proteins in the protein corona sample. In some embodiments, the CE may be Capillary Zone Electrophoresis (CZE) or Capillary Isoelectric Focusing (cIEF). In some embodiments, the protein-binding agents may be nanoparticles.
[0012] Additionally, the present disclosure provides a method for detecting one or more biomarkers or a pattern of one or more biomarkers associated with a disease following the method of characterizing proteins described herein, where the biological sample is obtained from the subject. The present disclosure also provides a method of diagnosing a disease in a subject, by following the characterization method described herein. In some embodiments, the disease may be cancer or a neurological disease.
[0013] Also provided herein are compositions including the protein corona sample produced by the method described herein, which may include the protein corona and substantially no protein-binding agent.
[0014] Further areas of applicability will become apparent from the description provided herein. The description and specific examples in this summary are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations and are not intended to limit the scope of the present disclosure.
[0016] FIG. 1 is an example schematic of the MS-based top-down proteomics (TDP) workflow for protein corona using polystyrene nanoparticles (PSNPs), a human plasma sample, and capillary zone electrophoresis (CZE)-tandem MS (MS/MS) (see, EXAMPLE 1).
[0017] FIG. 2 includes transmission electron microscopy (TEM) images of bare NPs, NPs after protein binding, and NPs after protein elution with 0.4% SDS (see, EXAMPLE 1).
[0018] FIG. 3 includes examples of total ion current (TIC) electropherograms of eluted protein corona after CZE-MS/MS analyses in “high-high” mode. Three protein corona samples (samples 1-3) were prepared in parallel and analyzed by CZE-MS/MS. Each sample was measured in triplicate (see, EXAMPLE 1).
[0019] FIG. 4 shows proteoform intensity correlations between any two samples. The data are from “high-high” mode. Log-log plots are shown in the figure (see, EXAMPLE 1).
[0020] FIG. 5 includes heatmaps of detected proteoform intensity across three samples. Proteoform intensity was log2 transformed and used to create the heat map using GraphPad Prism (see, EXAMPLE 1).
[0021] FIG. 6 shows extracted ion electropherograms (EIE) of three large proteins (1-3). The m/z ion with the highest intensity for each protein was used for the peak extraction with a mass tolerance of 200 ppm (see, EXAMPLE 1).
[0022] FIG. 7 shows the averaged mass spectrum of each protein across the peak and the corresponding deconvoluted masses of various proteoforms. The mass deconvolution was performed using the UniDec (Universal Deconvolution) software with default settings (see, EXAMPLE 1).
[0023] FIG. 8 is a bar graph representing numbers of proteoform identifications, proteoform- spectrum matches (PrSMs), and proteoform family identifications from TDP (see, EXAMPLE 1).
[0024] FIG. 9 is a bar graph representing numbers of peptide identifications, peptide- spectrum matches (PSMs), and protein group identifications from BUP (see, EXAMPLE 1).
[0025] FIG. 10 includes four example proteoforms of SAA1 identified by TDP with sequences and fragmentation patterns and the five proteins included in the protein group SAA1 identified by BUP (see, EXAMPLE 1).
[0026] FIG. 11 is a bar graph representing numbers of proteoform and proteoform families identified from the protein corona by using different workflows. The protein corona sample was from treating the protein corona-coated PSNPs by 1% SDS for 3 h at 60 °C (see, EXAMPLE 1).
[0027] FIG. 12 is a Venn diagram of identified proteoforms from the protein corona by three different methods (CZE-MS/MS, CZE-FAIMS-MS/MS, and RPLC-MS/MS). The protein corona sample was from treating the protein corona-coated PSNPs by 1% SDS for 3 h at 60 °C (see, EXAMPLE 1).
[0028] FIG. 13 is an example schematic workflow of cIEF-MS/MS-based TDP for NP protein corona. Polystyrene NPs (PSNPs) were used. The figure was created using BioRender™ and used here with permission (see, EXAMPLE 2).
[0029] FIG. 14 includes base peak electropherograms of duplicate cIEF-MS/MS runs (see, EXAMPLE 2).
[0030] FIG. 15 is a Venn diagram of proteoform overlaps between duplicate measurements (see, EXAMPLE 2).
[0031] FIG. 16 is a line graph showing proteoform intensity correlation between the duplicate runs. Log2 (proteoform intensity) was used, and the proteoforms having proteoform feature intensities in both runs were used (see, EXAMPLE 2).
[0032] FIG. 17 shows sequence and fragmentation pattern of one APOA1 proteoform (proteoform 1), having one N-terminal acetylation and diphosphorylation (see, EXAMPLE 2).
[0033] FIG. 18 is a bar graph representing mass distribution of proteoforms identified in the two replicate runs (see, EXAMPLE 2).
[0034] FIG. 19A shows a base peak electropherogram of protein corona by cIEF-MS/MS using an Orbitrap Exploris™ 480 mass spectrometer in low-high mode. Deconvoluted masses of detected large proteoforms of proteins 1 (FIG. 19C and FIG. 19D) and 4 (FIG. 19B) are shown. UniDec software was used for mass deconvolution with default settings (see, EXAMPLE 2).
[0035] FIG. 20 shows a brief example workflow of preparing the protein corona sample for TDP after incubating the PSNPs with a human plasma sample to form the protein corona on the surface of PSNPs (see, EXAMPLE 3).
[0036] FIG. 21 is a graph representing DLS analysis of bare PSNPs (uncoated) and protein corona-coated.NPs (protein corona) (see, EXAMPLE 3).
[0037] FIG. 22 includes an example schematic design of the high-throughput cIEF-MS/MS for protein corona analysis (see, EXAMPLE 3).
[0038] FIG. 23 shows total ion current (TIC) electropherograms of protein corona proteoforms by CIEF-MS/MS in “High-High (HH)” and “Low-High (LH)” modes. Six selected electropherograms from runs #4, #8, #14, #16, #20, and #23 in HH and LH modes are shown (see, EXAMPLE 3).
[0039] FIG. 24 includes intensity correlations of overlapped proteoforms between any two cIEF-MS/MS runs. Six runs were randomly selected for this analysis. Proteoform intensities were log2-transformed for the plot, and Pearson’s correlation coefficient (r) values were labeled (see, EXAMPLE 3).
[0040] FIG. 25 is a graph representing mass distribution of the identified proteoforms from 25 cIEF-MS/MS runs (High-High) (see, EXAMPLE 3).
[0041] FIG. 26 includes sequences and fragmentation patterns of four distinct proteoforms of Apolipoprotein A-I (APOA1) identified using cIEF-MS/MS-based TDP in "high-high" mode (see, EXAMPLE 3).
DETAILED DESCRIPTION
A. Introduction
[0042] Conventional mass spectrometry (MS)-based bottom-up proteomics (BUP) analysis of protein corona [i.e., an evolving layer of biomolecules, mostly proteins, formed on the surface of nanoparticles (NPs) during their interactions with biomolecular fluids] enabled the nanomedicine community to partly identify the biological identity of NPs. Such an approach, however, fails to pinpoint the specific proteoforms — distinct molecular variants of proteins, which is essential for prediction of the biological fate and pharmacokinetics of nanomedicines.
[0043] Recognizing this limitation, this disclosure describes a robust and reproducible MSbased top-down proteomics (TDP) technique for characterizing proteins (including one or more proteoforms) in the protein corona by adding protein-binding agents (e.g., NPs) to a biological sample to form a complex including the protein-binding agent and protein corona, separating the complex from the biological sample, separating the protein-binding agents from the protein corona to generate a protein corona sample, and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS). The present TDP
approach has successfully identified about 900 proteoforms in the protein corona of polystyrene NPs, ranging from 2-70 kDa, revealing proteoforms of 48 protein biomarkers with combinations of post-translational modifications, signal peptide cleavages, and/or truncations — details that bottom-up proteomics (BUP) could not fully discern. This advancement in MS-based TDP offers a more advanced approach to characterize NP protein coronas, deepening the understanding of NPs' biological identities.
[0044] Example embodiments are provided so that this disclosure will be thorough and will fully convey the scope to those who are skilled in the art. Numerous specific details are set forth such as examples of specific compositions, components, devices, and methods, to provide a thorough understanding of embodiments of the present disclosure. It will be apparent to those skilled in the art that specific details need not be employed, that example embodiments may be embodied in many different forms and that neither should be construed to limit the scope of the disclosure. In some example embodiments, well-known processes, well-known device structures, and well-known technologies are not described in detail.
B. Definitions
[0045] The terminology used herein is for the purpose of describing particular example embodiments only and is not intended to be limiting. As used herein, the singular forms “a,” “an,” and “the” may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms “comprises,” “comprising,” “including,” and “having,” are inclusive and therefore specify the presence of stated features, elements, compositions, steps, integers, operations, and/or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and/or groups thereof. Although the open-ended term “comprising,” is to be understood as a non-restrictive term used to describe and claim various embodiments set forth herein, in certain aspects, the term may alternatively be understood to instead be a more limiting and restrictive term, such as “consisting of’ or “consisting essentially of.” Thus, for any given embodiment reciting compositions, materials, components, elements, features, integers, operations, and/or process steps, the present disclosure also specifically includes embodiments consisting of, or consisting essentially of, such recited compositions, materials, components, elements, features, integers, operations, and/or process steps. In the case of “consisting of,” the alternative embodiment excludes any additional compositions, materials, components, elements, features, integers, operations, and/or process steps, while in the case of “consisting essentially of,” any additional compositions, materials, components, elements, features, integers, operations, and/or process steps that materially affect
the basic and novel characteristics are excluded from such an embodiment, but any compositions, materials, components, elements, features, integers, operations, and/or process steps that do not materially affect the basic and novel characteristics can be included in the embodiment.
[0046] Any method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order discussed or illustrated, unless specifically identified as an order of performance. It is also to be understood that additional or alternative steps may be employed, unless otherwise indicated.
[0047] The use of the term "a" or "an" when used in conjunction with the term "comprising" in the claims and/or the specification may mean "one," but it is also consistent with the meaning of "one or more," "at least one," and "one or more than one." As such, the terms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise. Thus, for example, reference to "a compound" may refer to one or more compounds, two or more compounds, three or more compounds, four or more compounds, or greater numbers of compounds.
[0048] The use of the term "at least one" will be understood to include one as well as any quantity more than one, including but not limited to, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, etc. The term "at least one" may extend up to 100 or 1000 or more, depending on the term to which it is attached; in addition, the quantities of 100/1000 are not to be considered limiting, as higher limits may also produce satisfactory results. In addition, the use of the term "at least one of X, Y, and Z" will be understood to include X alone, Y alone, and Z alone, as well as any combination of X, Y, and Z. The use of ordinal number terminology (i.e., "first," "second," "third," "fourth," etc.) is solely for the purpose of differentiating between two or more items and is not meant to imply any sequence or order or importance to one item over another or any order of addition, for example.
[0049] The use of the term "or" in the claims is used to mean an inclusive "and/or" unless explicitly indicated to refer to alternatives only or unless the alternatives are mutually exclusive. For example, a condition "A or B" is satisfied by any of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0050] As used herein, any reference to "one embodiment," "an embodiment," "some embodiments," "one example," "for example," or "an example" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearance of the phrase "in some embodiments" or "one example" in various places in the specification is not necessarily all referring to the same embodiment, for
example. Further, all references to one or more embodiments or examples are to be construed as non-limiting to the claims.
[0051] Throughout this disclosure, the term "about" is used to indicate that a value includes the inherent variation of error for a composition/apparatus/device, the method being employed to determine the value, or the variation that exists among the study subjects. For example, but not by way of limitation, when the term "about" is utilized, the designated value may vary by plus or minus twenty percent, or fifteen percent, or twelve percent, or eleven percent, or ten percent, or nine percent, or eight percent, or seven percent, or six percent, or five percent, or four percent, or three percent, or two percent, or one percent from the specified value, as such variations are appropriate to perform the disclosed methods and as understood by persons having ordinary skill in the art. Particularly in reference to a given quantity, number or percentage, “about” is meant to encompass deviations of plus or minus ten percent (± 10). For example, about 5% encompasses any value of 4.5% to 5.5%, such as 4.5, 4.6, 4.7, 4.8, 4.9, 5, 4.1, 5.2, 5.3, 5.4, or 5.5. Accordingly, unless otherwise indicated, the numerical parameters set forth in this specification and attached claims are approximations that can vary depending upon the desired properties sought to be obtained by the presently disclosed subject matter.
[0052] The term "or a combination thereof" as used herein refers to all permutations and combinations of the listed items preceding the term. For example, "A, B, C, or combinations thereof" is intended to include at least one of: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, CBA, BCA, ACB, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AAB, BBC, AAABCCCC, CBBAAA, CABABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context.
[0053] As will be understood by one skilled in the art, for any and all purpose, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Furthermore, as will be understood by one skilled in the art, a range includes each individual member.
[0054] A “protein” as used herein includes all proteoforms of the protein. A “proteoform” as used herein includes all different molecular forms in which the protein product of a single gene may be found. In some embodiments, proteoforms may be different forms of a protein produced from the genome with a variety of sequence variations, splice isoforms, post-translational
modifications, or a combination thereof. Thus, in some embodiments, referring to a protein may include referring to one or more proteoforms of the protein.
[0055] ‘Protein corona,” often referred to as “PC”, as used herein is a biomolecular shell that forms on the surface of nanoparticles (NPs) during their interactions with biological fluids, which changes over time. The term "biomolecule corona" refers to multiple different biomolecules that are able to bind to a protein-binding agent having a surface, such as a nanoparticle. The term "biomolecule corona" encompasses "protein corona" which is a term used in the art to refer to the proteins, lipids and other plasma components that bind a protein-binding agent, such as nanoparticles when they come into contact with biological fluids. For use herein, the term "protein corona" encompasses both the soft and hard protein corona as referred to in the art, see, e.g., Milani, et al., "Reversible versus Irreversible Binding of Transferrin to Polystyrene Nanoparticles: Soft and Hard Corona," ACS NANO, 2012, 6(3), pp. 2532-2541; Mirshafiee, et al., "Impact of protein pre-coating on the protein corona composition and nanoparticle cellular uptake," Biomaterials, vol. 75, Jan. 2016 pp. 295-304, Mahmoudi, et al., "Emerging understanding of the protein corona at the nano-bio interfaces," Nanotoday, 11(6) Dec. 2016, pp. 817-832, and Mahmoudi, et al., "Protein-Nanoparticle Interactions: Opportunities and Challenges," Chem. Rev., 2011, 111(9), pp. 5610-5637, the contents of which are incorporated by reference in their entireties. As described in the art, adsorption curve shows the build-up of a strongly bound monolayer up to the point of monolayer saturation (at a geometrically defined protein-to- nanoparticle ratio), beyond which a secondary, weakly bound layer is formed. While the first layer is irreversibly bound (hard corona), the secondary layer (soft corona) exhibits dynamic exchange. Proteins that adsorb with high affinity form what is known as the “hard” corona, consisting of tightly bound proteins that do not readily desorb, and proteins that adsorb with low affinity form the “soft” corona, consisting of loosely bound proteins. Soft and hard corona can also be defined based on their exchange times. Hard corona usually shows much larger exchange times in the order of several hours. See, e.g., M. Rahman, et al., Protein-Nanoparticle Interactions, Spring Series in Biophysics 15, 2013, incorporated by reference in its entirety.
[0056] A “protein-binding agent” as used herein is any agent that binds, interacts with, or attracts one or more protein in a biological sample, and has a surface for protein binding. In some embodiments, the protein-binding agent may be referred to as a protein-interacting agent, a protein- attracting agent, and/or a protein corona-forming agent. In some embodiments, the protein-binding agent may be a nanoscale material and/or a microscale material.
[0057] A “nanoscale material” or “nanomaterial” as used herein is a material of which a single unit is sized (in at least one dimension) between 1 nanometer (nm) and 999 nm. A non-limiting list of nanomaterials includes nanoparticles, nanorods, nanospheres, nanodisks, nanoclusters, nanofibers, and nanotubes.
[0058] A “microscale material” or “micromaterial” as used herein is a material of which a single unit is sized (in at least one dimension) between 1000 nm and 100,000 nm. A non-limiting list of micromaterials includes microparticles, microrods, microspheres, and microbeads.
[0059] Examples of suitable nanoscale and/or microscale materials include, but are not limited to, organic materials, non-organic materials, or combinations thereof. In some embodiments, the materials are micelles, liposomes, iron oxide, graphene, silica, protein-based materials, polystyrene, silver, and gold materials, such as colloidal gold, quantum dots, palladium, platinum, titanium, and combinations thereof. In some embodiments, nanoparticles are liposomes. One skilled in the art is able to select and prepare suitable nanoscale and/or microscale material(s).
[0060] A “small molecule” as used herein is a molecule having a molecular weight of less than 5 kilodaltons (kDa). For the purposes of this disclosure, molecular weight may be calculated by the following formula: Molecular weight (in the dalton unit (Da) or the unified atomic mass unit (u)) = ^((atomic mass of element)n x (# of atoms of that element)n). The small molecule can be synthetic or natural. In other words, the small molecule may be synthetically generated or it can be naturally produced. The small molecules described in the methods herein may also be referred to as protein-recruitment agents because they encourage recruitment of different proteins (and other types of biomolecules) to the surface of materials, such as nanoparticles. In some embodiments, the small molecule may also be referred to as a “high-abundance protein-binding agent” or a “high- abundance protein-interacting agent.”
[0061] “High-abundance protein” as used herein refers to one of the seven most abundant proteins (albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and fibrinogen) in human plasma. Such proteins collectively represent 85% of the total protein mass in human plasma.
[0062] ‘Low-abundance protein” as used herein refers to any proteins present in human plasma that are not included in the seven high- abundance proteins.
[0063] A “biomolecule” as used herein is a molecule produced by a living organism and essential to one or more biological processes. A biomolecule may also be synthetically generated to mimic the function of a naturally generated biomolecule. Examples of suitable biomolecules include, but are not limited to, macromolecules (such as proteins, carbohydrates, lipids, nucleic acids, etc.) as well as smaller molecules (such as vitamins, hormones, etc.). A biomolecule may
also be referred to herein as a biological material. Thus, a small molecule as used herein can be considered a “small biomolecule” if it fits the definition of a small molecule (a molecule having a molecular weight of less than 5 kilodaltons (kDa)) and the definition of a biomolecule (a molecule produced by a living organism and essential to one or more biological processes or a molecule synthetically generated to mimic the function of a naturally generated bio molecule).
[0064] A “metabolite” as used herein is an intermediate or end product of metabolism. It is intended to include all natural, synthetic, and biological small molecules, such as amino acids, alcohols, polyols, alkaloids, organic acids, sugars (e.g., glucose) as well as nucleotides (e.g., inosine-5'-monophosphate and guanosine-5'-monophosphate).
[0065] “Lipids” as used herein are a group of organic compounds including all natural, synthetic, and biological fatty compounds, such as glycerolipids (e.g., triacylglycerols also known as triglycerides, TG or TAG; diacylglycerols also known as diglycerides, DG or DAG, such as 1,2- diacylglycerols and 1,3-diacylglycerols), glycerophospholipids (e.g., phosphatidylcholine, phosphatidylethanolamine, L-a-phosphatidylinositol), sterol lipids, very-low-density lipoprotein (VLDL), low density lipoprotein (LDL), and high-density lipoprotein (HDL), fatty acids, prenol lipids, and sphingolipids.
[0066] ‘Nutrient” as used herein is a substance used by an organism to survive, grow, and reproduce. In some cases a metabolite can also be considered a nutrient, such as amino acids, fatty acids, vitamins (e.g., vitamin B complex), minerals and choline.
[0067] ‘Plant-derived molecule” as used herein refers to any molecule derived from a plant. Suitable plant-derived molecules include small molecules such as auxin, gibberellic acid, alkaloids, phenylpropanoids; plant-made biologies, such as anti-cancer biologies; and phytopharmaceutical drugs, such as those with properties against human health problems such as allergy, inflammation, etc.
[0068] The term “endogenous” as used herein refers to a substance, e.g. a nucleic acid, protein, enzyme, small molecule, etc., that is produced from within a host organism (e.g., a human) and/or that is naturally occurring or naturally found inside a host organism. Thus, an endogenous substance refers to a substance produced by and found inside a host organism. In some embodiments, an endogenous substance may be encoded by the genome of the host organism. In some embodiments, the substance may be encoded by an autonomously replicating plasmid carried by the host organism. In some embodiments, an endogenous substance is a substance that was present in a host organism’s biological sample when the biological sample was originally isolated from nature, i.e., the substance is native to the organism. For example, an “endogenously
produced” substance may be expressed/generated by a host organism’s own machinery. In other words, the host organism has not been genetically engineered to produce the substance. In other words, a host organism may endogenously produce a native or non-native substance.
[0069] In contrast, an “exogenous” substance, e.g., a nucleic acid, protein, enzyme, small molecule, etc., as used herein, refers to a substance that is not encoded by or produced by the host organism (human or non-human, such as a mouse, rat, pig, etc.), and which is therefore added to the host organism or biological sample taken from the host organism from outside of the host organism. For example, “exogenously added” may refer to adding a substance, such as a small molecule, to the host organism or a biological sample taken from the host organism. A nucleic acid sequence encoding a variant (i.e., mutant) polypeptide, when added to a host organism or host organism cell, is one example of an exogenous nucleic acid sequence. The exogenous nucleic acid sequence can encode a polypeptide or an enzyme that is also otherwise endogenous or native to the cell. Such an encoded polypeptide or enzyme can be considered “exogenously expressed.” For example, to achieve overexpression of an endogenous gene, additional copies of the gene can be introduced into the cell (e.g., in a vector, such as a plasmid). Such additional copies of the endogenous gene can be considered as “exogenous” (e.g., exogenous gene(s) or an exogenous nucleic acid sequence(s)), because the additional copies are introduced into the cell from outside the cell. An “exogenous gene” or “exogenous nucleic acid sequence” also refers to a native (or endogenous) gene or nucleic acid sequence that is deregulated (e.g., upregulated, downregulated, or attenuated) or otherwise altered or modified, for example, by operably linking it to a regulatory element. An exogenous nucleic acid sequence or exogenous gene can also be used to express or overexpress a heterologous polypeptide or enzyme in a cell. Thus, an exogenous nucleic acid sequence or an exogenous gene can encode a polypeptide (e.g., an enzyme) that is native to the cell, that is otherwise endogenous to the cell, or that is heterologous to the cell.
[0070] As used herein the term “native” refers to the form of a composition, such as a small molecule, that is isolated from nature, or to a composition that is in its natural state without intentionally introduced mutations in the structural sequence and/or without any engineered changes in expression such as e.g., changing a developmentally regulated gene to a constitutively expressed gene. As used herein, “native” also refers to “wildtype” or “wild-type,” in which the composition is present in both sequence, quantity, and relative quantity as typically found in the organism as naturally found. Wild-type organisms may serve as a control and/or reference for determination of cellular functions. A native molecule, e.g. a lipid, gene, nucleic acid sequence, polypeptide, or enzyme, for example, is typically endogenous to a cell, i.e., found in or produced by the cell. An exogenous nucleic acid sequence or an exogenous gene can encode a native
polypeptide or enzyme, for example, where additional copies of a native gene or nucleic acid sequence are added to the cell from outside the cell, or where a native gene or nucleic acid sequence is deregulated or altered, e.g., by operably coupling it to a regulatory element that is not native or endogenous to the cell.
[0071] The term “non-native” is used herein to refer to nucleic acid sequences, amino acid sequences, polypeptide sequences, enzymes, and/or small molecules that do not occur naturally in the host. Heterologous genes and polypeptides are considered “non-native.” A nucleic acid sequence or amino acid sequence that has been removed from a host cell, subjected to laboratory manipulation, and introduced or reintroduced into a host cell, is also considered “non-native.” Synthetic or partially synthetic genes introduced into a host cell are “non-native.” Non-native genes further include genes that are endogenous and/or native to the host microorganism but that are operably linked to one or more heterologous regulatory sequences that have been recombined into the host genome. A naturally occurring gene under the control of a heterologous regulatory sequence is considered “non-native.” In some embodiments, an organism comprising a non-native gene may be utilized as a control and/or reference for an organism having additional and/or different variations from wild-type organisms.
[0072] "Sample" as used herein refers to a biological sample or a complex biological sample obtained from a subject. Suitable biological samples include, but are not limited to, biological fluids, such as systemic blood, plasma, serum, lung lavage, cell lysates, menstrual blood, urine, processed tissue samples, amniotic fluid, cerebrospinal fluid, tears, saliva, semen and the like. In a particular embodiment, the sample is a whole blood, plasma or serum sample.
[0073] ‘Health spectrum” as used herein covers a broad range of health states, including complete well-being, minor health issues, chronic conditions, and severe illnesses. The health spectrum adopts a holistic and dynamic view of health, emphasizing prevention and the continuum of health states. The health spectrum is concerned with overall well-being and the factors that influence it, whether they lead to optimal health or contribute to illness. The health spectrum emphasizes preventive care, lifestyle modifications, and overall health maintenance. Whereas a “disease/disorder” focuses on specific pathological conditions affecting health. A disease/disorder adopts a more specific, diagnostic approach, focusing on identifying and treating particular conditions. A disease/disorder is focused on the pathological aspects of health, identifying specific illnesses or dysfunctions to address them effectively. Disease/disorder emphasizes medical intervention, treatment plans, and symptom management for specific conditions.
[0074] A disease/disorder is one example of a health spectrum condition. Other conditions include pre-diagnosis states where one may be at risk of developing a disease or disorder, but has not reached the level of a diagnosed disease yet.
[0075] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. In particular, this disclosure utilizes routine techniques in the field of nanoparticles and protein characterization.
C. Methods
Method for characterizing proteins in a biological sample
[0076] In one embodiment, the present disclosure provides a method for characterizing proteins (including one or more proteoforms of the protein) in a biological sample. The method includes adding one or more protein-binding agents, having a surface capable of binding proteins, to the biological sample and allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one or more protein-binding agents. The method also includes separating the complex from the biological sample to generate a complex sample including the complex and separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample. The method includes characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
[0077] In some embodiments, characterizing proteins may include determining one or more protein characteristics, such as but not limited to, detecting/identifying/quantifying proteins, type of proteins, post-translational modifications in proteins, proteins with signal peptide cleavage, truncations, etc., protein biomarkers, protein migration time, protein intensity, number of protein -spectrum matches, protein amino acid sequence, protein isoelectric point, and molecular weight of protein.
[0078] Also provided herein is a composition including a protein corona sample produced by the method described herein for characterizing proteins in a biological sample. Thus, provided herein is a composition prepared by adding one or more protein-binding agents, having a surface capable of binding proteins, to a biological sample and allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample comprising the complex; and separating the one or more protein-binding agents from the protein corona in the complex sample to generate the protein corona sample.
Biological sample
[0079] In some embodiments, a biological sample may include, but is not limited to, biological fluids, such as blood (whole blood), plasma, serum, lung lavage, a cell lysate, menstrual blood, urine, a tissue, amniotic fluid, cerebrospinal fluid, tears, a liquid biopsy, saliva, semen and the like. In a particular embodiment, the biological sample may be a blood, plasma, or serum sample. In another particular embodiment, the biological sample may be a plasma sample. In some embodiments, the biological sample may include more than one biological sample. For example, the biological sample includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, or more of the same or different biological samples.
[0080] In some embodiments, biological fluids or complex biological samples may be prepared by methods and kits known in the art. For example, some biological samples (e.g., menstrual blood, whole blood, semen, etc.) may first be centrifuged at low speed to remove cell debris, blood clots and other cellular components that may interfere with the methods described herein. In other embodiments, for example, tissue specimens may be processed, e.g., tissue samples may be minced or homogenized, treated with enzymes to break up the tissue and/or centrifuged to remove cellular debris allowing for the assaying and extraction of molecules within the tissue samples. Suitable methods of isolating and/or properly preparing and storing biological samples are known in the art, and may include, but are not limited to, the addition of an anticoagulant agent.
[0081] In some embodiments, the method for detecting proteins in a biological sample may include a step for depleting one or more proteins in a biological sample. For example, depleting one or more proteins in a biological sample may include running the biological sample through a depletion column or through a spin column with resin. In some embodiments, running the biological sample through a depletion column or spin column with resin may reduce the complexity of biological samples for analysis via antibody-based techniques or proteomics techniques. For example, the complexity of biological samples may be reduced for top-down proteomics analysis. In some embodiments, depletion columns or spin columns may be used to reduce the complexity of biological samples, such as serum, plasma, etc., which contain high concentrations of albumin and immunoglobulins. For example, depletion columns or spin columns may be used to remove highly abundant proteins, such as albumin and IgG, from biological samples. In some embodiments, a suitable depletion or spin column may be a High Select™ Depletion Spin Column (Thermo Scientific™). In some embodiments, protein depletion methods may be used for applications in drug delivery and/or imaging. In some embodiments, protein
depletion methods, such as those including depletion or spin columns, may include particles, such as lipid-based NPs, with a small molecule, such as phosphatidylcholine (PtdChos), on their surface. Thus, in some embodiments, proteins, such as albumin, may be attracted and/or bound to the surface of the small molecule-coated NPs, which may enhance the small molecule-coated NP’s blood circulation time and/or allow them to be removed quickly by the immune system.
Biomolecule detection
[0082] The method described herein may be used to detect biomolecules in a biological sample. Biomolecules that may be detected by the disclosed method may include small molecules (such as lipids, fatty acids, glycolipids, sterols, monosaccharides, vitamins, hormones, neurotransmitters, metabolites, etc.), monomers (such as amino acids, monosaccharides, isoprene, nucleotides, etc.), oligomers (such as oligopeptides, oligosaccharides, terpenes, oligonucleotides, etc.), and polymers (such as polypeptides, proteins and/or their proteoforms, polysaccharides, polyterpenes, polynucleotides, nucleic acids, etc.). In some embodiments, the method described herein may be used to detect proteins (including one or more proteoforms). In some embodiments, proteoforms of proteins may be different forms of a protein produced from the genome with a variety of biological variations, such as sequence variations, splice isoforms, post-translational modifications, etc., which may alter the primary sequence and composition at the whole-protein level. In some embodiments, proteoforms may carry different biological functions.
[0083] In some embodiments, the method described herein may be used to detect a number of unique biomolecules in a biological sample. For example, the method may be used to detect a number of different proteins and/or proteoforms in a biological sample. Thus, the method described herein may be used to detect one type of protein and/or the method may be used to detect more than one type of protein. In some embodiments, the method may be used to detect an amount of one type of protein present in a biological sample and/or present in a protein corona. In some embodiments, the method may be used to detect an amount of more than one type of protein present in a biological sample and/or present in a protein corona. In some embodiments, the method may be used to detect how many distinct proteins are present in a biological sample and/or are present in a protein corona.
Small molecules
[0084] In some embodiments, the method may include adding one or more small molecules along with the one or more protein-binding agents to the biological sample.
[0085] As described herein, “one or more small molecule” covers a “small molecule combination” (i.e., a combination of one or more small molecules). In some embodiments, the
method includes adding one or more small molecule or small molecule combination to the biological sample. In some embodiments, the method may include adding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. small molecules to the biological sample. In some embodiments, the method may include adding a combination of one or more small molecules to the biological sample. In some embodiments, the small molecule combination may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. different small molecule types. In some embodiments, the method may include adding one or more small molecule combination to the biological sample. In some embodiments, the method may include adding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, etc. small molecule combinations to the biological sample. In some embodiments, the number of small molecules or small molecule combinations added to the biological sample may be any number to reach a desired concentration of small molecules or small molecule combinations in the biological sample.
[0086] In some embodiments, the small molecule or small molecule combination may be added to the biological sample at a concentration of about Ipg/ml to about Ig/ml. For example, in some embodiments, the small molecule or small molecule combination may be added to the biological sample at a concentration of about 1 pg/ml, about 2 pg/ml, about 3 pg/ml, about 4 pg/ml, about 5 pg/ml, about 6 pg/ml, about 7 pg/ml, about 8 pg/ml, about 9 pg/ml, about 10 pg/ml, about 20 pg/ml, about 30 pg/ml, about 40 pg/ml, about 50 pg/ml, about 60 pg/ml, about 70 pg/ml, about 80 pg/ml, about 90 pg/ml, about 100 pg/ml, about 200 pg/ml, about 300 pg/ml, about 400 pg/ml, about 500 pg/ml, about 600 pg/ml, about 700 pg/ml, about 800 pg/ml, about 900 pg/ml, about 1,000 pg/ml, about 2,000 pg/ml, about 3,000 pg/ml, about 4,000 pg/ml, about 5,000 pg/ml, about 6,000 pg/ml, about 7,000 pg/ml, about 8,000 pg/ml, about 9,000 pg/ml, about 10,000 pg/ml., about 11,000 pg/ml, about 12,000 pg/ml, about 13,000 pg/ml, about 14,000 pg/ml, about 15,000 pg/ml, about 16,000 pg/ml, about 17,000 pg/ml, about 18,000 pg/ml, about 19,000 pg/ml, about 20,000 pg/ml, about 10 pg/ml to about 10,000 pg/ml, about 20 pg/ml to about 10,000 pg/ml, about 30 pg/ml to about 10,000 pg/ml, about 40 pg/ml to about 10,000 pg/ml, about 50 pg/ml to about 10,000 pg/ml, about 60 pg/ml to about 10,000 pg/ml, about 70 pg/ml to about 10,000 pg/ml, about 80 pg/ml to about 10,000 pg/ml, about 90 pg/ml to about 10,000 pg/ml, about 100 pg/ml to about 10,000 pg/ml, about 200 pg/ml to about 9,000 pg/ml, about 300 pg/ml to about 8,000 pg/ml, about 400 pg/ml to about 7,000 pg/ml, about 500 pg/ml to about 6,000 pg/ml, about 600 pg/ml to about 5,000 pg/ml, about 700 pg/ml to about 4,000 pg/ml, about 800 pg/ml to about 3,000 pg/ml, about 900 pg/ml to about 2,000 pg/ml, about 950 pg/ml to about 1,500 pg/ml, about 10 pg/ml to about 1,000 pg/ml, about 20 pg/ml to about 1,000 pg/ml, about 30 pg/ml to about 1,000 pg/ml,
about 40 pg/ml to about 1,000 pg/ml, about 50 pg/ml to about 1,000 pg/ml, about 60 pg/ml to about 1,000 pg/ml, about 70 pg/ml to about 1,000 pg/ml, about 80 pg/ml to about 1,000 pg/ml, about 90 pg/ml to about 1,000 pg/ml, about 10 pg/ml to about 900 pg/ml, about 10 pg/ml to about 800 pg/ml, about 10 pg/ml to about 700 pg/ml, about 10 pg/ml to about 600 pg/ml, about 10 pg/ml to about 500 pg/ml, about 10 pg/ml to about 400 pg/ml, about 10 pg/ml to about 300 pg/ml, about 10 pg/ml to about 200 pg/ml, about 10 pg/ml to about 100 pg/ml, about 10 pg/ml to about 50 pg/ml, about 1,000 pg/ml to about 20,000 pg/ml, about 2,000 pg/ml to about 19,000 pg/ml, about 3,000 pg/ml to about 18,000 pg/ml, about 4,000 pg/ml to about 17,000 pg/ml, about 5,000 pg/ml to about 16,000 pg/ml, about 6,000 pg/ml to about 15,000 pg/ml, about 7,000 pg/ml to about 14,000 pg/ml, about 8,000 pg/ml to about 13,000 pg/ml, about 9,000 pg/ml to about 12,000 pg/ml, about 9,000 pg/ml to about 11,000 pg/ml, or about 9,500 pg/ml to about 10,500 pg/ml. In some embodiments, about 10 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, about 100 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, about 1,000 pg/ml of a small molecule or small molecule combination may be added to the biological sample. In some embodiments, one or more small molecule or small molecule combination may be added to the biological sample at diverse concentrations. In some embodiments, the diverse concentrations may be from about 1 pg/ml to about Ig/ml.
[0087] In some embodiments, a small molecule may have a molecular weight of less than about 5 kDa, about 4.9 kDa, about 4.8 kDa, about 4.7 kDa, about 4.6 kDa, about 4.5 kDa, about
4.4 kDa, about 4.3 kDa, about 4.2 kDa, about 4.1 kDa, about 4.0 kDa, about 3.9 kDa, about 3.8 kDa, about 3.7 kDa, about 3.6 kDa, about 3.5 kDa, about 3.4 kDa, about 3.3 kDa, about 3.2 kDa, about 3.1 kDa, about 3.0 kDa, about 2.9 kDa, about 2.8 kDa, about 2.7 kDa, about 2.6 kDa, about
2.5 kDa, about 2.4 kDa, about 2.3 kDa, about 2.2 kDa, about 2.1 kDa, about 2.0, about 1.9, about 1.8, about 1.7, about 1.6, about 1.5, about 1.4, about 1.3, about 1.2, about 1.1 kDa, about 1.0 kDa, about 0.9 kDa, about 0.8 kDa, about 0.7 kDa, about 0.6 kDa, about 0.5 kDa, about 0.4 kDa, about 0.3 kDa, about 0.2 kDa, or about 0.1 kDa. In other words, a small molecule may be any molecule that weighs less than about 5 kDa.
[0088] A small molecule as used herein may be considered a “small biomolecule” if it is a molecule having a molecular weight of less than 5 kilodaltons (kDa) and is produced by a living organism and essential to one or more biological processes (or is synthetically generated to mimic the function of a naturally generated biomolecule). In some embodiments, the small molecule may be synthetically generated (i.e., artificial). Alternatively, in some embodiments, the small molecule may be naturally produced. In some embodiments, the small molecule may be
endogenous to the host organism from which the biological sample is taken. In other embodiments, the small molecule may be exogenous to the host organism from which the biological sample is taken. Additionally or alternatively, the small molecule may be native or nonnative to the host organism. For example, the small molecule may be native or non-native to the host organism from which the biological sample is taken, e.g. a lipid, protein, or nucleic acid, and exogenously added to the biological sample.
[0089] In some embodiments, the small molecule may encourage recruitment of different proteins (and other types of biomolecules) to the surface of protein-binding agents, such as nanoparticles. In some embodiments, the small molecule may be capable of altering a protein corona. For example, the small molecule may be capable of altering a protein corona that has formed on the surface of a protein-binding agent, such as a nanoparticle. In some embodiments, the small molecule may be capable of depleting high- abundance plasma proteins, such as albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and fibrinogen, in a biological sample. Thus, in some embodiments, the small molecule may be referred to as a “high-abundance protein depleting agent.” In some embodiments, the small molecule may be capable of depleting at least one, at least two, at least three, at least four, at least five, at least six, or all seven of the high-abundance plasma proteins. In a specific embodiment, the small molecule may be capable of depleting albumin.
[0090] In some embodiments, a small molecule may physically or chemically interact with one or more proteins in a biological sample. Thus, in some embodiments, the small molecule may be referred to as a “protein-interacting agent.” In some embodiments, the small molecule may bind proteins in the biological sample. In some embodiments, the small molecule may bind albumin, IgG, antitrypsin, IgA, transferrin, haptoglobin, and/or fibrinogen. In specific embodiments, the small molecule may bind one or more abundant protein types. Thus, in some embodiments, binding the abundant proteins may deplete the number of abundant proteins in a biological sample. In some embodiments, a small molecule may change the conformation of one or more proteins in the biological sample.
[0091] In some embodiments, the small molecule may be a metabolite and/or a derivative thereof. In some embodiments, a metabolite may be any natural, synthetic, or biological small molecule, such as an amino acid, alcohols, polyols, alkaloids, organic acids, sugars (e.g., glucose) as well as nucleotides (e.g., inosine-5'-monophosphate and guanosine-5'-monophosphate).
[0092] In some embodiments, the small molecule may be a lipid and/or a derivative thereof. In some embodiments, lipids may be a group of organic compounds including all natural,
synthetic, and biological fatty compounds, such as glycerolipids (e.g., triacylglycerols also known as triglycerides, TG or TAG; diacylglycerols also known as diglycerides, DG or DAG, such as 1,2- diacylglycerols and 1,3-diacylglycerols), glycerophospholipids (e.g. phosphatidylcholine, phosphatidylethanolamine, L-a-phosphatidylinositol), sterol lipids, very-low-density lipoprotein (VLDL), low density lipoprotein (LDL), and high-density lipoprotein (HDL), fatty acids, prenol lipids, and sphingolipids.
[0093] In some embodiments, the small molecule may be a nutrient and/or a derivative thereof. In some embodiments, a metabolite can also be considered a nutrient, such as amino acids, fatty acids, vitamins (e.g. vitamin B complex), minerals and choline.
[0094] In some embodiments, the small molecule may be a plant-derived molecule and/or a derivative thereof, such as a small molecule (e.g., auxin, gibberellic acid, alkaloid, phenylpropanoid), a plant-made biologic (e.g., anti-cancer biologies), or phytopharmaceutical drugs (e.g., those with properties against human health problems, such as allergy, inflammation, etc.).
[0095] In some embodiments, the small molecule may be a metabolite and/or a derivative thereof, lipid and/or a derivative thereof, nutrient and/or a derivative thereof, plant-derived molecule and/or a derivative thereof, or a combination thereof.
[0096] In some embodiments, the small molecule may be a triacylglycerol and/or a derivative thereof, diacylglycerol, such as 1,2 and 1,3 diacylglycerol, and/or a derivative thereof, glycerophospholipid and/or a derivative thereof, glucose and/or a derivative thereof, inosine 5’- monophosphate and/or a derivative thereof, vitamin B complex and/or a derivative thereof, phosphatidylcholine and/or a derivative thereof, phosphatidylethanolamine and/or a derivative thereof, phosphatidylserine and/or a derivative thereof, phosphatidic acid and/or a derivative thereof, phosphatidylinositol and/or a derivative thereof, phosphatidylglycerol and/or a derivative thereof, cardiolipin and/or a derivative thereof, L-a-phosphatidylinositol and/or a derivative thereof, or a combination thereof.
[0097] In some embodiments, adding a small molecule into a biological sample (e.g., human plasma sample) can change the protein composition in the protein corona. In some embodiments, the small molecule may be phosphatidylcholine (PtdChos) (see, e.g., ASHKARRAN, A. et al., (2024) Small molecule modulation of protein corona for deep plasma proteome profiling. Nature Communications. 15(9638), incorporated by reference in its entirety).
[0098] Combinations of small molecules are also contemplated for use herein. For example, in some embodiments, the small molecules may be a combination of a triacylglycerol and/or a
derivative thereof, a diacylglycerol and/or a derivative thereof, and a glycerophospholipid and/or a derivative thereof. Additionally or alternatively, the small molecules may be a combination of glucose and/or a derivative thereof, inosine 5 ’-monophosphate and/or a derivative thereof, and vitamin B complex and/or a derivative thereof. Additionally or alternatively, the small molecules may be a combination of a triacylglycerol and/or a derivative thereof, a diacylglycerol and/or a derivative thereof; a glycerophospholipid and/or derivative thereof, glucose and/or a derivative thereof, inosine 5 ’-monophosphate and/or a derivative thereof, and vitamin B complex and/or a derivative thereof Additionally or alternatively, the small molecules may be a combination of glucose and/or a derivative thereof, a triacylglycerol and/or a derivative thereof, a diacylglycerol and/or a derivative thereof; and phosphatidylcholine and/or a derivative thereof. Additionally or alternatively, the small molecules may be a combination of phosphatidylethanolamine and/or a derivative thereof, L-a-phosphatidylinositol and/or a derivative thereof, inosine 5’- monophosphate and/or a derivative thereof, and vitamin B complex and/or a derivative thereof.
[0099] The one or more small molecules and one or more protein-binding agent may be added to the biological sample in any order. For example, in some embodiments, the one or more small molecule may be added to the biological sample, then the one or more protein-binding agent may be added to the biological sample containing the one or more small molecule. Alternatively, in some embodiments, the one or more protein-binding agent may be added to the biological sample, then the one or more small molecule may be added to the biological sample containing the one or more protein-binding agent. Alternatively, in some embodiments, the one or more small molecule and one or more protein-binding agent may be added to the biological sample at or approximately at the same time.
Protein-binding agent
[00100] As mentioned herein, the method includes adding one or more protein-binding agents to the biological sample. The protein-binding agents have a surface capable of binding proteins (including one or more proteoforms). The protein-binding agent also allows a protein-corona to form on the surface of the protein-binding agent, forming a complex including the protein corona and the protein-binding agent. In some embodiments, adding the protein-binging agent to the biological sample may generate a colloidal suspension.
[00101] In some embodiments, the one or more protein-binding agents may be added to the biological sample at a concentration of about Ipg/ml to about Ig/ml. Thus, in some embodiments, the one or more protein-binding agents may be added to the biological sample at a concentration of about 1 pg/ml, about 2 pg/ml, about 3 pg/ml, about 4 pg/ml, about 5 pg/ml, about 6 pg/ml,
about 7 pg/ml, about 8 pg/ml, about 9 pg/ml, about 10 pg/ml, about 20 pg/ml, about 30 pg/ml, about 40 pg/ml, about 50 pg/ml, about 60 pg/ml, about 70 pg/ml, about 80 pg/ml, about 90 pg/ml, about 100 pg/ml, about 200 pg/ml, about 300 pg/ml, about 400 pg/ml, about 500 pg/ml, about 600 pg/ml, about 700 pg/ml, about 800 pg/ml, about 900 pg/ml, about 1,000 pg/ml, about 2,000 pg/ml, about 3,000 pg/ml, about 4,000 pg/ml, about 5,000 pg/ml, about 6,000 pg/ml, about 7,000 pg/ml, about 8,000 pg/ml, about 9,000 pg/ml, about 10,000 pg/ml., about 11,000 pg/ml, about 12,000 pg/ml, about 13,000 pg/ml, about 14,000 pg/ml, about 15,000 pg/ml, about 16,000 pg/ml, about 17,000 pg/ml, about 18,000 pg/ml, about 19,000 pg/ml, about 20,000 pg/ml, about 10 pg/ml to about 10,000 pg/ml, about 20 pg/ml to about 10,000 pg/ml, about 30 pg/ml to about 10,000 pg/ml, about 40 pg/ml to about 10,000 pg/ml, about 50 pg/ml to about 10,000 pg/ml, about 60 pg/ml to about 10,000 pg/ml, about 70 pg/ml to about 10,000 pg/ml, about 80 pg/ml to about 10,000 pg/ml, about 90 pg/ml to about 10,000 pg/ml, about 100 pg/ml to about 10,000 pg/ml, about 200 pg/ml to about 9,000 pg/ml, about 300 pg/ml to about 8,000 pg/ml, about 400 pg/ml to about 7,000 pg/ml, about 500 pg/ml to about 6,000 pg/ml, about 600 pg/ml to about 5,000 pg/ml, about 700 pg/ml to about 4,000 pg/ml, about 800 pg/ml to about 3,000 pg/ml, about 900 pg/ml to about 2,000 pg/ml, about 950 pg/ml to about 1,500 pg/ml, about 10 pg/ml to about 1,000 pg/ml, about 20 pg/ml to about 1,000 pg/ml, about 30 pg/ml to about 1,000 pg/ml, about 40 pg/ml to about 1,000 pg/ml, about 50 pg/ml to about 1,000 pg/ml, about 60 pg/ml to about 1,000 pg/ml, about 70 pg/ml to about 1,000 pg/ml, about 80 pg/ml to about 1,000 pg/ml, about 90 pg/ml to about 1,000 pg/ml, about 10 pg/ml to about 900 pg/ml, about 10 pg/ml to about 800 pg/ml, about 10 pg/ml to about 700 pg/ml, about 10 pg/ml to about 600 pg/ml, about 10 pg/ml to about 500 pg/ml, about 10 pg/ml to about 400 pg/ml, about 10 pg/ml to about 300 pg/ml, about 10 pg/ml to about 200 pg/ml, about 10 pg/ml to about 100 pg/ml, about 10 pg/ml to about 50 pg/ml, about 1,000 pg/ml to about 20,000 pg/ml, about 2,000 pg/ml to about 19,000 pg/ml, about 3,000 pg/ml to about 18,000 pg/ml, about 4,000 pg/ml to about 17,000 pg/ml, about 5,000 pg/ml to about 16,000 pg/ml, about 6,000 pg/ml to about 15,000 pg/ml, about 7,000 pg/ml to about 14,000 pg/ml, about 8,000 pg/ml to about 13,000 pg/ml, about 9,000 pg/ml to about 12,000 pg/ml, about 9,000 pg/ml to about 11,000 pg/ml, or about 9,500 pg/ml to about 10,500 pg/ml. In some embodiments, about 1,000 pg/ml (0.1 mg/ml) of one or more protein-binding agents may be added to the biological sample. In some embodiments, about 2,000 pg/ml (0.2 mg/ml) of one or more proteinbinding agents may be added to the biological sample.
[00102] The protein-binding agents may bind, interact with, or attract one or more proteins in a biological sample. Thus, the protein-binding agents may also be referred to as a proteininteracting agents and/or a protein- attracting agents. A protein-binding agent may be any agent or
material that provides a surface for protein binding. In other words, the protein-binding agent may have a surface capable of binding proteins and forming a protein corona on its surface.
[00103] In some embodiments, the protein-binding agent may be an inorganic agent, metalbased agent, metal oxide-based agent, polymer-based agent, lipid-based agent, carbon-based agent, core-shell agent, composite agent, mesoporous agent, or a combination thereof. Additionally or alternatively, the protein-binding agent may include one or more nanoscale or microscale material.
[00104] A nanoscale material may be any material of which a single unit is sized (in at least one dimension) at about 1 nm and about 999 nm. For example, the nanoscale material (or nanomaterial) may be about 1 nm, about 5 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 55 nm, about 60 nm, about 65 nm, about 70 nm, about 75 nm, about 80 nm, about 85 nm, about 90 nm, about 95 nm, about 100 nm, about 150 nm, about 200 nm, about 250nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, or about 999 nm. In a specific embodiment, the nanoscale material may be about an 80 nm nanoparticle.
[00105] A microscale material may be any material of which a single unit is sized (in at least one dimension) at about 1,000 nm and about 100,000 nm. For example, the microscale material (or micromaterial) may be about 1,000 nm, 2,000 nm, 3,000 nm, 4,000 nm, 5,000 nm, 6,000 nm, 7,000 nm, 8,000 nm, 9,000 nm, 10,000 nm, 20,000 nm, 30,000 nm, 40,000 nm, 50,000 nm, 60,000 nm, 70,000 nm, 80,000 nm, 90,000 nm, or 100,000 nm.
[00106] Thus, the protein-binding agent may be any material of which a single unit is sized (in at least one dimension) at about 1 nm and about 100,000 nm. For example, the protein-binding agent may be about 1 nm, about 50 nm, about 100 nm, about 500 nm, about 1,000 nm, about 5,000nm about 10,000 nm, about 50,000 nm, about 100,000 nm, about 50 nm to about 100,000 nm, about 100 nm to about 100,000 nm, about 500 nm to about 100,000 nm, about 1,000 nm to about 100,000 nm, about 5,000 nm to about 100,000 nm, about 10,000 to about 100,000 nm, about 50,000 to about 100,000 nm, about 1 nm to about 50,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 10,000 nm, about 1 nm to about 5,000 nm, about 1 nm to about 1,000 nm, about 1 nm to about 500 nm, about 1 nm to about 100 nm, about 1 nm to about 50 nm, about 50 nm to about 50,000 nm, about 100 nm to about 10,000 nm, or about 500 nm to about 5,000 nm.
[00107] A non-limiting list of nanoscale materials includes nanoparticles, nanorods, nanospheres, nanodisks, nanoclusters, nanofibers, and nanotubes. A non-limiting list of
microscale materials includes microparticles, microrods, microspheres, and microbeads. The one or more protein-binding agents may include one or more nanoscale materials, one or more microscale materials, or a combination thereof. For example, in some embodiments, the one or more protein-binding agents may be a combination of one or more nanoparticles and one or more nanodisks. In certain embodiments, the one or more protein binding agents may be one or more nanoparticles.
[00108] The protein-binding agent may be made from an organic material, a non-organic material, or a combination thereof. In some embodiments, the material may be micelles, liposomes, iron oxide, graphene, silica, protein-based materials, polystyrene, silver, and gold materials, quantum dots, palladium, platinum, titanium, and a combination thereof.
[00109] In some embodiments, more than one protein-binding agent may include at least two to at least 1,000 protein-binding agents, which may be the same or different. In some embodiments, the number of protein-binding agents added to the biological sample may be any number to reach a desired concentration of protein-binding agents in the biological sample. In some embodiments, the number of protein-binding agents may be about 1, about 5, about 10, about 20, about 30, about 40, about 50, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, or about 1,000 per biological sample.
[00110] In some embodiments, the one or more protein-binding agents used in the method may be all of the same type. For example, in some embodiments, the protein-binding agents added to the biological sample may be polystyrene nanoparticles. In other embodiments, the proteinbinding agent may be different types. For example, the protein-binding agents added to the biological sample may be a mixture of polystyrene nanoparticles and silica microbeads. Any combination of protein-binding agents having a surface to bind proteins may be used herein. The protein-binding agent selection is not of particular importance as long as it is able to form a protein corona on its surface. One skilled in the art can determine one or more suitable materials capable of forming a protein corona on its surface.
[00111] In some embodiments, the protein-binding agent may have a poly dispersity index (PDI) of about 0.01 to about 10. Thus, the protein-binding agent may have a PDI of about 0.01, about 0.02, about 0.03, about 0.04, about 0.05, about 0.06, about 0.07, about 0.08, about 0.09, about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.02 to about 0.6, about 0.03 to about 0.5, about 0.04 to about 0.4, 0.05 to about 0.3, 0.06 to about 0.2, 0.07 to about 0.1, 0.08 to about 0.09, about 0.01 to about 0.6, about 0.01 to about 0.5, about 0.01 to about 0.4, about 0.01 to about 0.3, about 0.01 to about 0.2, about 0.01 to about 0.1, about 0.01 to about 0.09,
about 0.01 to about 0.08, about 0.01 to about 0.07, about 0.01 to about 0.06, about 0.01 to about 0.05, about 0.01 to about 0.04, about 0.01 to about 0.03, about 0.01 to about 0.02, about 0.02 to about 0.7, about 0.03 to about 0.7, about 0.04 to about 0.7, about 0.05 to about 0.7, about 0.06 to about 0.7, about 0.07 to about 0.7, about 0.08 to about 0.7, about 0.09 to about 0.7, about 0.1 to about 0.7, about 0.2 to about 0.7, about 0.3 to about 0.7, about 0.4 to about 0.7, about 0.5 to about 0.7, or about 0.6 to about 0.7. A PDI of about 0.01 may be referred to as mono-dispersed.
[00112] In some embodiments, the PDI of the protein-binding agent may be about 0.01 to about 0.7. In a specific embodiment, the protein-binding agent PDI may be about 0.7. In another embodiment, the protein-binding agent PDI may be about 0.3. In another embodiment, the proteinbinding agent PDI may be about 0.2.
[00113] The one or more protein-binding agents have a surface capable of binding proteins to the biological sample and allows a protein corona to form on the surface of the one or more protein-binding agents. This can produce a complex including the protein corona and the one or more protein-binding agent. Thus, the protein-binding agent may also be referred to herein as a protein-corona forming agent or a protein-corona attracting agent.
Protein corona
[00114] The disclosed method includes adding a protein-binding agent to a biological sample to generate a protein corona. In some embodiments, the protein corona may form on a proteinbinding agent substantially spontaneously/immediately upon addition of one or more proteinbinding agents to a biological sample. In some embodiments, the protein corona may form on a protein-binding agent upon addition of one or more protein-binding agent to a biological sample, followed by incubation of said biological sample containing one or more protein-binding agent. Thus, in some embodiments, the protein-binding agent and biological sample may be incubated for a time which allows a protein corona to form on the surface of the one or more protein-binding agent. In particular embodiments, a protein corona may be formed around a nanoparticle.
[00115] Although protein coronas may include mostly proteins, in some embodiments, the protein corona may include molecules in addition to proteins, such as lipids and other biological sample components, that are able to bind to a protein-binding agent as described herein. In this case, the protein corona may be referred to as a biomolecule corona. In some embodiments, the proteins in the protein corona may be present in the same ratio as the proteins in the untreated biological sample (i.e., the biological sample before addition of small molecules and/or proteinbinding agents). In some embodiments, the proteins in the protein corona may be present in a different ratio than the proteins in the untreated biological sample. In some embodiments, the
protein corona profile may be different in different subjects and/or in different biological samples. In some embodiments, the protein corona profile may be the same in different subjects and/or in different biological samples.
[00116] In some embodiments, a protein corona may form in different patterns and have different compositions depending on the protein and/or protein-binding agent size, shape, composition, charge, and surface functional groups. In some embodiments, protein corona properties may vary in different environmental factors such as temperature, pH, shearing stress, immersed media composition, and exposing time. In some embodiments, protein corona compositions may change according to the biochemical and physiochemical surface interactions with the protein-binding agent. In some embodiments, the protein corona may include a hard corona and/or a soft corona. The hard corona may include higher- affinity proteins that may be irreversibly bonded to the protein-binding agent’s surface and/or to other proteins in the protein corona. The soft corona may include lower-affinity proteins that are reversibly bound to the protein-binding agent’s surface and/or to other proteins in the protein corona. In some embodiments, proteins in the soft corona may be exchanged or detached over time. In some embodiments, larger proteins with lower affinities may aggregate to the protein-binding agent’s surface first, and over time, smaller proteins with higher affinities may replace them (i.e., “hardening” the corona).
[00117] Additionally or alternatively, protein compositions in protein corona may change upon adding one or more small molecules to a biological sample along with one or more protein-binding agents as described herein. For example, the addition of phosphatidylcholine to a biological sample along with one or more protein-binding agents may alter protein composition in protein corona. For example, phosphatidylcholine may itself bind specific proteins based on its properties, thus changing the type and/or number of proteins that may bind to the protein-binding agents. For example, a small molecule may reduce the amount of one or more proteins that are free-floating in the biological sample. Thus, in some embodiments, the type and/or number of proteins available to bind a protein-binding agent may be altered by one or more small molecules in the biological sample. In some embodiments, phosphatidylcholine may bind high-abundance proteins, such as albumin, thus altering the protein corona composition to include less high-abundance proteins, such as albumin.
Protein corona formation
[00118] After addition of the one or more protein-binding agents to the biological sample, the sample containing the one or more protein-binding agents may be incubated to allow a protein
corona to form on the surface of the one or more protein-binding agents. In some embodiments, the sample may be incubated with the one or more protein-binding agent. In some embodiments, the biological sample may be incubated for at least 10 seconds to about 24 hours. For example, the biological sample may be incubated for at least about 10 seconds, at least about 15 seconds, at least about 20 seconds, at least about 25 seconds, at least about 30 seconds, at least about 40 seconds, at least about 50 seconds, at least about 60 seconds, at least about 90 seconds, at least about 2 minutes, at least about 3 minutes, at least about 4 minutes, at least about 5 minutes, at least about 6 minutes, at least about 7 minutes, at least about 8 minutes, at least about 9 minutes, at least about 10 minutes, at least about 15 minutes, at least about 20 minutes, at least about 25 minutes, at least about 30 minutes, at least about 45 minutes, at least about 50 minutes, at least about 60 minutes, at least about 90 minutes, at least about 2 hours, at least about 3 hours, at least about 4 hours, at least about 5 hours, at least about 6 hours, at least about 7 hours, at least about 8 hours, at least about 9 hours, at least about 10 hours, at least about 12 hours, at least about 14 hours, at least about 15 hours, at least about 16 hours, at least about 17 hours, at least about 18 hours, at least about 19 hours, at least about 20 hours, at least about 21 hours, at least about 22 hours, at least about 23 hours, at least about 24 hours, and include any time and increment in between. In a specific embodiment, the biological sample may be incubated for about 1 hour. In some embodiments, the biological sample may be incubated more than once, for example, the biological sample may be incubated one, two, three, four, or five times.
[00119] The incubation temperature can be determined by one skilled in the art, and includes temperatures of about 4° C to about 40° C, about 4° C to about 20° C, about 10° C to about 15° C, about 10° C to about 40° C, about 4° C, about 5° C, about 6° C, about 7° C, about 8° C, about 9° C, about 10° C, about 11° C, about 12° C, about 13° C, about 14° C, about 15° C, about 16° C, about 17° C, about 18° C, about 19° C, about 20° C, about 21° C, about 22° C, about 25° C, about 30° C, about 35°, about 37° C, etc. In some embodiments, the method may be performed at room temperature (e.g., about 37° C.; e.g., about 35° C to about 40° C).
[00120] In some embodiments, the biological sample may be diluted. In some embodiments, the biological sample may be diluted using a suitable buffer, such as phosphate buffer saline (PBS), Tris-HCl buffer, ammonium bicarbonate buffer, HEPES (4-(2-hydroxyethyl)-l- piperazineethanesulfonic acid) buffer, MOPS (3-(N-morpholino)propanesulfonic acid) buffer, PIPES (1,4-Piperazinediethanesulfonic acid) buffer, or EPPS (4-(2-Hydroxyethyl)-l- piperazinepropanesulfonic acid) buffer. In a specific embodiment, the biological sample may be diluted using PBS. For example, the biological sample may be diluted to a final concentration of about 25% to about 85%, about 30% to about 80%, about 35% to about 75%, about 40% to about
10%, about 45% to about 65%, about 50% to about 60%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, or about 85%. In a specific embodiment, the biological sample may be diluted to a final concentration of about 55% using PBS.
Separating the complex from the biological sample
[00121] In some embodiments, separating the complex including the protein corona and the one or more protein binding agents from the biological sample may include one or more rounds of washing and centrifuging to remove the supernatant including the biological sample and any proteins not bound or that have low affinity for the one or more protein -binding agents. In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 washing and/or centrifugation steps may be performed. In certain embodiments, six washes may be performed. Proteins that are removed from the surface of protein-binding agents after centrifugation with at least 5000 g may be considered “low affinity proteins,” i.e., proteins which have low affinity for binding to the surface of proteinbinding agents.
[00122] In some embodiments, the complex may be separated from the remainder of the biological sample, for example, by centrifugation (such as gradient centrifugation), size exclusion chromatography, magnetic separation, field-flow fractionation, etc. In some embodiments, the complex may be washed and resuspended. In some embodiments, the proteins from the complex may be reduced, alkylated, and/or digested.
[00123] In some embodiments, proteins in the hard protein corona, soft protein corona, or a combination of the hard and soft protein corona may be detected. In some embodiments, proteins present in the protein corona may be detected after a biological sample has been incubated with one or more protein-binding agent, and optionally one or more small molecule. In some embodiments, the protein corona composition may change over time. For example, in some embodiments, a protein corona composition after 10 minutes of incubation may be different than a protein corona composition after 20 minutes of incubation. In some embodiments, molecules other than proteins may be detected that have been bound to the protein-binding agent.
Separating the one or more protein-binding agents from the protein corona
[00124] In some embodiments, separating the one or more protein-binding agents from the protein corona in the complex sample may include eluting the protein corona from the one or more protein-binding agents. Thus, in some embodiments, eluting the protein corona from the one or more protein-binding agents may include adding an elution buffer containing a detergent and/or an ionic liquid to the complex sample. An ionic liquid with detergent-like capabilities may be
used, for example an ionic liquid with 4-12 carbons in its hydrocarbon chain. A suitable detergent and/or ionic liquid can easily be selected by one skilled in the art, and includes, but is not limited to, sodium dodecyl sulphate (SDS), sodium deoxycholate (SDC), NP-40, 1 -butyl- 3 -methyl imidazolium tetrafluoroborate (BMIM BF4), l-dodecyl-3-methylimidazolium chloride (C12Im- Cl), or a combination thereof. In some embodiments, detergent concentration (i.e., sodium dodecyl sulfate, SDS), elution temperature, time, or a combination thereof may influence protein recovery.
[00125] The amount of detergent and/or ionic liquid in the elution buffer can be determined by one skilled in the art, and includes percentages of about 0.1%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, or about 5%. In some embodiments, the elution buffer may contain about 0.1% to about 5%, about 0.5% to about 4.5%, about 1% to about 4%, about 1.5% to about 3.5%, about 2% to about 3%, about 0.1% to about 4.5%, about 0.1% to about 4%, about 0.1% to about 3.5% about 0.1% to about 3%, about 0.1% to about 2.5%, about 0.1% to about 2%, about 0.1% to about 1.5%, about 0.1% to about 1%, about 0.1% to about 0.5%, about 0.5% to about 5%, about 1% to about 5%, about 1.5% to about 5%, about 2% to about 5%, about 2.5% to about 5%, about 3% to about 5%, about 3.5% to about 5%, about 4% to about 5%, or about 4.5% to about 5% detergent and/or ionic liquid. In certain embodiments, the elution buffer may contain about 0.1% to about 5% detergent and/or ionic liquid.
[00126] In some embodiments eluting the protein corona from the one or more protein-binding agents may include adding the elution buffer including detergent and/or an ionic liquid to the complex sample and incubating the complex sample. The incubation time can be determined by one skilled in the art, and includes times of about 0.5 hours, about 1 hour, about 1.5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, or about 5 hours. In some embodiments, the complex sample may be incubated for about 0.5 to about 5 hours, about 1 hour to about 4.5 hours, about 1.5 hours to about 4 hours, about 2 hours to about 3.5 hours, about 2.5 hours to about 3 hours, about 0.5 hours to about 4.5 hours, about 0.5 hours to about 4 hours, about 0.5 hours to about 3.5 hours, about 0.5 hours to about 3 hours, about 0.5 hours to about 2.5 hours, about 0.5 hours to about 2 hours, about 0.5 hours to about 1.5 hours, about 0.5 hours to about 1 hour, about 1 hour to about 5 hours, about 1.5 hours to about 5 hours, about 2 hours to about 5 hours, about 2.5 hours to about 5 hours, about 3 hours to about 5 hours, about 3.5 hours to about 5 hours, about 4 hours to about 5 hours, or about 4.5 hours to about 5 hours. In certain embodiments, the complex sample may be incubated with the elution buffer for about 0.5 to about 5 hours.
[00127] The incubation temperature can be determined by one skilled in the art, and includes temperatures of about 20°C, about 25°C, about 30°C, about 35°C, about 40°C, about 45°C, about 50°C, about 55°C, about 60°C, about 65°C, about 70°C, about 75°C, about 80°C, about 85°C, about 90°C, 95°C, or about 100°C. In some embodiments, the complex sample may be incubated at about 20°C to about 100°C, about 25°C to about 95°C, about 30°C to about 90°C, about 35°C to about 85°C, about 40°C to about 80°C, about 45°C to about 75°C, about 50°C to about 70°C, about 55°C to about 65°C, about 20°C to about 95°C, about 20°C to about 90°C, about 20°C to about 85°C, about 20°C to about 80°C, about 20°C to about 75°C, about 20°C to about 70°C, about 20°C to about 65°C, about 20°C to about 60°C, about 20°C to about 55°C, about 20°C to about 50°C, about 20°C to about 45°C, about 20°C to about 40°C, about 20°C to about 35°C, about 20°C to about 30°C, about 20°C to about 25°C, about 25°C to about 100°C, about 30°C to about 100°C, about 35°C to about 100°C, about 40°C to about 100°C, about 45°C to about 100°C, about 50°C to about 100°C, about 55°C to about 100°C, about 60°C to about 100°C, about 65°C to about 100°C, about 70°C to about 100°C, about 75°C to about 100°C, about 80°C to about 100°C, about 85°C to about 100°C, about 90°C to about 100°C, or about 95°C to about 100°C. In certain embodiments, the complex sample may be incubated at about 20°C to about 100°C.
[00128] Thus, in some embodiments, eluting the protein corona from the one or more proteinbinding agents may include adding the elution buffer, containing about 0.1% to about 5% detergent and/or ionic liquid, to the complex sample and incubating the complex sample for about 0.5 to about 5 hours at about 20°C to about 100°C.
Separating the one or more protein-binding agents from the protein corona
[00129] In some embodiments, separating the one or more protein-binding agents from the protein corona may include dissolving the one or more protein-binding agents. In some embodiments, different protein-binding agents may be dissolved by different solutions. In some embodiments, the dissolving solution may be selected based off the type of protein-binding agent used. For example, in some embodiments, one or more gold protein-binding agents may be dissolved by a potassium-iodide/iodine etching solution. In some embodiments, one or more polystyrene protein-binding agents may be dissolved by ethyl acetate. In some embodiments, one or more iron oxide nanoparticles may be dissolved by ethylenediaminetetraacetic acid (EDTA). In some embodiments, the dissolving solutions may digest the protein-binding agents while keeping the protein corona shell intact. Also provided herein are compositions including the complex (including the protein corona and the one more protein-binding agents) and the dissolving solution. For example, provided herein is a composition including the complex
(including the protein corona and the one more protein-binding agents) and a potassium- iodide/iodine etching solution. Additionally, for example, provided herein is a composition including the complex (including the protein corona and the one more protein-binding agents) and ethyl acetate. Also provided herein is a composition including the complex (including the protein corona and the one more protein-binding agents) and EDTA.
[00130] Also provided herein is a composition including one or more proteins obtained from one or more complex (including one or more proteins and a protein-binding agent) where the protein-binding agent has been removed from the composition. In other words, in some embodiments, the composition may contain substantially no protein-binding agent. In some embodiments, the protein-binding agent may be removed from the composition by adding a detergent and/or an ionic liquid as described herein. In some embodiments the protein-binding agent may be dissolved in the composition. In some embodiments, the dissolved protein-binding agent, such as nanoparticles, may stay in the composition during CE. Alternatively, in some embodiments, the dissolved protein-binding agent, such as nanoparticles, may be removed from the composition before CE.
Buffer exchange
[00131] In some embodiments, the protein corona sample may be buffer-exchanged to a mass spectrometry (MS)-compatible buffer prior to characterizing the proteins in the protein corona sample.
[00132] In some embodiments, the protein samples may contain a high-concentration of detergent (i.e., SDS) and/or ionic liquid in the elution buffer, which can be incompatible with mass spectrometry. In some embodiments, the detergent and/or ionic liquid is removed from the composition. In some embodiments, removing the detergent and/or ionic liquid maintains high protein recovery during the cleanup. In some embodiments, the buffer exchange approach is used for sample cleanup with a molecular weight cutoff membrane.
[00133] In certain embodiments, the buffer-exchange may include about a 3-kDa, 5-kDa, 10- kDa, 15-kDa, 20-kDa, 25-kDa, or 30-kDa molecular weight cutoff filter. In some embodiments, the buffer-exchange may include about a 3-kDa to about a 30-kDa, about a 5-kDa to about a 25- kDa, about a 10-kDa to about a 20-kDa, about a 3-kDa to about a 25-kDa, about a 3-kDa to about a 20-kDa, about a 3-kDa to about a 15-kDa, about a 3-kDa to about a 10-kDa, about a 3-kDa to about a 5-kDa, about a 5-kDa to about a 30-kDa, about a 10-kDa to about a 30-kDa, about a 15- kDa to about a 30-kDa, about a 20-kDa to about a 30-kDa, or about a 25-kDa to about a 30-kDa molecular weight cutoff filter. In certain embodiments, the buffer-exchange may include about a
3-kDa to about a 30-kDa molecular weight cutoff filter. In certain embodiments, the bufferexchange may include about a 10-kDa molecular weight cutoff. In some embodiments, the molecular weight cutoff membrane may be a filter, such as a Amicon® Ultra Centrifugal Filter.
[00134] In some embodiments, rounds of washing and/or centrifuging may be included in the buffer exchange sample clean up. In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 washing/centrifugation steps may be performed. In certain embodiments, six washes may be performed. In some embodiments, urea may be used as the washing buffer. In some embodiments, about 6, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, 7.0, 7.1, 7.2, 7.3, 7.4, 7.5, 7.6, 7.7, 7.8, 7.9, 8.0, 8.1, 8.2, 8.3, 8.4, 8.5, 8.6, 8.7, 8.8, 8.9, 9.0, 9.1, 9.2, 9.3 9.4, 9.5, 9.6, 9.7, 9.8, 9.9, or 10M washing buffer may be used. In certain embodiments, about 8M washing buffer may be used. In certain embodiments, about 8M urea may be used as the washing buffer.
Capillary Electrophoresis
[00135] In some embodiments, the CE may be Capillary Zone Electrophoresis (CZE) or Capillary Isoelectric Focusing (cIEF). In some embodiments, CE may be used to detect and/or characterize proteins from protein coronas. In some embodiments, CE may be used to determine protein isoforms and/or study post-translational protein modifications. CZE and cIEF are described in SUN, L. et al. (2014) Capillary zone electrophoresis for analysis of complex proteomes using an electrokinetically pumped, sheath flow nanospray interface. Proteomics. 14(0): 622-628 and XU, T. and SUN, L. (2021) A Mini Review on Capillary Isoelectric Focusing- Mass Spectrometry for Top-Down Proteomics. Front Chem. 9:651757, which are both incorporated by reference in their entirety.
[00136] In some embodiments, cIEF separates molecules based on their isoelectric point (pl), the pH at which a molecule carries no net charge. In some embodiments, a pH gradient may be established in the capillary, and molecules may migrate to the point where their net charge is zero.
[00137] In some embodiments, CZE separates molecules based on the size and/or charge. In some embodiments, molecules may migrate through a capillary filled with an electrolyte solution under the influence of an electric field, with smaller and more highly charged molecules moving faster.
[00138] In some embodiments, about 100 ng, about 150 ng, about 200 ng, about 250 ng, about 300 ng, about 350 ng, about 400 ng, about 450 ng, about 500 ng, about 550 ng, about 600 ng, about 650 ng, about 700 ng, about 750 ng, about 800 ng, about 850 ng, about 900 ng, about 950 ng, or about 1,000 ng of proteins and in the protein corona sample may be loaded into a capillary for CE. In some embodiments, about 100 ng to about 1,000 ng, about 200 ng to about 900 ng,
about 300 ng to about 800 ng, about 400 ng to about 700 ng, about 500 ng to about 600 ng, about 100 ng to about 900 ng, about 100 ng to about 800 ng, about 100 ng to about 700 ng, about 100 ng to about 600 ng, about 100 ng to about 500 ng, about 100 ng to about 400 ng, about 100 ng to about 300 ng, about 100 ng to about 200 ng, about 200 ng to about 1,000 ng, about 300 ng to about 1,000 ng, about 400 ng to about 1,000 ng, about 500 ng to about 1,000 ng, about 600 ng to about 1,000 ng, about 700 ng to about 1,000 ng, about 800 ng to about 1,000 ng, or about 900 ng to about 1,000 ng of proteins in the protein corona sample may be loaded into a capillary for CE. In certain embodiments, about 100 to about 1,000 ng of proteins in the protein corona sample may be loaded into a capillary for CE.
Increased Protein Detection
[00139] In some embodiments, the method described herein may detect more proteins (including one or more proteoforms) compared to using a biological sample without the one or more protein-binding agent. In other words, the present method may result in an increase in the number of detected proteins compared to using a biological sample without the one or more protein-binding agent. For example, in some embodiments, when using the disclosed method, there may be at least a 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 0.1 to 20, 0.2 to 20, 0.3 to 20, 0.4 to 20, 0.5 to 20, 0.6 to 20, 0.7 to 20, 0.8 to 20, 0.9 to 20, 1 to 20, 2 to 19, 3 to 18, 4 to 17, 5 to 16, 6 to 15, 7 to 14, 8 to 13, 9 to 12, 10 to 11, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 20, 3 to 20, 4 to 20, 5 to 20, 6 to 20, 7 to 20, 8 to 20, 9 to 20, 10 to 20, 11 to 20, 12 to 20, 13 to 20, 14 to 20, 15 to 20, 16 to 20, 17 to 20, 18 to 20, or 19 to 20-fold increase in the number of proteins compared to a biological sample without the one or more protein-binding agent and, optionally the one or more small molecule or small molecule combination. In some embodiments, when using the disclosed method, there may be at least a 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 0.1 to 20, 0.2 to 20, 0.3 to 20, 0.4 to 20, 0.5 to 20, 0.6 to 20, 0.7 to 20, 0.8 to 20, 0.9 to 20, 1 to 20, 2 to 19, 3 to 18, 4 to 17, 5 to 16, 6 to 15, 7 to 14, 8 to 13, 9 to 12, 10 to 11, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 20, 3 to 20, 4 to 20, 5 to 20, 6 to 20, 7 to 20, 8 to 20, 9 to 20, 10 to 20, 11 to 20, 12 to 20, 13 to 20, 14 to 20, 15 to 20, 16 to 20, 17 to 20, 18 to 20, or 19 to 20- fold increase in the number of proteins compared to a biological sample including the one or more protein-binding agent. In a specific embodiment, there may be at least a 2-fold increase in the number of proteins compared to a biological sample without the one or more protein-binding agent.
[00140] In some embodiments, the disclosed method may increase the detection of low- abundance proteins. For example, when one or more small molecules are added to the biological sample, they may bind one or more high- abundance proteins, such as albumin. Thus, in some embodiments, less high-abundance proteins may be present in the biological sample to bind the protein-binding agents. So in some embodiments, more low-abundance proteins may attach to the protein-binding molecule and may be detected by the disclosed method. Thus, in some embodiments, when detecting proteins in the protein corona, adding one or more small molecules to the biological sample with one or more protein-binding agents may reduce the detection of one or more high-abundance proteins by at least about 1%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 99%.
[00141] Thus, the present disclosure can provide a method of increasing the number of proteins detected in a biological sample compared to other known methods.
[00142] In some embodiments, protein coronas may be used to detect one or more protein types and/or the amount of one or more protein types present in a biological sample. In some embodiments, to detect more than one protein and, multiple protein-binding agent types may be used. In some embodiments, different protein-binding agents may attract different protein types. In some embodiments, using more than one type of protein-binding agent may increase the number of detected proteins in a biological sample. In some embodiments, when more than one protein-binding agent may be used, each protein-binding agent type may have distinct physicochemical properties. Thus, in some embodiments, the protein corona formed around the protein-binding agents may be different for different protein-binding agents.
[00143] Along with detecting proteins in a biological sample, the present disclosure also provides a method for detecting biomarkers in a biological sample.
Method for detecting biomarkers
[00144] In another embodiment, the present disclosure provides a method for detecting one or more biomarkers or a pattern of one or more biomarkers associated with a disease, including adding one or more protein-binding agents, having a surface capable of binding proteins, to at least two biological samples, where each biological sample from different subjects diagnosed with the disease. The method includes allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents. The method also includes separating the
complex from the biological sample to generate a complex sample, separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample, and detecting the one or more biomarkers or the pattern of one or more biomarkers in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
[00145] In some embodiments, the one or more biomarkers or pattern of one or more biomarkers may be associated with a health spectrum condition, such as a disease or disorder. In some embodiments, the method may include adding one or more small molecules, as described herein, to at least two biological samples, where each biological sample may be from different subjects diagnosed with the disease. The method may also include adding one or more proteinbinding agent (which may all be the same type or a combination of different types), as described herein, having a surface capable of binding proteins to the biological samples.
[00146] In some embodiments, a biomarker pattern may include multiple types of biomarkers. In some embodiments, a biomarker pattern may include a certain amount of one or more biomarkers in a biological sample. In some embodiments, individual organisms or biological samples may have different biomarker patterns. For example, a biomarker pattern, similar to a biomarker, may be used to detect and/or diagnose a disease or disorder in a subject.
[00147] A biomarker may be a measurable indicator of some biological state or condition. In some embodiments, bio markers may be upregulated or downregulated according to different disease types or disease stages. In some embodiments, the biomarker may be a molecular, physiologic, histologic, and/or radiographic biomarker. Additionally or alternatively, the biomarker may be predictive, prognostic, or diagnostic. In some embodiments, predictive biomarkers may help optimize ideal treatments. Examples of predictive biomarkers may include HER2/neu in breast cancer or EGFR1 mutations in non-small cell lung cancer. In some embodiments, diagnostic biomarkers can be a traceable substance that is introduced into an organism as a means to examine aspects of health. In other embodiments, a diagnostic biomarker may be used as a substance whose detection indicates a particular disease state. For example, the presence of an antibody may indicate an infection. For example, a diagnostic biomarker may be prostate-specific antigen (PSA), which may be used as a proxy of prostate size with rapid changes potentially indicating cancer.
[00148] In some embodiments, to detect more than one biomarker and/or more than one biomarker pattern, multiple protein-binding agent types may be used. In some embodiments, different protein-binding agents may attract different biomarker types. In some embodiments,
using more than one type of protein-binding agent may increase the number of detected biomarkers in a biological sample.
[00149] In some embodiments, the health spectrum condition may be any health state, including complete well-being, minor health issues, chronic conditions, and severe illnesses. The health spectrum refers to overall well-being and the factors that influence it, whether they lead to optimal health or contribute to illness. In some embodiments, a health spectrum may include a condition a subject may have which is not considered a disease or disorder yet for medical treatment. For example, in some embodiments, a health spectrum condition may include a predisposition for a disease or disorder. In other words, the presence of biomarker(s) may indicate a subject is at risk of developing a health spectrum condition and/or a disease/disorder. In some embodiments, the biomarker(s) for a disease/disorder may be the same for a disease or disorder predisposition. Additionally or alternatively, the biomarker(s) for a disease/disorder may be different for the same disease or disorder. Alternatively, the health spectrum condition may be a disease or a disorder. Some health spectrum condition examples include, but are not limited to, a predisposition to, or risk of developing, or diagnosed obesity, heart disease, liver disease, kidney disease, depression, cancer, etc.
[00150] In some embodiments, the disease may be any condition that adversely affects the structure or function of all or part of an organism and is not immediately due to any external injury. In some embodiments, the disease the biomarker may be associated with may be a neoplastic disease, a cardiovascular disease, a metabolic disease, an infectious disease, an inflammatory disease, a congenital disease, a hereditary disease, a degenerative disease, a neurological disease, or a combination thereof. For example, in some embodiments, the disease may be a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof. In some embodiments, the disease may be a neurological disease selected from Alzheimer’s disease, brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke, frontotemporal dementia, demyelinating disease, multiple sclerosis, Devic's disease, central pontine myelinolysis, progressive multifocal leukoencephalopathy, leukodystrophies, Guillain-Barre syndrome, progressing inflammatory neuropathy, Charcot-Marie-Tooth disease, chronic inflammatory demyelinating polyneuropathy, anti-MAG peripheral neuropathy, and a combination thereof.
[00151] Suitable biomarkers for Alzheimer’s disease include, but are not limited to, for example, the amyloid beta (AP)42/40 ratio, phosphorylated tau (p-tau), serum neurofilament light chain (NfL), and glial fibrillary acidic protein (GFAP).
[00152] Suitable cancer biomarkers include, but are not limited to, for example, AHSG (a2- HS-Glycoprotein), AKR7A2 (Aflatoxin Bl aldehyde reductase), AKT3 (PKB y), ASGR1 (ASGPR1), BDNF, BMP1 (BMP-1), BMPER, C9, CA6 (Carbonic anhydrase VI), CAPG (CapG), Carcino-embryonic antigen, CDH1 (Cadherin-1), CHRDL1 (Chordin-Like 1), CKB-CKM-(CK- MB), CLIC1 (chloride intracellular channel 1), CM Al (Chymase), CNTN1 (Contactin- 1), COL18A1 (Endostatin), CRP, CTSL2 (Cathepsin V), DDC (dopa decarboxylase), EGFR (ERBB1), FGA-FGB-FGG (D-dimer), FN1 (Fibronectin FN1.4), GHR (Growth hormone receptor), GPI (glucose phosphate isomerase), HMGB1 (HMG-1), HNRNPAB (hnRNP A/B), HP (Haptoglobin, Mixed Type), HSP90AA1 (HSP 90a), HSPA1A (HSP 70), IGFBP2 (IGFBP-2), IGFBP4 (IGFBP-4), IL12B-IL23A (IL-23), ITIH4 (Inter- a-tryp sin inhibitor heavy chain H4), KIT (SCF sR), KLK3-SERPINA3 (PSA-ACT), L1CAM (NCAM-L1), LRIG3, MMP12 (MMP-12), MMP7 (MMP-7), NME2 (NDP kinase B), PA2G4 (ErbB3 binding protein Ebpl), PLA2G7 (LpPLA2/PAFAH), PLAUR (suPAR), PRKACA (PRKA C-a), PRKCB (PKC-P-II), PROK1 (EG-VEGF), PRSS2 (Trypsin-2), PTN (Pleiotrophin), SERPINA1 (al-Antitrypsin), STC1 (Stanniocalcin-1), STX1A (Syntaxin 1A), TACSTD2 (GA733-1 protein), TFF3 (Trefoil factor 3), TGFBI (PIGH3), TPI1 (Triosephosphate isomerase), TPT1 (Fortilin), YWHAG (14-3-3 protein y), YWHAH (14-3-3 protein eta), prostate cancer biomarkers, for example, p63 protein, PSA, ProPSA, Pro2PSA, PHI, PCA3, TMPRSS3:ERG, PCMT, MTEN, breast cancer markers, for example, epidermal growth factor receptor 2 (HER2) oncogene, melanoma biomarker BRAF, lung cancer biomarker EML4-ALK, A2ML1, BAX, C10orf47, Clorfl62, CSDA, EIFC3, ETFB, GABARAPL2, GUK1, GZMH, HIST1H3B, HLA-A, HSP90AA1, NRGN, PRDX5, PTMA, RABAC1, RABAGAP1L, RPL22, SAP 18, SEPW1, SOX1, EGFR, EGFRvIII, apolipoprotein Al, apolipoprotein CIII, myoglobin, tenascin C, MSH6, claudin-3, claudin-4, caveolin-1, coagulation factor III, CD9, CD36, CD37, CD53, CD63, CD81, CD136, CD147, Hsp70, Hsp90, Rabl3, Desmocollin-1, EMP-2, CK7, CK20, GCDF15, CD82, Rab-5b, Annexin V, MFG-E8, HLA-DR, a miR200 microRNA, MDC, NME-2, KGF, PIGF, Flt-3L, HGF, MCP1, SAT-1, MIP-l-b, GCLM, OPG, TNF RII, VEGF-D, ITAC, MMP-10, GPI, PPP2R4, AKR1B1, Amyl A, MIP-lb, P-Cadherin, EPO and the like. For example, biomarkers for breast cancer include, but are not limited to, Circulating Tumor Cells (EpCAM, CD45, cytokeratins 8, 18+, 19+), ER/PR, HER- 2/neu, CA15-3, CA27.29, and the like. Biomarkers for colorectal cancer include, but are not limited to, for example, EGFR, KRAS, UGT1A1, Fibrin/ fibrinogen degradation product (DR-
70), Human hemoglobin (fecal occult blood), and the like. Biomarkers associated with leukemia/lymphoma include, but are not limited to, e.g., CD20 antigen, CD30, FIP1L1- PDGFRalpha, PDGFR, Philadelphia Chromosome (BCR/ABL), PML/RAR alpha, TPMT, UGT1A1, and the like. Biomarkers associated with lung cancer include but are not limited to, e.g., ALK, EGFR, KRAS and the like. Biomarkers associated with ovarian cancer include but are not limited to, e.g., ROMA (HE4+CA-125), OVA1 (multiple proteins), HE4, CA-125, and the like. Biomarkers associated with hepatocellular cancer include but are not limited to AFP-L3%, and the like. Biomarkers associated with gastrointestinal stromal tumors include but are not limited to c-Kit, and the like. Biomarkers associated with pancreatic cancer include but are not limited to CAI 9-9, and the like. Biomarkers are known in the art, and can be found in, for example, Bigbee W, Herberman R B. Tumor markers and immunodiagnosis. In: Bast R C Jr., Kufe D W, Pollock R E, et al., editors. Cancer Medicine. 6th ed. Hamilton, Ontario, Canada: BC Decker Inc., 2003; Andriole G, Crawford E, Grubb R. et al. Mortality results from a randomized prostate-cancer screening trial. New England Journal of Medicine 2009; 360( 13): 1310- 1319; Schroder F H, Hugosson J, Roobol M J, et al. Screening and prostate-cancer mortality in a randomized European study. New England Journal of Medicine 2009; 360(13):1320-1328; Buys S S, Partridge E, Black A. et al. Effect of screening on ovarian cancer mortality: the Prostate, Lung, Colorectal and Ovarian (PECO) Cancer Screening Randomized Controlled Trial. JAMA 2011; 305(22):2295-2303; Cramer D W, Bast R C Jr, Berg C D, et al. Ovarian cancer biomarker performance in prostate, lung, colorectal, and ovarian cancer screening trial specimens. Cancer Prevention Research 2011; 4(3):365-374; Sparano J A, Gray R J, Makower D F, et al. Prospective validation of a 21 -gene expression assay in breast cancer. New England Journal of Medicine 2015; First published online Sep. 28, 2015. doi: 10.1056/NEJMoal510764, and clinicalproteomicsjoumal.biomedcentr al.com/artic les/10.1186/1559-0275- 10- 13/tables/l incorporated by reference in their entireties.
[00153] Biomarkers associated with a cardiovascular disease may include, but are not limited to, lipid profile, glucose, and hormone level and physiological biomarkers based on measurement of levels of important biomolecules such as serum ferritin, triglyceride to HDLp (high density lipoproteins) ratio, lipophorin-cholesterol ratio, lipid-lipophorin ratio, LDL cholesterol level, HDLp and apolipoprotein levels, lipophorins and LTPs ratio, sphingolipids, Omega-3 Index, and ST2 level, among others. Suitable biomarkers for cardiovascular disease can be found in the art, for example, but not limited to, in van Holten et al. “Circulating Biomarkers for Predicting Cardiovascular Disease Risk; a Systemic Review and Comprehensive Overview of MetaAnalyses” PLoS One, 2013 8(4): e62080, incorporated by reference in its entirety.
[00154] Biomarkers associated with a neurological disease may include, but are not limited to, e.g., Api-42, t-tau andp-tau 181, a-synuclein, among others. See, e.g., Chintamaneni and Bhaskar “Biomarkers in Alzheimer's Disease: A Review” ISRN Pharmacol. 2012. 2012: 984786. Published online 2012 Jun. 28, incorporated by reference in its entirety.
Method for diagnosing a disease or identifying a health spectrum condition
[00155] In another embodiment, a method of diagnosing a disease or identifying another health spectrum condition, such as a predisposition for a disease in a subject is disclosed herein. The method includes adding one or more protein-binding agents, having a surface capable of binding proteins, to a biological sample from the subject and allowing a protein corona, including one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex including the protein corona and the one more protein-binding agents. The method also includes separating the complex from the biological sample to generate a complex sample, separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample, and detecting one or more biomarkers or a pattern of one or more biomarkers associated with the disease in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
[00156] As described above, the “health spectrum” covers a range of health states, including complete well-being, minor health issues, chronic conditions, and severe illnesses. For example, some health spectrum condition examples include both mental and physical conditions at a predisease or pre-disorder level all the way to severe illness, such as, but are not limited to, obesity, heart disease, liver disease, kidney disease, depression, cancer, etc.
[00157] As described above, the disease may be a neoplastic disease, a cardiovascular disease, a metabolic disease, an infectious disease, an inflammatory disease, a congenital disease, a hereditary disease, a degenerative disease, a neurological disease, or a combination thereof. For example, in some embodiments, the disease may be a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof. In some embodiments, the disease may be a neurological disease selected from Alzheimer’s disease, brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke, frontotemporal dementia, demyelinating disease,
multiple sclerosis, Devic's disease, central pontine myelinolysis, progressive multifocal leukoencephalopathy, leukodystrophies, Guillain-Barre syndrome, progressing inflammatory neuropathy, Charcot-Marie-Tooth disease, chronic inflammatory demyelinating polyneuropathy, anti-MAG peripheral neuropathy, and a combination thereof.
[00158] In some embodiments, the disclosed method may be used to provide a predisposition (e.g. a chance or likelihood) of developing the disease. In some embodiments, the disclosed method may be used to detect early onset of a disease (i.e., early disease diagnosis). For example, in some embodiments, apolipoproteins may be detected, which are indicators of cardiovascular and neurodegenerative disorders. Thus, the present method may detect protein categories that are important in disease onset and progression.
[00159] As mentioned herein, biomarkers(s) for a disease/disorder may be the same and/or different for a predisposition of a disease/disorder. Thus, in some embodiments, the same biomarkers that are detected for diagnosing a disease/disorder may be detected to identify a predisposition for that disease/disorder. In some embodiments, the biomarker concentration or amount may be lower in the sample for identifying a predisposition for a disease compared to diagnosing the disease. In some embodiments, diagnosing a disease, or identifying another health spectrum condition, such as a predisposition for a disease or health spectrum condition may help prevent or reduce the risk of developing conditions such as obesity, heart disease, liver disease, kidney disease, depression, cancer, etc. Thus, in some embodiments, subjects diagnosed with a disease, a health spectrum condition, or identified as having a predisposition for a disease or health spectrum condition may take measures to prevent or slow the progression of the disease or condition.
EXAMPLES
[00160] The following examples are merely illustrative, and do not limit this disclosure in any way.
EXAMPLE 1: CZE-MS characterization of proteins in protein corona
Experimental Methods
[00161] Chemicals and Materials: Ammonium bicarbonate (ABC), 3 -(trimethoxy silyl) propyl methacrylate, dithiothreitol (DTT), iodoacetamide (IAA), and Amicon® Ultra (0.5 mL, 10 kDa cutoff size) centrifugal filter units were ordered from Sigma- Aldrich® (St. Louis, MO). LC/MS grade water, acetonitrile (ACN), HPLC-grade acetic acid (AA), and fused silica capillaries (50 mm i.d., 360 mm o.d., Polymicro Technologies) were purchased from Fisher Scientific™
(Pittsburgh, PA). Acrylamide was obtained from Acros Organics™ (Fair Lawn, NJ). Healthy human plasma protein was purchased from Innovative Research™ (www.innov-research.com) and diluted to 55% using phosphate buffer solution (PBS, lx). Plain polystyrene nanoparticles (PSNPs, ~100 nm) were obtained from Polysciences® (www.polysciences.com).
[00162] Formation of Protein Corona on the Surface of NPs: The protein corona was formed on the PSNPs according to the procedure in ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities. Nat. Commun. 13(6610) with minor modifications. PSNPs (25 mg/mL, 75 pL) were incubated with 55% human plasma (1 mL) for 1 h at 37 °C with constant stirring to form protein coronas. Then, the mixture was centrifugated at 14,000g for 20 min to remove unbound and loosely attached plasma proteins to the surface of NPs. The protein-NP complexes underwent two cold PBS washes under identical conditions. Subsequently, two-thirds of the resulting protein-NP complexes were collected for the TDP experiment, while the remaining protein-NP complexes were used for the BUP experiment.
[00163] Sample Preparation for TDP: The protein-NP complexes were treated with four different conditions to elute the protein corona from the surface of the NPs. For the first condition, the protein corona-coated PSNPs were incubated in a 0.4% (w/v) SDS solution for 1.5 h at 60°C with constant agitation. Subsequently, the supernatant, containing the protein corona in a 0.4% SDS solution, was separated from the PSNPs through centrifugation at 19,000g for 20 min at 4°C. The resulting supernatant underwent an additional centrifugation step under the same condition to ensure complete removal of the PSNPs. The final protein corona sample was cleaned through a buffer exchange step. An Amicon® Ultra Centrifugal Filter with a molecular weight cutoff (MWCO) of 10 kDa was employed for the buffer exchange, effectively eliminating SDS from the protein samples. For the other three conditions, 1% SDS with a 1.5 h incubation, 1% SDS with a 3 h incubation, and 2% SDS with a 3 h incubation were employed using the same procedure as the 0.4% SDS and 1.5 h incubation with one minor change. For the 1 and 2% SDS conditions, the SDS and PSNP solution was centrifuged at 14,000g for 20 min at room temperature to remove the PSNPs followed by a second centrifugation at 19,000g at 4°C for 20 min. The resulting supernatant was then transferred to a different tube and underwent an additional centrifugation step at 19,000g for 20 min at 4°C to ensure complete removal of the PSNPs.
[00164] The buffer exchange protocol started with the initial wetting of the filter using 20 pL of 100 mM ammonium bicarbonate (pH 8.0) followed by centrifugation at 14,000g for 10 min. Subsequently, 200 pg of proteins was added to the filter, and centrifugation was carried out for 20
min at 14,000g. A total of 200 pL of 8 M urea in 100 mM ammonium bicarbonate solution was added followed by centrifugation at 14,000g for 20 min. This step was repeated twice under the same conditions to ensure the complete removal of SDS and other small interferences. To eliminate urea from the purified protein, the filter underwent three additional rounds of buffer exchange. Precisely, 100 mM ammonium bicarbonate was added to the filter, bringing the final volume to 200 pL. All steps were executed with centrifugation at 4°C, ensuring the thorough removal of urea from the protein corona.
[00165] After buffer exchange, the concentration of total proteins was determined using a bicinchoninic acid (BCA) kit (Fisher Scientific™) in accordance with the manufacturer’s instructions, and the sample was stored at 4°C overnight. The final protein solution, comprising 60 pL of 100 mM ammonium bicarbonate with a protein concentration of 2.4 mg/mL, was collected for CZE-MS/MS analysis.
[00166] Sample Preparation for BUP: One-third of the protein corona-coated PSNPs prepared in the “Formation of Protein Corona on the Surface of NPs” Section (~70 pg of total proteins) was dispersed in 35 pL of 100 mM ABC buffer (pH 8.0) containing 8 M urea and led the protein corona to be denatured at 37 °C for 30 min. Then, the protein corona was reduced by adding 5 pL of 70 mM DTT at 37°C for 30 min and alkylated by adding 12.5 pL of 70 mM IAA for 20 min in the dark at room temperature. The reaction was quenched by adding 1 pL of 70 mM DTT. The protein samples were diluted four times using 100 mM ammonia bicarbonate followed by trypsin (1.5 pg, bovine pancreas TPCK-treated) digestion at 37°C overnight. The digestion was finally terminated by adding formic acid (0.6% (v/v) final concentration). The samples were desalted with Sep-Pak® Cl 8 Cartridge (Waters, Milford, MA) according to manufacturer’s protocol. The eluates were lyophilized in a vacuum concentrator and then redissolved in 70 pL of 100 mM ABC buffer (pH 8.0).
[00167] CZE-MS/MS: Linear polyacrylamide (LPA)-coated fused silica capillaries (50 pm i.d., 360 pm o.d.) were prepared according to CHEN, D. et al. (2017) Capillary zone electrophoresis-mass spectrometry with microliter-scale loading capacity, 140 min separation window and high peak capacity for bottom-up proteomics. Analyst. 142:2118-2127 and ZHU, G. et al. (2016) Thermally-initiatedfree radical polymerization for reproducible production of stable linear polyacrylamide coated capillaries, and their application to proteomic analysis using capillary zone electrophoresis— mass spectrometry. Taianta. 146:839-843. After making the LPA coating, one end of the separation capillary was etched by hydrofluoric acid to reduce its outer diameter to around 100 pm.
[00168] The CZE-MS/MS system configuration involved the integration of a CESI 8000 Plus CE system (Beckman Coulter) with an Orbitrap Exploris™ 480 mass spectrometer (Thermo Fisher Scientific™), employing an in-house-built electrokinetically pumped sheath-flow CE-MS nanospray interface. The interface featured a glass spray emitter with an orifice size of 30-35 pm, filled with sheath buffer composed of 0.2% (v/v) formic acid and 10% (v/v) methanol. The spray voltage was about 2 kV. The length of the LPA-coated CZE capillary was 100 cm. The capillary’s inlet was securely affixed within the cartridge of the CE system, while its outlet was inserted into the emitter of the interface. The capillary outlet-emitter orifice distance was maintained at approximately 0.5 mm.
[00169] For TDP, a 5 psi pressure was applied to load ~240 ng of corona proteins (2.4 mg/mL, injection volume of 100 nL) into the capillary and then the inlet of the capillary was inserted into the background electrolyte (BGE, 5% (v/v) acetic acid) for CZE separation with a separation voltage of 30 kV. For BUP, 115 nL of each corona peptide sample was loaded for CZE-MS/MS. Following this, the capillary’s inlet was immersed into the BGE (5% (v/v) acetic acid), initiating the CZE separation process under a separation voltage of 30 kV.
[00170] For the mass spectrometer, all experiments were conducted using an Orbitrap Exploris™ 480 mass spectrometer (Thermo Fisher Scientific™) in data-dependent acquisition (DDA) mode. For BUP, full MS scans were acquired in the Orbitrap mass analyzer over the m/ z 300-1500 range with a resolution of 60,000 (at 200 m/z). Only precursor ions with an intensity exceeding 5E4 and a charge state of 2 to 7 were fragmented in the higher-energy collisional dissociation (HCD) cell and analyzed by the Orbitrap mass analyzer with a resolution of 60,000 (at 200 m/z). One microscan was used. The normalized collision energy was set at 28%. For MS and MS/MS spectra acquisition, the maximum ion injection times were set as 50 and 100 ms, respectively. The precursor isolation width was 2 m/z. The dynamic exclusion was applied with a duration of 15 s, and the exclusion of isotopes was enabled.
[00171] Two conditions (high-high and low-high) were implemented for the full MS parameters of TDP to enhance the precursor ion abundance. In the “high-high” condition, the full MS parameters included a high mass resolution of 480,000 (at m/z 200) with a single microscan, covering a scan range of 600-2000 m/z. Conversely, the “low-high” condition employed the least full MS resolution (7,500 at m/z 200) with 10 microscans. Other MS parameters remained consistent between the two conditions, encompassing a normalized AGC target value of 300% and an auto maximum injection time. For both conditions, precursor ions in full MS spectra were isolated with a 2 m/z window and subjected to fragmentation through higher-energy collisional
dissociation (HCD) with a normalized collision energy (NCE) of 25%. Only precursor ions with an intensity exceeding 1 x 104 and a charge state ranging from 5 to 60 nm were selected for fragmentation. Product ions were detected with a resolution of 120,000 (at 200 m/z), utilizing three microscans and maintaining a normalized AGC target value of 100% for both conditions. Dynamic exclusion was enabled with a duration of 30 s and a mass tolerance of 10 ppm (parts per million). Additionally, the “Exclude isotopes” function was activated.
[00172] For CZE-FAIMS (high field asymmetric waveform ion mobility spectrometry)- MS/MS, the FAIMS Pro™ Duo interface (Thermo Fisher Scientific™) was used. The FAIMS interface was set to standard resolution with a nitrogen carrier gas flow rate of 4.6 E/min. The distance between the ESI spray emitter orifice and the FAIMS inlet was maintained at 2-3 mm. Five different compensation voltages (CVs), -60, -40, -20, 0, and +20 V, were applied to cover small and large proteoforms according to XU, T. et al. (2023) Coupling High-Field Asymmetric Waveform Ion Mobility Spectrometry with Capillary Zone Electrophoresis-Tandem Mass Spectrometry for Top-Down Proteomics. Anal. Chem. 95:9497-9504. One CZE-MS/MS run was performed for each CV in “high-high” mode.
[00173] RPLC-MS/MS for TDP of Protein Corona: The RPEC separation was performed using an EASY-nEC 1200 system from Fisher Scientific™. A 1 pF aliquot of the protein corona sample (0.3 mg/mE) was loaded onto a home-packed C4 capillary column (75 pm i.d. x 360 pm o.d., 20 cm in length, 3 pm particles, 300 A, Bio-C4, Sepax) and separated at a flow rate of 400 nE/min. A gradient composed of mobile phase A (2% ACN in water containing 0.1% FA) and mobile phase B (80% ACN with 0.1% FA) was used for separation. The gradient profile consisted of an 80 min program: 0-60 min, 20-100% B; 60-80 min, 100% B. The EC system required an additional 30 min for column equilibration and sample loading between runs, resulting in approximately 110 min per EC-MS run. The sample was run in triplicate.
[00174] NP Characterization: Dynamic light scattering (DES) and zeta potential analyses were performed to measure the size distribution and surface charge of the NPs before and after protein corona formation using a Zetasizer® Nano Series DES instrument (Malvern Panalytical™). A Helium Neon laser with a wavelength of 632 nm was used for the size distribution measurement at room temperature. Protein corona profiles at the surface of the NPs were studied by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) analysis. Transmission electron microscopy (TEM) was carried out using a JEM-2200FS (JEOL® Ltd.) operated at 200 kV. The instrument was equipped with an in-column energy filter and an Oxford Instruments™ X-ray energy-dispersive spectroscopy (EDS) system. A total of 20 pL of
the bare PSNPs was deposited onto a copper grid and used for imaging. For protein corona-coated NPs, 20 pL of the sample was negatively stained using 20 pL of uranyl acetate (1%) and finally washed with deionized (DI) water, deposited onto a copper grid, and used for imaging on the same day.
[00175] Data Analysis: For BUP, database searching of the raw files was performed in Proteome Discoverer 2.2 with the SEQUEST HT search engine against the UniProt proteome database of human (UP000005640, 82697 entries, version 12/29/2023). Database searching of the reversed database was also performed to evaluate the false discovery rate (FDR). The database searching parameters included full tryptic digestion and allowed up to two missed cleavages, a precursor mass tolerance of 50 ppm, and a fragment mass tolerance of 0.05 Da. Carbamidomethylation (C) was set as a fixed modification. Oxidation (M), deamidated (NQ), and acetyl (protein N-term) were set as variable modifications. The data was filtered with a 1% peptide-level FDR. The protein grouping was enabled.
[00176] For TDP, proteoform identification and quantification were performed using the TopPIC (top-down mass spectrometry-based proteoform identification and characterization) pipeline. In the first step, RAW files were converted to mzML files using the Msconvert tool. The spectral deconvolution that converted precursor and fragment isotope clusters into the monoisotopic masses and proteoform features was then performed using TopFD (top-down mass spectrometry feature detection, version 1.6.3). The resulting mass spectra and proteoform feature information were stored in msalign and text files, respectively. The database search was performed using TopPIC (version 1.6.3) against a home-built protein database (~ 1,000 protein sequences), in which the proteins identified in the BUP data in this work and in literature were included. The maximum number of unexpected mass shifts was one. The mass error tolerance for precursors and fragments was 50 ppm. There was a maximum mass shift of 500 Da for unknown mass shifts. To estimate FDRs of proteoform identifications, the target-decoy approach was used, and proteoform identifications were filtered by a 1 and 5% FDR at the PrSM level and proteoform level, respectively. The lists of identified proteoforms (TDP) and protein groups (BUP) from all runs are shown in the Supporting Information of SADEGHI, S.A. et al., (2024) Mass Spectrometry- Based Top-Down Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona. ACS Nano. 18:26024-26036, which is hereby incorporated by reference in its entirety. The TopDiff (top-down mass spectrometry-based identification of differentially expressed proteoforms, version 1.6.3) software was used to perform label-free quantification of identified proteoforms using default settings.
[00177] For TDP, the complex sample data was analyzed using Xcalibur™ software (Thermo Fisher Scientific™) to get the intensity and migration time of proteoforms. For the final figures, the electropherograms were exported from Xcalibur™ and formatted using Adobe Illustrator®.
CZE-MS-based TDP workflow for the characterization of protein corona
[00178] To develop an efficient top-down proteomics (TDP) workflow for protein corona, polystyrene nanoparticles (PSNPs) and a commercially available healthy human plasma sample were employed to form protein corona and utilized dynamic pH junction-based capillary zone electrophoresis (CZE)-MS/MS for online proteoform separation, detection, and identification. FIG. 1 illustrates a detailed example TDP workflow. PSNPs were incubated with a human plasma sample to form protein corona. The protein corona was then eluted from the PSNP surface using an elution buffer containing sodium dodecyl sulfate (SDS). The eluted protein corona sample was buffer-exchanged to an MS-compatible buffer (100 mM ammonium bicarbonate, pH 8), followed by dynamic pH junction-based CZE-MS/MS analysis using an Orbitrap Exploris™ 480 mass spectrometer (Thermo Scientific™).
[00179] Both “high-high” and “low-high” MS modes were used to measure proteoforms in the protein corona. In the “high-high” mode, proteoform parent ions (MS) were detected using a 480,000 resolution (at m/z 200) and fragment ions (MS/MS) were measured using a 120,000 resolution (at m/z 200). The “high-high” mode was mainly used to identify proteoforms smaller than 30 kDa. In the “low-high” mode, proteoform parent ions (MS) were detected using a 7,500 resolution (at m/z 200) and fragment ions (MS/MS) were still measured using a 120,000 resolution (at m/z 200). Since the 480,000 resolution (at m/z 200) has difficulties in achieving isotopically resolved peaks for proteoforms larger than 30 kDa, the low-resolution MS was employed to measure the large proteoforms for determining their average masses. Finally, the TopPIC (Top- down mass spectrometry-based Proteoform Identification and Characterization) software (developed by KOU, Q. et al. (2016) TopPIC: a software tool for top-down mass spectrometrybased proteoform identification and characterization. Bioinformatics, 32:3495-3497) was used for database search to identify and quantify proteoforms from the “high-high” mode. For the “low- high” mode data, the UniDec software (MARTY, M. T. et al. (2015) Bayesian Deconvolution of Mass and Ion Mobility Spectra: From Binary Interactions to Polydisperse Ensembles. Anal Chem, 87:4370-4376) was employed for mass deconvolution to determine the average masses of large proteoforms and employed the ProSight Lite bioinformatics tool (FELLERS, R. T. et al. (2015) ProSight Lite: Graphical software to analyze top-down mass spectrometry data. Proteomics 15:1235-1238) to perform the identification of large proteoforms based on the masses of
proteoforms and corresponding fragments in a targeted fashion.
Efficient and. reproducible TDP measurements of protein corona
[00180] Protein corona was formed on the surface of PSNPs after incubating the PSNPs and human plasma sample. The particle size increased from 78.9 ± 0 nm to 105.3 ± 3.8 nm after the formation of protein corona based on the dynamic light scattering (DLS) measurement. The particle size distribution was reasonably narrow according to the standard deviation (SD) and polydispersity index (PDI) data. Interestingly, the particle-size SD and PDI of PSNPs were higher after protein binding compared to the bare NPs, suggesting some extent of heterogeneity of protein corona on individual PSNPs. The protein corona was then eluted from the surface of PSNPs using a 0.4% SDS buffer for 1.5 hours at 60°C initially. After protein corona elution, the size of particles decreased to 92.9 ± 0 nm and the PDI of 0.042, suggesting that the elution procedure using a 0.4% (w/v) SDS solution can efficiently elute proteins in the outer layer of protein corona. The data also indicates that some proteins bound to PSNPs strongly and closely (due to, e.g., non-specific adsorption) are difficult to recover. This may be, at least in large part, due to the high affinity of proteins to the surface of NPs through various physical and chemical forces. The zeta potential of PSNPs has a substantial change after protein binding and protein elution compared to bare NPs, agreeing with the particle size data. Transmission electron microscopy (TEM) was further employed to characterize the PSNPs before protein binding (bare NPs), after protein binding, and after protein elution (FIG. 2). After incubating with human plasma, the surface of PSNPs was clearly covered by a darker shell, indicating the formation of protein corona. After protein elution, there are much less proteins on the PSNPs. All the measurement data on PSNPs demonstrate that protein corona was successfully formed on the surface of PSNPs and the protein elution procedure is capable of eluting the outer layer of protein corona.
[00181] The eluted protein corona sample was measured by SDS-PAGE and CZE-MS/MS (FIG. 3, FIG. 4, and FIG. 5). Three samples were prepared in parallel starting from the PSNP and human plasma incubation for evaluating the overall reproducibility of the TDP technique. The SDS-PAGE data shows consistent proteoform profiles across the three samples with strong bands at around 25 kDa and between 50-75 kDa. The CZE-MS/MS data (“high-high” mode) also show consistent total ion current (TIC) electropherograms, the number of proteoform identifications (98+13), and the number of proteoform- spectrum matches (PrSMs, 650+26) across the three samples (FIG. 3).
[00182] To assess the quantitative reproducibility of CZE-MS/MS, the correlation coefficients of proteoform intensities between any two samples were examined (FIG. 4). The proteoform intensity was obtained using the TopDiff software (LUBECKYJ, R. A. et al., (2019) Large-Scale
Qualitative and Quantitative Top-Down Proteomics Using Capillary Zone Electrophoresis- Electrospray Ionization-Tandem Mass Spectrometry with Nanograms of Proteome Samples. J. Am. Soc. Mass Spectrom. 30:1435-1445) from the “high-high” mode data. The shared proteoforms among any two samples were utilized for the analysis. The intensities of the shared proteoforms from technical triplicate measurements were averaged and used to generate the plots. The strong linear correlations (Pearson’s r = 0.88-0.93) indicate the high quantitative reproducibility of the TDP technique developed in this work for measuring proteoforms in the protein corona samples. The proteoform intensity heatmap in FIG. 5 further illustrates the reproducibility of this TDP technique in measuring proteoforms (3-70 kDa) across three protein corona samples prepared in parallel. The CZE-MS/MS in “high-high” mode only enabled the identification of proteoforms around 10 kDa or smaller. The proteoforms larger than 28 kDa in FIG. 5 are from the “low-high” mode.
[00183] There were consistent base peak electropherogram profiles for the three protein corona samples in the “low-high” mode. Subsequently, extracted ion electropherograms (EIEs) of three large proteins (~28 kDa, ~51 kDa, and ~66 kDa) from the three samples were obtained (FIG. 6). The separation profiles and peak intensities of the three proteins are relatively consistent across the three samples. The peak 1 shows obvious intensity variations across the three samples, which may be due to the limited number of data points across the top part of the peak, arising from the relatively long data acquisition cycle in TDP measurement. FIG. 7 shows the mass spectra and deconvoluted masses of the three large proteins. Interestingly, multiple proteoforms for each protein were detected and the deconvoluted data shows the masses and relative abundance of different proteoforms in each protein peak. For example, three clear proteoforms of protein 1 with masses 66,560 Da, 66,872 Da (66,560+312 Da), and 67,184 Da (66,560+312+312 Da) were observed with the 66,560-Da proteoform as the most abundant one. Based on the mass and the list of proteins in the corona sample from BUP measurement, the protein was presumably identified as HSA. The MS/MS data also provided strong support about the identification of the HSA with a series of matched b-type fragment ions close to the N-terminus of the protein. The theoretical mass of HSA is 66,438 Da considering the 17 disulfide bonds (native form). The most abundant proteoform (66,560 Da) represents a +122-Da mass shift compared to the native form, presumably due to a combination of one phosphorylation (+80 Da) and one acetylation (+42 Da) or cysteinylation (+119 Da). The other two proteoforms of HSA most likely have additional glycosylation PTMs according to literature data and the information in the UniProt knowledgebase (www.uniprot.org/uniprotkb/P02768/entry).
[00184] For the 28-kDa protein in FIG. 7, five proteoforms were clearly detected and the most
abundant one had a mass of 28,120 Da. According to the proteoform mass and the BUP measurement results, the protein was presumably identified as Apolipoprotein A-I (APOA1). The MS/MS data offers strong evidence about the identification with a series of matched y-type fragment ions close to the C-terminus of the protein. Interestingly, the five different proteoforms of intact APOA1 show a continuous increase in proteoform mass from the highest abundant to the lowest abundant one with a 266-Da mass difference between any neighboring proteoforms. The 266-Da mass difference could be due to lipidation, for example, adding a stearic acid (octadecanoic acid, 284 Da) molecule through reaction between the carboxyl group of lipid and amine group of protein via removing a H2O molecule. The five different proteoforms could represent APO Al with 0-4 stearic acid modification sites. One previous TDP study provided strong evidence that APOA1 proteoforms are closely related to the indices of cardiometabolic health. APOA1 is a prognostic marker in renal and liver cancers according to the human protein atlas (www.proteinatlas.org/ENSGOOOOOH8137-APOAl). The data here suggests that coupling PSNP-based protein corona and MS-based TDP could be a valuable strategy for discovering novel APOA1 proteoform biomarkers of diseases (i.e., cancers and cardiometabolic diseases) using human plasma samples.
[00185] For the 51-kDa protein, two proteoforms were detected with 51,200 Da and 51,860 Da. Based on the mass and the BUP data of the protein corona sample, the protein might be Fibrinogen beta chain (50,763 Da in the mature form without PTMs) or Clusterin (50,062 Da in the mature form without PTMs). The MS/MS spectra of those two detected proteoforms do not match well with theoretical b- or y-type of fragment ions from Fibrinogen beta chain or Clusterin. More additional studies may be conducted in the future to provide more information about the identities of those two proteoforms.
[00186] To ensure that the detected proteoforms primarily originate from the protein corona and not from the processes involved in protein dissociation from the surface of PSNPs, the procedure was run using a standard protein mixture. In this case, a standard protein mixture (ubiquitin (Ub), myoglobin (Mb), and carbonic anhydrase (CA)) was analyzed by CZE-MS under two different conditions. The experimental details were described in SADEGHI, S.A. et al., (2024) Mass Spectrometry-Based Top-Down Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona. ACS Nano. 18:26024-26036, which is hereby incorporated by reference in its entirety. The charge state distributions may have slight changes between the two conditions due to the significant conformational differences of proteoforms. The masses of proteoforms between the two conditions are consistent with only ±1 Da difference, most likely due to potential errors from MS measurement using a Q-TOF mass spectrometer in this
experiment. According to this study, the possibility of proteoform changes (e.g., artificial modifications) in the protein corona is low during sample preparation. More systematic investigations can be performed on this topic using a complex system in the future to make additional conclusions.
Comparisons between TDP and. BUP for the characterization of protein corona
[00187] The protein corona samples were analyzed by both BUP and TDP using CZE-MS/MS. For the BUP, the protein corona on PSNPs was prepared through the on-bead digestion procedure in BLUME, J. E. et al. Rapid, deep and precise profiling of the plasma proteome with multi- nanoparticle protein corona. Nat. Commun. 2020, 11, 3662 and ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities . Nat. Commun, 13:6610. For TDP, the protein corona was eluted from PSNPs and cleaned prior to CZE-MS/MS analysis, as shown in FIG. 1. Three protein corona samples prepared in parallel starting from the same human plasma sample were analyzed.
[00188] FIGs. 8-12 summarize the protein corona data from TDP (FIG. 8) and BUP (FIG. 9). Single-shot TDP analysis of the protein corona sample by CZE-MS/MS consistently identified nearly 600 PrSMs, 100 proteoforms, and 20 proteoform families across the three protein corona samples (FIG. 8). The proteoform family represents a set of proteoforms from the same gene. In total, 263 proteoforms and 50 proteoform families were identified from the three corona samples. Single-shot BUP analysis of the protein corona sample by CZE-MS/MS identified around 8,000 peptide- spectrum matches (PSMs), 3,000 peptides, and 200 protein groups with high reproducibility (FIG. 9). A protein group can contain multiple proteins sharing the same set of identified peptides and those proteins usually have similar protein sequences. BUP analyses of the three protein corona samples using CZE-MS/MS identified, in total, 280 protein groups and 5,606 peptides.
[00189] TDP produced a much lower proteome coverage than BUP in terms of the number of identified genes (50 vs. -280) due to its lower sensitivity, resulting from much wider charge state distributions of intact proteoforms compared to peptides after electrospray ionization. To better highlight the sensitivity differences between BUP and TDP, two calibration curves were created on CZE-MS-based TDP and BUP analysis of intact bovine serum albumin (BSA) and BSA tryptic digest samples. The slope of the calibration curve from BUP is about three orders of magnitude larger than that from TDP (E9 vs. E6), indicating drastically higher sensitivity of BUP compared to TDP. However, TDP offered more advanced measurements of proteins in a proteoform- specific manner. For example, 18 different proteoforms of the gene Serum amyloid A-l (SAA I) were
identified by CZE-MS/MS-based TDP (“high-high” mode) from the protein corona samples. The SAA1 protein is a prognostic marker of renal cancer according to the Human Protein Atlas (www.proteinatlas.org/ENSG00000173432-SAAl). Four of the identified SAA1 proteoforms are shown in FIG. 10. They have varied length of sequences due to variations in signal peptide cleavage. For example, proteoform 1 has the whole protein sequence without signal peptide cleavage and proteoforms 2-4 have signal peptide cleavage at slightly different positions. Additionally, various PTMs occurred on those proteoforms. For example, proteoform 1 carries a N-terminal acetylation and one +463-Da mass shift close to the N-terminus; proteoform 2 has a roughly 14-Da mass shift, corresponding to a methylation PTM; proteoform 3 contains a 16-Da mass shift, corresponding to an oxidation modification; proteoform 4 does not have any PTMs. The location of those PTMs on the proteoforms is at the underlined regions. Those proteoforms were identified with high confidence evidenced by the E- Value and the number of matched fragment ions and were reasonably well characterized according to the five-level classification system described in SMITH, E. M. et al. (2019) A five-level classification system for proteoform identifications. Nat. Methods, 16:939-940.
[00190] The proteoform 4 is level- 1 identification, proteoforms 2 and 3 belong to level-2a identifications, and proteoform 1 is a level-3 proteoform. The TDP approach also enabled the determination of relative abundance of proteoforms from the same gene. For example, proteoform 1 had a much lower abundance than the other three proteoforms, suggested by their intensities. On the other hand, the BUP measurement identified a protein group SAA1, containing five proteins, which have similar protein sequences. The main issue of the BUP data is that the protein identification is ambiguous, which means which protein(s) in the protein group exist in the protein corona sample is unresolved due to the “peptide-to-protein” inference problem.
[00191] To advance the number of proteoform and proteoform family identifications by TDP, the protein corona elution from the surface of PSNPs was changed by adjusting the concentration of SDS (0.4%, 1%, and 2%) and incubation time (1.5 h and 3 h). It is clear that 1% SDS with 3-h incubation and 2% SDS with 3-h incubation produced better recovery of large proteoforms than 0.4% SDS and 1% SDS with 1.5-h incubation. Interestingly, after protein corona elution, the size of the PSNPs has a minimal change across the four different elution conditions. Even though the protein corona was eluted with 2% SDS and a 3-h incubation at 60°C, the size of the PSNPs was still almost the same as that after being treated with 0.4% SDS for 1.5 h (~ 93 nm), suggesting that it is challenging to recover the proteoforms bound to PSNPs strongly and closely due to, e.g., non-specific adsorption. Considering that the proteoforms on the surface of protein corona most likely have a stronger impact on the outcome of nanomedicine, additional efforts were not made
to improve the protein corona elution and 1% SDS was chosen with a 3-h incubation time as the conditions for the following studies.
[00192] Three different platforms were then employed to analyze one protein corona sample from 1% SDS with a 3-h incubation to maximize the proteoform identifications. First, the same CZE-MS/MS condition as before was used to analyze the new protein corona sample in technical triplicate. The sample was then analyzed using CZE-FAIMS (high field asymmetric waveform ion mobility spectrometry)-MS/MS with five different compensation voltages according to Xu, T. et al. (2023) Coupling High-Field Asymmetric Waveform Ion Mobility Spectrometry with Capillary Zone Electrophoresis-Tandem Mass Spectrometry for Top-Down Proteomics. Anal. Chem, 95:9497-9504. Then the sample was analyzed by RPLC-MS/MS. In all the experiments, the “high-high” mode was employed on an Orbitrap Exploris™ 480 mass spectrometer (Thermo Scientific™). As shown in FIG. 11, CZE-MS/MS analysis of the new sample identified more proteoforms and proteoform families than that from the sample eluted by 0.4% SDS (FIG. 8). CZE-FAIMS-MS/MS identified nearly 50% more proteoforms and 64% more proteoform families than CZE-MS/MS alone. CZE-MS/MS and RPLC-MS/MS identified a comparable number of proteoform and proteoform family identifications. Combining the data from all the three platforms, 883 proteoforms and 73 proteoform families were identified with on average about 10 proteoforms per family. The proteoform overlap between CZE and RPLC as well as CZE-FAIMS and RPLC is low (FIG. 12) suggesting the nice complementarity between CZE- MS/MS and RPLC-MS/MS for proteoform identifications. Interestingly, only less than 40% of the proteoforms from CZE-MS/MS were covered by CZE-FAIMS-MS/MS, indicating some potential proteoform loss in the FAIMS interface. The identified proteoforms were reasonably well characterized. About 300 proteoforms have no PTMs and no additional unexplained mass shifts and they are well characterized, belonging to the level 1 identifications. About 130 proteoforms were identified with specific PTMs (i.e., methylation, acetylation, phosphorylation, and oxidation) with or without an accurate localization, belonging to either level 1 or 2a identifications. For all other proteoforms with mass shifts (-450), their sequences and genes were identified, and the PTMs and localizations were not determined. Those proteoforms belong to level 3 identifications.
[00193] Additional data analyses were performed in terms of the identified protein biomarkers in this study. TDP identified specific proteoforms of 48 protein biomarkers (Table 1). The number of proteoforms ranged from 1 to 131 for those protein biomarkers. The BUP data of those protein biomarkers showed that each protein group can have a range of 1-16 proteins with an average of nearly 4 proteins per protein group. Interestingly, ten protein biomarkers were identified by TDP
but not by BUP, most likely due to potential sample loss during the sample preparation of BUP. The data agrees with Xu, T. et al. (2020) Automated. Capillary Isoelectric Focusing-Tandem Mass Spectrometry for Qualitative and Quantitative Top-Down Proteomics. Anal. Chem, 92:15890-15898, on comparing BUP and TDP for quantitative proteomics of zebrafish brains. The result here indicates the valuable and more advanced information that TDP can offer on protein biomarkers of human plasma.
[00194] Table 1: Summary of disease-related protein biomarkers identified by TDP and BUP. The protein biomarkers were determined according to the information in the Human Protein Atlas (www.proteinatlas.org/) except proteins labelled by *, which were determined based on literature data. N/A represents not identified.
EXAMPLE 2: cIEF-MS characterization of proteins in protein corona cIEF-MS-based TDP workflow for the characterization of protein corona
[00195] In this study, an automated cIEF-MS/MS method was developed to measure NP protein corona using TDP (FIG. 13). The protein corona was prepared on polystyrene NPs (PSNPs) according to the procedure in ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities. Nat. Commun, 13:6610 and SADEGHI, S.A. et al., (2024) Mass Spectrometry-Based Top-Down
Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona. ACS Nano. 18:26024-26036. PSNPs were used due to experience in changing the parameters involved in the formation of a pure protein corona, ensuring highly accurate and reproducible MS results. Full details on PSNP parameter changing and characterization for protein corona formation are available in ASHKARRAN, A. A. et al. (2022) Measurements of heterogeneity in proteomics analysis of the nanoparticle protein corona across core facilities. Nat. Commun, 13:6610; SADEGHI, S.A. et al., (2024) Mass Spectrometry-Based Top-Down Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona. ACS Nano. 18:26024- 26036; ASHKARRAN, A. A. et al. (2024) Standardizing Protein Corona Characterization in Nanomedicine: A Multicenter Study to Enhance Reproducibility and Data Homogeneit. Nano Lett. 24:9874-9881; and SHEIBANI, S. et al. (2021) Nanoscale characterization of the biomolecular corona by cryo-electron microscopy, cryo-electron tomography, and image simulation. Nat. Comm. 12(1):573. Detailed information on protein corona formation is in the Supporting Information of ZHU, G. et al., (2024) Deciphering nanoparticle protein coronas by capillary isoelectric focusing-mass spectrometry based top-down proteomics. Chem. Commun. 60: 11528, which is hereby incorporated by reference in its entirety. Briefly, PSNPs were incubated with healthy human plasma to form protein coronas. After washing with PBS, the protein corona was eluted from PSNPs using a 0.4% (w/v) SDS solution, followed by buffer exchange to a 100 mM NH4HCO3 buffer for cIEF-MS/MS.
[00196] First, ampholyte concentration for cIEF-MS/MS was determined. Higher ampholyte concentration can achieve better separation resolution but also can lead to unavoidable ionization suppression of proteoforms. Three concentrations of ampholytes, 1.5%, 1%, and 0.5%, were studied using a standard protein mixture containing cytochrome c (cyt c, pl 10.8), myoglobulin (Mb, pl 6.9) and carbonic anhydrase (CAs, pl 5.4). Automated cIEF-MS was carried out using the sandwich injection approach, the electrokinetically pumped sheath flow CE-MS interface, and an AGILENT® 6545XT Q-TOF mass spectrometer. The three proteins were all baseline separated under the three conditions. CIEF with a higher concentration of ampholyte could reach a better separation resolution. CIEF with a higher ampholyte concentration tends to need a longer analysis time due to the higher buffering capacity of ampholytes, requiring a longer time for titration. Considering the analysis time, separation resolution, and instrument contamination from ampholytes, cIEF-MS with 0.5% ampholytes was employed for the analysis of protein coronas.
[00197] The electropherograms of cIEF-MS runs of three protein corona samples (SI, S2, S3) were prepared in parallel and each sample was analyzed in technical duplicates. The separation profile and base peak intensity are reasonably consistent across all runs, demonstrating
reproducible protein corona analyses. The cIEF-MS observed clear proteoform peaks of large proteins (a, b, and c) and small proteins (d). For example, three and four proteoforms were detected for the 28 kDa (a) and 66 kDa (b) proteins with the relative abundance of those proteoforms resolved. The data demonstrates that cIEF-MS can delineate large and small proteoforms in protein coronas.
[00198] To identify proteoforms based on MS/MS, cIEF was coupled to an Orbitrap Exploris™ 480 mass spectrometer (Thermo Scientific™). One protein corona sample was analyzed in technical duplicates by a high-high mode, employing high mass resolution for both MSI and MS2. The duplicate cIEF-MS/MS runs generated a consistent separation profile and similar numbers of proteoform (63 +/- 1, n = 2) and protein (25 +/- 0, n = 2) identifications (FIG. 14). In total, 82 proteoforms and 31 proteins were identified. The two runs shared 43 proteoforms, representing nearly 70% of the number of identified proteoforms in one run (FIG. 15). The proteoform intensity between the duplicate runs has a clear linear correlation (Pearson’s r = 0.99) (FIG. 16).
[00199] An example of identified proteoforms of gene APOA1 is shown in FIG. 17. Proteoform 1 is 28091.238 Da and has one N-terminal acetylation and one 157.947 Da mass shift (FIG. 17). According to the dbPTM database, the S and T amino acid residues in this specific amino acid sequence (position 52-66) can be phosphorylated. The deconvoluted MS/MS spectrum of the proteoform shows clear signals of ions corresponding to losses of H2O and H3PO4. Therefore, the 157.947 Da mass shift should correspond to two phosphorylation events. Proteoform 1 belongs to the level 2A identification. Proteoform 2 is 22519.954 Da and has N- terminal truncation and a 144.354 Da mass shift between position 195 and 232. Multiple acetylation (i.e., K) and phosphorylation (i.e., S or T) could happen in this region. The 144.354 Da may be from the combination of phosphorylation, acetylation, and other PTMs. Proteoform 3 is 18431.319 Da and has N-terminal truncation and one 264.751 Da mass shift. Proteoforms 2 and 3 are level 3 identifications. The mass errors of matched fragment ions of the three APO Al proteoforms are smaller than 10 ppm, and for most fragment ions, especially proteoforms 2 and 3, the mass error is close to 0. The high mass accuracy of matched fragment ions ensures the high confidence of identifications. The results demonstrate that this cIEF-MS/MS-based TDP could measure diverse proteoforms of the same gene (i.e., APOA1) in the protein corona. This technique could provide a relative abundance of proteoforms from the same gene. For example, proteoform 1 of gene APOA1 has a substantially higher abundance than others, evidenced by its much higher intensity (2E10 vs.< 5E6).
[00200] APOA1 is a prognostic marker of cancer (www.proteinatlas.org/). 12 proteoforms of APO Al were identified. Overall, over 70 proteoforms of 16 cancer-related genes were identified. cIEF-MS/MS-based TDP provides an advanced view of the diverse proteoforms in the protein corona, including variations such as truncations and PTMs, as well as their combinations. This proteoform-centric TDP approach has the potential to offer more detailed and accurate information about protein corona composition compared to the traditional peptide-centric BUP. This enhanced accuracy is fundamental for developing and improving safer and more efficient nanomedicines. The data also implies that TDP profiling of protein corona could be useful for discovering novel proteoform biomarkers of diseases, e.g., cancers.
[00201] Most of the proteoforms identified in this study using the high-high mode (B80%) are smaller than 10 kDa (FIG. 18). The other 20% of the proteoforms are in the mass range of 11-30 kDa. It is challenging for TDP to identify large proteoforms (430 kDa) from complex samples due to their substantially lower measurement sensitivity compared to small proteoforms. To improve the measurement quality of large proteoforms, a low-high approach was employed, utilizing low- resolution MSI and high-resolution MS2. Twenty-four proteoforms were detected close to or larger than 28 kDa from 4 proteins (FIG. 19). Nine proteoforms were detected from protein 1 (466 kDa) and 2 proteoforms from protein 4 (443 kDa) (FIG. 19). Based on the capillary zone electrophoresis (CZE)-MS/MS data in EXAMPLE 1, protein 1 should be human serum albumin (HSA). CZE-MS/MS detected three HSA proteoforms and, here, cIEF-MS/MS observed nine HSA proteoforms in a mass range of 66436-67 625 Da, and the 66 820 Da proteoform is the most abundant one. The theoretical mass of HSA with 17 disulfide bonds (native form) is 66 438 Da. The smallest HSA proteoform detected here (66 436 Da) should be the native form. HSA can be modified by various PTMs, e.g., phosphorylation and glycosylation. The HSA proteoforms detected here must be due to the combinations of PTMs and/or sequence variations. cIEF-MS/MS detected two proteoforms of protein 4 (about 43 kDa), not observed in the CZE-MS/MS study in EXAMPLE 1. For protein 2, cIEF separated it into two peaks (2 and 20), and each peak has two proteoforms. The EXAMPLE 1 CZE-MS/MS study only detected the two highly abundant proteoforms of protein 2 (51 200 and 51 860 Da) in one peak. The nine proteoforms of protein 3 with masses of about 28 kDa correspond to the products of gene APO Al based on the high-high mode data (FIG. 17). The most abundant proteoform of intact APOA1 has an average mass of 28 110 Da, which should be the proteoform in FIG. 17, having a monoisotopic mass of 28 091 Da (average mass 28 108 Da). The nine proteoforms were separated into three peaks (3, 30, and 300) by cIEF. Only five intact APOA1 proteoforms were observed by CZE-MS/MS in one peak.
[00202] In summary, these findings demonstrate that cIEF-MS/MS is a superior technique for TDP characterization of protein coronas. It surpasses CZE-MS/MS in large proteoform analysis due to its exceptionally high separation resolution and greater sample loading capacity (400-1000 nL vs. 100 nL). This study marks the first investigation of cIEF-MS/MS for TDP of protein coronas. It is anticipated that cIEF-MS/MS will significantly advance the field of nanomedicine by providing efficient measurement of small and large proteoforms in protein coronas.
EXAMPLE 3: cIEF-MS characterization of proteins in protein corona
Experimental Methods
[00203] Chemicals and materials: The following materials were purchased from Sigma- Aldrich® (St. Louis, MO): ammonium bicarbonate (ABC), 3-(trimethoxysilyl) propyl methacrylate (y-MAPS), dithiothreitol (DTT), ammonia hydroxide (NH3H2O), ammonium acetate (NH4AC), ammonium persulfate (APS), Pharmalytes with pl range of 3-10, 5-8 and 8-10.5 (GE HEALTHCARE™). HPLC-grade acetic acid (AA), MS-grade water, methanol (MeOH), formic acid (FA), Amicon® Ultra (0.5 mF, 10 kDa cut-off size) centrifugal filter units, and fused silica capillaries (50 pm i.d./360 pm o.d., Polymicro Technologies) were purchased from Fisher Scientific™ (Pittsburgh, PA). Acrylamide was purchased from Acros Organics™ (Fair Lawn, NJ). A healthy human plasma sample was purchased from Innovative Research™ (www.innov- research.com) and diluted to 55% using phosphate buffer solution (PBS, IX). Polystyrene NPs (PSNPs, -100 nm) were obtained from Polysciences® (www.polysciences.com).
[00204] Sample preparation and characterization: Briefly, PSNPs were mixed with 55% human plasma. This mixture was stirred constantly for one hour at a temperature of 37 °C to allow the formation of a protein corona. After an hour, the protein-NP complexes were separated by centrifugation at 14,000 xg for 20 minutes to remove unbound proteins. The resulting pellet was then washed twice with cold PBS.
[00205] Dynamic Light Scattering (DLS) analysis was performed to measure the size distribution of the PSNPs before and after protein corona formation. The measurements were conducted at room temperature using a Zetasizer® Nano Series DLS instrument (Malvern Panalytical™) equipped with a Helium-Neon laser at a wavelength of 632 nm.
[00206] For the collected protein corona coated PSNPs, the proteins were extracted from the NP surface by incubating the pellet in a 0.4% SDS solution with agitation for 1.5 hours at 60°C, and the extracted protein corona-containing supernatant was separated by centrifugation. An
Amicon® Ultra centrifugal filter with a lOkDa molecular weight cutoff was used to exchange the buffer and remove the SDS. Finally, the protein corona sample in 100 mM ammonium bicarbonate (NH4HCO3) was measured using a BCA assay to determine the protein concentration, and it was adjusted to 1.5 mg/mL for MS analysis.
[00207] cIEF-MS/MS analysis: An automated cIEF-MS/MS system was built by combining a CESI 8000 Plus CE system (Beckman Coulter) with an Orbitrap Exploris™ 480 mass spectrometer (Thermo Fisher Scientific™) using an in-house electrokinetically pumped sheathflow CE-MS nanospray interface. The cIEF separation was carried out using an 80 cm long linear polyacrylamide (LPA)-coated capillary (50 pm i.d./360 pm o.d.). One end of the separation capillary was etched using hydrofluoric acid to reduce its outer diameter to approximately 100 pm. The interface featured a glass spray emitter with an orifice size of 30-35 pm, filled with sheath buffer composed of 0.2% (v/v) formic acid and 10% (v/v) methanol. The spray voltage was set to 2 kV, and the capillary outlet to emitter orifice distance was maintained at approximately 0.5 mm. The distance between the emitter orifice and MS inlet was about 2 mm.
[00208] The automated cIEF-MS system was based on the "sandwich" injection approach. The injection sequence involved three steps: first, a 6 cm catholyte plug was injected at 10 psi for 8 seconds containing 0.3% NH4OH, followed by a 20 cm mixture of sample and ampholyte plug containing 0.6% ampholytes (3-10, 5-8, and 8-10.5, GE HEALTHCARE™), injected at 10 psi for 27 seconds. Approximately 600 ng of corona proteins (1.5 mg/mL, injection volume of 400 nL) were loaded into the capillary, and finally, a 50 cm anolyte plug was injected at 10 psi for 67 seconds containing 5% acetic acid. This combination provided efficient focusing and mobilization of the protein corona samples under a separation voltage of 30 kV.
[00209] The Orbitrap Exploris™ 480 mass spectrometer was used to analyze the proteoforms separated by cIEF in data-dependent acquisition (DDA) mode. Two approaches were used for data acquisition to detect both small and large (>30 kDa) proteoforms. For small proteoforms (<30 kDa), a “high-resolution MSI and high-resolution MS/MS” mode, i.e., “High-High” mode was employed. The detailed parameters for the “High-High” mode include MSI resolution 480,000 at m/z 200 with a single microscan across a m/z range of 700-3000. Maximum ion injection time was set to 50 ms for MS and 100 ms for MS/MS. Normalized AGC target 300%, Ions with an intensity of over 1E4 and charge states varying from 5 to 60 were isolated with a 2 m/z window, followed by fragmentation through higher-energy collision dissociation (HCD) at 25% normalized collision energy (NCE). Dynamic exclusion was enabled with a duration of 30 seconds and a mass tolerance of 10 ppm, and isotope exclusion was activated. The fragment ions were
detected with a resolution of 120,000 at m/z 200 and normalized AGC 100%. For large proteoforms (>30 kDa), a “low-resolution MSI and high-resolution MS/MS” mode, i.e., “Low- High” mode, was employed. MSI resolution of 7,500 at m/z 200 was used. The microscan setting is 3. The other parameters are the same as the “High-High” mode.
[00210] Data analysis: Data processing was conducted using the TopPIC software to identify and quantify proteoforms in the "high-high" mode. For the "low-high" mode, the UniDec software facilitated mass deconvolution, determining the average masses of larger proteoforms. The cIEF- MS/MS data analysis began with converting RAW files to mzML format using MSconvert. The converted data was then processed using TopFD (version 1.7.0) software to convert isotope clusters into monoisotopic masses and identifiable proteoform features, with the results stored in msalign and text files. The deconvoluted mass spectra and proteoform features were then searched against a home-built protein database of approximately 1,000 sequences using TopPIC software (version 1.7.0), which included proteins previously identified in bottom-up proteomics (BUP) data. TopPIC was configured to accommodate a single unexpected mass shift per proteoform with a maximum shift of 500 Da and maintained a mass error tolerance of 50 ppm for both precursor and fragment ions. A target-decoy approach was used to estimate and control the false discovery rate (FDR), setting it at 1% at the proteoform- spectrum match (PrSM) level and 5% at the proteoform level. Finally, the identified proteoforms were quantified using TopDiff software to enable label-free quantification across technical replicates. The quantification aggregated the intensities of each proteoform's peaks across all scans and charge states. The raw mass spectrometry data files were processed using Xcalibur™ Qual Browser (Thermo Fisher Scientific™) to extract proteoform intensity values and migration time information. Base peak chromatograms and extracted ion chromatograms were generated to visualize the separation profiles. The electropherograms are graphically refined using Adobe Illustrator® for figure preparation. cIEF-MS-based TDP workflow for the characterization of protein corona
[00211] A high-throughput automated cIEF-MS/MS technique was developed that took 30 minutes or less per run for TDP of NP protein coronas (FIGs. 20-22). The protein corona sample was prepared according to Sun, L. et al. (2013) Ultrasensitive and Fast Bottom-up Analysis of Femtogram Amounts of Complex Proteome Digests. Angewandte Chemie - International Edition, 52(51): 13661-13664. Briefly, proteoforms in the protein corona of PSNPs were eluted using a 0.4% SDS buffer and cleaned up by buffer exchange, followed by cIEF-MS/MS (FIG. 20). The protein coronas were uniformly formed on the surface of PSNPs evidenced by the dark shells of
the particles (FIG. 21). The advanced cIEF-MS/MS technique for high-throughput TDP analysis of protein coronas was carried out by employing a short separation capillary for cIEF-MS with a commercial CE system (FIG. 22). An 80-cm long LPA-coated capillary was used, and the effective capillary length for cIEF separation was shorter than 30 cm because a “sandwich” injection approach was used. The “sandwich” method includes injecting a plug of catholyte (0.3% NH4OH, pH ~11), a plug of the sample with ampholyte in 100 mM NH4HCO3, and a long plug of anolyte (5% acetic acid, pH 2.4). A 6 cm plug of catholyte, a 20 cm sample plug containing 0.6% ampholytes (pl 3-10, 5-8, and 8-10.5 with ratios 1:1:1), and a 50 cm plug of anolyte using a standard protein mixture were used. Because of the short effective capillary length for cIEF (<30 cm), the analysis could be carried out in a high-throughput fashion. Also, because the total capillary length was 80 cm, a regular commercial CE system could be used, allowing the technique to be adopted easily by other researchers.
Reproducibility of high-throughput cIEF -MS/MS -based TDP for protein corona
[00212] The protein corona sample of PSNPs was analyzed using the high-throughput cIEF- MS/MS technique for 50 runs (FIG. 23). Each cIEF-MS run took less than 30 minutes, producing a 2-6-fold improvement in analysis throughput compared to the previous cIEF- MS/MS-based TDP studies. Twenty-five runs were performed in "high-high" mode and twenty- five runs in "low-high" mode to evaluate the technique for both small and large proteoform measurements.
[00213] The cIEF-MS/MS technique produced reproducible separation, detection, and identification of proteoforms. The electropherograms in FIG. 23 show consistent separation profiles of proteoforms in both “high-high” and “low-high” modes. In the "high-high" mode, a normalized level (NL) of 4.0+0.8 E10 (n=25) was obtained for the total ion current (TIC) electropherograms, corresponding to a relative standard deviation (RSD) of 20%. In the "low- high" mode, the NL was 8.2+0.8E09, corresponding to an RSD of about 10%. The numbers of proteoform and proteoform- spectrum match (PrSM) identifications are also consistent across the “high-high” runs, with 71+10 (n=25) proteoforms and 196+30 (n=25) PrSMs. In the “low-high” runs, three large proteins (28 kDa, 51 kDa, and 66 kDa) with multiple proteoforms per protein were consistently detected. Those three proteins correspond to human serum albumin (HSA, 66 kDa), Apolipoprotein A-I (APOA1, 28 kDa), and an unknown protein (51 kDa). The data agreed reasonably with the results in EXAMPLES 1 and 2. Seven proteoforms were randomly selected from seven genes and their migration times were determined across the 25 “high-high” runs from the database search results to further evaluate the separation reproducibility (Table 2). The RSDs
of migration time of those proteoforms were less than 4% across 25 cIEF-MS/MS runs, indicating excellent reproducibility of the technique for proteoform separation. To validate the consistency of the high-throughput cIEF-MS/MS methodology for protein corona analysis regarding proteoform intensity, six cIEF-MS/MS runs (“High-High”) were randomly chosen and the intensity of overlapped proteoforms was plotted between any two runs (FIG. 24). The proteoform intensity showed strong linear correlations between runs, evidenced by the high Pearson’s correlation coefficient (r) of 0.92+0.06, underscoring the quantitative reproducibility of the TDP technique for protein corona analysis. FIG. 25 shows the mass distribution of identified proteoforms from all the “high-high” runs. The mass of identified proteoforms ranged from ~2 kDa to ~30 kDa, and the majority of them were ~10 kDa or smaller. If the large proteoforms detected in “low-high” mode are included, the mass range of identified proteoforms will be extended to 2-66 kDa.
[00214] Table 2. Summary of migration time of seven selected proteoforms from seven proteins across 25 “High-High” runs.
Protein biomarkers identified by cIEF-MS/MS analysis of protein corona
[00215] The TDP analysis of protein corona identified 53 genes, and the number of detected proteoforms per gene ranged from 1 to 102 (Table 3). 33 out of the 53 genes are biomarkers, and they span various protein families and functional classes, including but not limited to apolipoproteins, complement proteins, immunoglobulins, and cytoskeletal proteins. Many of these proteins are associated with diverse diseases and pathological states, underscoring their potential utility as diagnostic or prognostic biomarkers. Particularly noteworthy is the prominence of the apolipoprotein family within the dataset, which includes APOA1, APOA2, APOA4, APOB,
AP0C2, AP0C3, APOE, and APOF. These proteins play critical roles in lipid metabolism and are strongly linked to cardiometabolic disorders such as dyslipidemia, metabolic syndrome, atherosclerosis, and cardiovascular diseases. This connection provides a significant opportunity for further research into their pathobiological mechanisms and applications in clinical diagnostics. Multiple proteoforms were identified for most of the apolipoprotein family members. For example, 102 proteoforms of the APO Al gene were identified, and four of them are shown in FIG. 26. Those proteoforms carry variations due to signal peptide cleavages, truncations, and PTMs. Proteoform 1 has an N-terminal cleavage of the first 26 amino acid residues, most likely corresponding to the signal peptide cleavage. Proteoform 1 also contains a mass shift of +288.535 Da in the highlighted region. Based on the PTM information in the dbPTM database (awi.cuhk.edu.cn/dbPTM/), three lysine residues in the highlighted region can be acetylated, corresponding to a +126 Da mass shift. The +288.535 Da may correspond to the combination of acetylation and other PTMs. For Proteoform 2, the first 70 amino acid residues were truncated, and it carries a mass shift of +340.875 Da. Proteoform 3 shows a truncation of the first 127 amino acid residues at the N-terminus. This proteoform also contains an unknown mass shift of +59.054 Da in the highlighted region. The exact nature of this modification requires further investigation. Proteoform 4 exhibits an N-terminal removal of the first 24 amino acid residues due to the signal peptide cleavage and a C-terminal truncation. This proteoform also contains an unknown mass shift of +143.988 Da in the highlighted region. Additionally, the dataset identifies biomarkers pertinent to inflammatory and autoimmune diseases, including complement proteins (C3, C9), serum amyloid A proteins (SAA1), and serpins (SERPINA1, SERPINC1). The catalog also highlights biomarkers associated with neurodegenerative conditions such as Alzheimer's disease (APOE, CEU) and Parkinson's disease (ABCB9), as well as proteins involved in cancer progression and metastasis, such as ACTB, KRT1, and RAB15. These findings demonstrate that TDP studies of protein coronas from large cohorts of plasma samples with various diseases using this high-throughput cIEF-MS/MS technique could discover novel proteoform biomarkers of diseases, facilitating the disease early diagnosis and drug development.
[00216] Table 3. Summary of the identified genes and corresponding number of proteoforms from the cIEF-MS/MS-based TDP analysis of nanoparticle protein coronas.
EXAMPLE 4: Dissolving nanoparticles for capillary electrophoresis (CE) characterization of proteins in protein corona
[00217] A complementary approach is introduced herein to enhance the efficacy of both top- down proteomics (TDP) and bottom-up proteomics (BUP) in improving the depth of proteomics using protein corona-coated nanoparticles. A significant challenge in applying TDP to analyze protein corona profiles is the difficulty in detaching intact proteins from the surface of nanoparticles without altering their structure or losing important post-translational modifications. To address this issue, a novel strategy that involves the dissolution of nanoparticles after the formation of the protein corona, thereby preserving the integrity of the entire protein corona for
subsequent analysis. This approach allows researchers to examine the full array of proteoforms and proteins within the corona without the interference or structural complications caused by the presence of nanoparticles.
[00218] The method of nanoparticle removal is tailored to the specific type of nanoparticles being used. For instance, gold nanoparticles can be dissolved using a potassium-iodide/iodine etching solution, which effectively removes the nanoparticles while leaving the protein corona intact. Similarly, for polystyrene nanoparticles, ethyl acetate can be used as a solvent to dissolve the nanoparticles without affecting the associated proteins. This strategic dissolution of nanoparticles enables a more comprehensive and accurate analysis of the protein corona, facilitating a deeper proteomics analysis of protein corona and also understanding of nanoparticle interactions within biological systems and enhancing the overall diagnostic and therapeutic potential of nanomedicines.
[00219] Specific concentrations of potassium iodide/iodine, ethyl acetate, and ethylenediaminetetraacetic acid (EDTA) for digesting the nanoparticle core while preserving the protein corona can vary based on the type and amount of nanoparticles and the nature of the protein corona. Below are example ranges of concentration:
[00220] Potassium Iodide/iodine
Typical Concentration Range:
KI: 0.1 M to 1 M
L: 0.01 M to 0.1 M
Consideration: the pH should be kept around neutral to slightly acidic to prevent denaturation of the protein layer.
[00221] Ethyl Acetate
Typical Concentration/Volume:
Volume: 1-5 mL per reaction, depending on the sample volume and required extraction efficiency.
Consideration: Performing the extraction at lower temperatures (e.g., 5C) can improve phase separation and reduce protein denaturation.
[00222] Ethylenediaminetetraacetic Acid (EDTA)
Typical Concentration Range:
EDTA: 0.5 mM to 5 mM
Consideration: a compatible buffer system with EDTA (e.g., Tris-HCl or PBS) should be used to maintain pH stability during the digestion process.
[00223] An example protocol outline is presented below:
[00224] Preparation:
(1) Dissolving EDTA in PBS to a final concentration of 1 mM.
(2) Preparing a KI/F solution with KI at 0.5 M and L at 0.05 M.
[00225] Digestion:
(1) Adding the KI/E solution to the nanoparticle (e.g., gold or iron oxide nanoparticles)-protein mixture under gentle stirring.
(2) Incubating the mixture at room temperature for 30 minutes to 1 hour, monitoring the digestion progress.
[00226] Extraction:
(1) Adding ethyl acetate (2 mF) to the digested mixture.
(2) Vortex and allow the layers to separate (typically 10-15 minutes).
(3) Carefully collecting the aqueous layer containing the intact protein corona.
[00227] Post-Digestion:
(1) Performing buffer exchange or dialysis to remove residual ethyl acetate and small molecules.
(2) Analyzing the protein corona integrity using techniques like SDS-PAGE, Western blotting, or dynamic light scattering (DLS).
[00228] Table 4: Sequence Listing.
Claims
1. A method for characterizing proteins in a biological sample, the method comprising: adding one or more protein-binding agents, having a surface capable of binding proteins, to the biological sample and allowing a protein corona, comprising one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex comprising the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample comprising the complex; separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample; and characterizing the proteins in the protein corona sample with Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
2. The method of claim 1, wherein separating the complex from the biological sample comprises one or more rounds of washing and centrifuging to remove a supernatant comprising the biological sample and any proteins not bound or that have a low affinity for the one or more protein-binding agents.
3. The method of claim 1 or claim 2, wherein separating the one or more protein-binding agents from the protein corona in the complex sample comprises eluting the protein corona from the one or more protein-binding agents.
4. The method of claim 3, wherein eluting the protein corona from the one or more proteinbinding agents comprises adding an elution buffer containing a detergent and/or an ionic liquid with 4-12 carbons in the hydrocarbon chain to the complex sample.
5. The method of claim 4, wherein the detergent and/or ionic liquid with 4-12 carbons in the hydrocarbon chain comprises sodium dodecyl sulphate (SDS), sodium deoxycholate (SDC), NP-40, 1 -butyl- 3 -methyl imidazolium tetrafluoroborate (BMIM BF4), l-dodecyl-3- methylimidazolium chloride (C12Im-Cl), or a combination thereof.
6. The method of claim 4 or claim 5, wherein eluting the protein corona from the one or more protein-binding agents comprises adding the elution buffer, containing about 0.1% to about
5% detergent and/or ionic liquid with 4-12 carbons in the hydrocarbon chain, to the complex sample and incubating the complex sample for about 0.5 to about 5 hours at about 20°C to about 100°C.
7. The method of claim 1 or claim 2, wherein separating the one or more protein-bindings agents from the protein corona comprises dissolving the one or more protein-binding agents.
8. The method of claim 7, wherein one or more gold protein-binding agents are dissolved by a potassium-iodide/iodine etching solution, one or more polystyrene protein-binding agents are dissolved by ethyl acetate, or one or more iron oxide nanoparticles are dissolved by ethylenediaminetetraacetic Acid (EDTA).
9. The method of any one of the previous claims, wherein the protein corona sample is buffer- exchanged to a mass spectrometry (MS) -compatible buffer prior to characterizing the proteins in the protein corona sample.
10. The method of claim 9, wherein the buffer-exchange comprises about a 3-kDa to about a 30- kDa molecular weight cutoff filter.
11. The method of any one of the previous claims, wherein the CE comprises Capillary Zone Electrophoresis (CZE) or Capillary Isoelectric Focusing (cIEF).
12. The method of any one of the previous claims, wherein about 100 to about 1,000 ng of proteins in the protein corona sample are loaded into a capillary.
13. The method of any one of the previous claims, wherein the one or more protein-binding agents comprise an inorganic agent, metal-based agent, metal oxide-based agent, polymer-based agent, lipid-based agent, carbon-based agent, core-shell agent, composite agent, mesoporous agent, or a combination thereof.
14. The method of any one of the previous claims, wherein the one or more protein-binding agents comprise one or more nanoscale or microscale material, such as a nanoparticle, nanorod, nanosphere, nanodisk, nanocluster, nanofiber, nanotube, microparticle, microrod, microsphere, microbead, or a combination thereof.
15. The method of any one of the previous claims, wherein the one or more protein-binding agents are of the same type.
16. The method of any one of the previous claims, wherein the biological sample comprises blood, plasma, serum, lung lavage, a cell lysate, menstrual blood, urine, a tissue, amniotic fluid, cerebrospinal fluid, tears, a liquid biopsy, saliva, or semen, preferably plasma.
17. A method for detecting one or more biomarkers or a pattern of one or more biomarkers associated with a disease, the method comprising: adding one or more protein-binding agents, having a surface capable of binding proteins, to at least two biological samples, each biological sample from different subjects diagnosed with the disease, and allowing a protein corona, comprising one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex comprising the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample; separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample; and detecting the one or more biomarkers or the pattern of one or more biomarkers in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
18. A method of diagnosing a disease in a subject, the method comprising: adding one or more protein-binding agents, having a surface capable of binding proteins, to a biological sample from the subject and allowing a protein corona, comprising one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex comprising the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample; separating the one or more protein-binding agents from the protein corona in the complex sample to generate a protein corona sample; and detecting one or more biomarkers or a pattern of one or more biomarkers associated with the disease in the protein corona by characterizing the proteins in the protein corona sample by Capillary Electrophoresis (CE) and Tandem Mass Spectrometry (MS/MS).
19. The method of claim 17 or 18, wherein the disease comprises a neoplastic disease, a cardiovascular disease, a metabolic disease, an infectious disease, an inflammatory disease, a congenital disease, a hereditary disease, a degenerative disease, a neurological disease, or a combination thereof.
20. The method of claim 19, wherein the disease comprises a neoplastic disease selected from lung cancer, pancreas cancer, myeloma, myeloid leukemia, meningioma, glioblastoma, breast cancer, esophageal squamous cell carcinoma, gastrointestinal cancer, liver cancer, prostate cancer, bladder cancer, cervical cancer, ovarian cancer, thyroid cancer, neuroendocrine cancer, and a combination thereof.
21. The method of claim 19, wherein the disease comprises a neurological disease selected from Alzheimer’s disease, brain tumors, epilepsy, Parkinson's disease, ALS, arteriovenous malformation, cerebrovascular disease, brain aneurysms, epilepsy, multiple sclerosis, Peripheral Neuropathy, Post-Herpetic Neuralgia, stroke, frontotemporal dementia, demyelinating disease, multiple sclerosis, Devic's disease, central pontine myelinolysis, progressive multifocal leukoencephalopathy, leukodystrophies, Guillain-Barre syndrome, progressing inflammatory neuropathy, Charcot-Marie-Tooth disease, chronic inflammatory demyelinating polyneuropathy, anti-MAG peripheral neuropathy, and a combination thereof.
22. A composition comprising a protein corona sample produced by: adding one or more protein-binding agents, having a surface capable of binding proteins, to a biological sample and allowing a protein corona, comprising one or more proteins, to form on the surface of the one or more protein-binding agents to produce a complex comprising the protein corona and the one more protein-binding agents; separating the complex from the biological sample to generate a complex sample comprising the complex; and separating the one or more protein-binding agents from the protein corona in the complex sample to generate the protein corona sample.
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463568036P | 2024-03-21 | 2024-03-21 | |
| US63/568,036 | 2024-03-21 | ||
| US202463686961P | 2024-08-26 | 2024-08-26 | |
| US63/686,961 | 2024-08-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025199014A1 true WO2025199014A1 (en) | 2025-09-25 |
Family
ID=97140235
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/020200 Pending WO2025199014A1 (en) | 2024-03-21 | 2025-03-17 | Mass spectrometry-based top-down methods of characterizing proteins in protein corona |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025199014A1 (en) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20180172694A1 (en) * | 2016-12-16 | 2018-06-21 | The Brigham And Women's Hospital, Inc. | System and Method for Protein Corona Sensor Array for Early Detection of Diseases |
| WO2024253966A2 (en) * | 2023-06-07 | 2024-12-12 | Board Of Trustees Of Michigan State University | Compositions and methods for detecting proteins in protein corona |
-
2025
- 2025-03-17 WO PCT/US2025/020200 patent/WO2025199014A1/en active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20180172694A1 (en) * | 2016-12-16 | 2018-06-21 | The Brigham And Women's Hospital, Inc. | System and Method for Protein Corona Sensor Array for Early Detection of Diseases |
| WO2024253966A2 (en) * | 2023-06-07 | 2024-12-12 | Board Of Trustees Of Michigan State University | Compositions and methods for detecting proteins in protein corona |
Non-Patent Citations (4)
| Title |
|---|
| EFFENDI MUKHTAR, RAHMAH FIQIH SILVIA, SUSANTI LUSI, ROCHMAN NURUL TAUFIQU: "Preparation of monodisperse polystyrene spheres by physical method", IOP CONFERENCE SERIES: MATERIALS SCIENCE AND ENGINEERING, INSTITUTE OF PHYSICS PUBLISHING LTD., GB, vol. 509, GB , pages 012088, XP093360879, ISSN: 1757-899X, DOI: 10.1088/1757-899X/509/1/012088 * |
| FASERL KLAUS, CHETWYND ANDREW J., LYNCH ISEULT, THORN JAMES A., LINDNER HERBERT H.: "Corona Isolation Method Matters: Capillary Electrophoresis Mass Spectrometry Based Comparison of Protein Corona Compositions Following On-Particle versus In-Solution or In-Gel Digestion", NANOMATERIALS, MDPI, vol. 9, no. 6, 20 June 2019 (2019-06-20), pages 898, XP093360873, ISSN: 2079-4991, DOI: 10.3390/nano9060898 * |
| SADEGHI SEYED AMIRHOSSEIN, ASHKARRAN ALI AKBAR, WANG QIANYI, ZHU GUIJIE, MAHMOUDI MORTEZA, SUN LIANGLIANG: "Mass Spectrometry-Based Top-Down Proteomics in Nanomedicine: Proteoform-Specific Measurement of Protein Corona", ACS NANO, AMERICAN CHEMICAL SOCIETY, US, 14 September 2024 (2024-09-14), US , XP093360881, ISSN: 1936-0851, DOI: 10.1021/acsnano.4c04675 * |
| TABATABAEIAN NIMAVARD REYHANE, SADEGHI SEYED AMIRHOSSEIN, MAHMOUDI MORTEZA, ZHU GUIJIE, SUN LIANGLIANG: "Top-Down Proteomic Profiling of Protein Corona by High-Throughput Capillary Isoelectric Focusing-Mass Spectrometry", JOURNAL OF THE AMERICAN SOCIETY FOR MASS SPECTROMETRY, ELSEVIER SCIENCE INC, US, vol. 36, no. 4, 2 April 2025 (2025-04-02), US , pages 778 - 786, XP093360886, ISSN: 1044-0305, DOI: 10.1021/jasms.4c00463 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Blume et al. | Rapid, deep and precise profiling of the plasma proteome with multi-nanoparticle protein corona | |
| JP7811240B2 (en) | Compositions, methods and systems for protein corona analysis and uses thereof | |
| Baimanov et al. | In situ analysis of nanoparticle soft corona and dynamic evolution | |
| US12241899B2 (en) | Systems and methods for sample preparation, data generation, and protein corona analysis | |
| Ashkarran et al. | Small molecule modulation of protein corona for deep plasma proteome profiling | |
| Wu et al. | Mag-Net: Rapid enrichment of membrane-bound particles enables high coverage quantitative analysis of the plasma proteome | |
| Arvizo et al. | Identifying new therapeutic targets via modulation of protein corona formation by engineered nanoparticles | |
| Pitek et al. | Transferrin coated nanoparticles: study of the bionano interface in human plasma | |
| Metatla et al. | Neat plasma proteomics: getting the best out of the worst | |
| Fang et al. | Magnetic metal oxide affinity chromatography-based molecularly imprinted approach for effective separation of serous and urinary phosphoprotein biomarker | |
| Ashkarran et al. | Deep plasma proteome profiling by modulating single nanoparticle protein corona with small molecules | |
| Timerbaev | How well can we characterize human serum transformations of magnetic nanoparticles? | |
| Onigbinde et al. | Optimization of glycopeptide enrichment techniques for the identification of clinical biomarkers | |
| EP4724810A2 (en) | Compositions and methods for detecting proteins in protein corona | |
| Kim et al. | Development of an online microbore hollow fiber enzyme reactor coupled with nanoflow liquid chromatography-tandem mass spectrometry for global proteomics | |
| WO2025199014A1 (en) | Mass spectrometry-based top-down methods of characterizing proteins in protein corona | |
| Wang et al. | Establishment and clinical application evaluations of a deep mining strategy of plasma proteomics based on nanomaterial protein coronas | |
| Ma et al. | Highly efficient TiO2-based one-step strategy for micro volume plasma-derived extracellular vesicles isolation and multiomics sample preparation | |
| Kuruvilla et al. | Surface proteomics on nanoparticles: a step to simplify the rapid prototyping of nanoparticles | |
| Zeng et al. | Advances in phosphoproteomics and its application to COPD | |
| Lee et al. | Development of a parallel microbore hollow fiber enzyme reactor platform for online 18O-labeling: Application to lectin-specific lung cancer N-glycoproteome | |
| EP3578653A2 (en) | Method of aptamer selection and method for purifying biomolecules with aptamers | |
| Sadeghi | Advancing Mass Spectrometry-Based Human Plasma Proteomics Using a Nanoparticle Protein Corona Strategy | |
| Shuken et al. | Next-Generation Multiplexed Targeted Proteomics Quantifies Post-Translational Modifications, Compound-Protein Interactions, and Disease Biomarkers with High Throughput | |
| Panchal et al. | NP–Protein Corona Interaction: Characterization Methods and Analysis |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25774537 Country of ref document: EP Kind code of ref document: A1 |