WO2022148455A1 - Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes - Google Patents

Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes Download PDF

Info

Publication number
WO2022148455A1
WO2022148455A1 PCT/CN2022/070948 CN2022070948W WO2022148455A1 WO 2022148455 A1 WO2022148455 A1 WO 2022148455A1 CN 2022070948 W CN2022070948 W CN 2022070948W WO 2022148455 A1 WO2022148455 A1 WO 2022148455A1
Authority
WO
WIPO (PCT)
Prior art keywords
cell
cov
sars
peptide
epitopes
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/070948
Other languages
French (fr)
Inventor
Matthew Robert MCKAY
Ahmed Abdul Quadeer
Syed Faraz Ahmed
Muhammad Saqib Sohail
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hong Kong University of Science and Technology
Original Assignee
Hong Kong University of Science and Technology
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hong Kong University of Science and Technology filed Critical Hong Kong University of Science and Technology
Priority to US18/260,767 priority Critical patent/US20250000965A1/en
Publication of WO2022148455A1 publication Critical patent/WO2022148455A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61KPREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K39/00Medicinal preparations containing antigens or antibodies
    • A61K39/12Viral antigens
    • A61K39/215Coronaviridae, e.g. avian infectious bronchitis virus
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61PSPECIFIC THERAPEUTIC ACTIVITY OF CHEMICAL COMPOUNDS OR MEDICINAL PREPARATIONS
    • A61P31/00Antiinfectives, i.e. antibiotics, antiseptics, chemotherapeutics
    • A61P31/12Antivirals
    • A61P31/14Antivirals for RNA viruses
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K14/00Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
    • C07K14/005Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from viruses
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K16/00Immunoglobulins [IG], e.g. monoclonal or polyclonal antibodies
    • C07K16/08Immunoglobulins [IG], e.g. monoclonal or polyclonal antibodies against material from viruses
    • C07K16/10RNA viruses
    • C07K16/102Coronaviridae (F)
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K16/00Immunoglobulins [IG], e.g. monoclonal or polyclonal antibodies
    • C07K16/08Immunoglobulins [IG], e.g. monoclonal or polyclonal antibodies against material from viruses
    • C07K16/10RNA viruses
    • C07K16/102Coronaviridae (F)
    • C07K16/104Severe acute respiratory syndrome coronavirus 2 [SARS‐CoV‐2]
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/53Immunoassay; Biospecific binding assay; Materials therefor
    • G01N33/569Immunoassay; Biospecific binding assay; Materials therefor for microorganisms, e.g. protozoa, bacteria, viruses
    • G01N33/56966Animal cells
    • G01N33/56972White blood cells
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61KPREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K38/00Medicinal preparations containing peptides
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K2317/00Immunoglobulins specific features
    • C07K2317/30Immunoglobulins specific features characterized by aspects of specificity or valency
    • C07K2317/34Identification of a linear epitope shorter than 20 amino acid residues or of a conformational epitope defined by amino acid residues
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2770/00MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssRNA viruses positive-sense
    • C12N2770/00011Details
    • C12N2770/20011Coronaviridae
    • C12N2770/20022New viral proteins or individual genes, new structural or functional aspects of known viral proteins or genes
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2770/00MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssRNA viruses positive-sense
    • C12N2770/00011Details
    • C12N2770/20011Coronaviridae
    • C12N2770/20034Use of virus or viral component as vaccine, e.g. live-attenuated or inactivated virus, VLP, viral protein
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N2333/00Assays involving biological materials from specific organisms or of a specific nature
    • G01N2333/005Assays involving biological materials from specific organisms or of a specific nature from viruses
    • G01N2333/08RNA viruses
    • G01N2333/165Coronaviridae, e.g. avian infectious bronchitis virus
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N2333/00Assays involving biological materials from specific organisms or of a specific nature
    • G01N2333/435Assays involving biological materials from specific organisms or of a specific nature from animals; from humans
    • G01N2333/705Assays involving receptors, cell surface antigens or cell surface determinants
    • G01N2333/70503Immunoglobulin superfamily, e.g. VCAMs, PECAM, LFA-3
    • G01N2333/70539MHC-molecules, e.g. HLA-molecules

Definitions

  • COVID-19 is caused by a novel coronavirus, Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) .
  • SARS-CoV-2 Severe Acute Respiratory Syndrome Coronavirus 2
  • SARS-CoV-2 Severe Acute Respiratory Syndrome Coronavirus 2
  • Monoclonal antibodies are laboratory-made molecules that act as substitute antibodies. They can help the immune system recognize and respond more effectively to the virus, making it more difficult for the virus to reproduce and cause harm. Like other infectious organisms, SARS-CoV-2 can mutate over time, resulting in genetic variation in the population of circulating viral strains. The risk-benefit assessment for using certain monoclonal antibodies may not be favorable due to the increased frequency of resistant variants. Global efforts to combat COVID-19 have led to the rapid development of multiple vaccines. As the virus continues to circulate worldwide, virus variants have emerged in several regions, raising concerns about their potential to escape vaccine-induced antibody responses.
  • the present disclosure provides new compositions and methods useful for eliciting an immune response to SARS-CoV-2 and/or other coronavirus infections, based on the discovery that SARS-CoV-2 and/or other coronavirus infections may elicit T-cell and/or B cell response in a subject.
  • the present disclosure is related to a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence.
  • the peptide comprises at least one T cell epitope set forth in Table 3 and at least one T cell epitope set forth in Tables 15-47. In some embodiments, the peptide comprises at least one T cell epitope set forth in Table 3 and at least one B cell epitope set forth in Table 4. In some embodiments, the peptide comprises at least one T cell epitope set forth in Tables 15-47 and at least one B cell epitope set forth in Table 4. In some embodiments, the peptide comprises at least one T cell epitope set forth in Table 3 and at least one T cell epitope set forth in Tables 15-47 and at least one B cell epitope set forth in Table 4.
  • the peptide may optionally further comprise at least one heterologous amino acid sequence.
  • the peptide comprises or consists of at least one of the T cell epitopes.
  • the peptide comprises or consists of at least one of the B cell epitopes.
  • the peptide comprises or consists of at least one of the T cell epitopes and at least one heterologous amino acid sequence.
  • the peptide comprises or consists of at least one of the B cell epitopes and at least one heterologous amino acid sequence.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids. In any cases, the peptide may further comprise at least one heterologous amino acid sequence.
  • the second aspect of the present disclosure provides a nucleic acid comprising a polynucleotide sequence encoding the peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence.
  • the polynucleotides encodes a peptide of at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the polynucleotides encodes a peptide of between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the polynucleotides encodes a peptide of about 20 amino acids. In some embodiments, the polynucleotides encodes a peptide of about 100 amino acids.
  • the polynucleotide sequence comprises deoxyribonucleic acid (DNA) .
  • the polyneucleotide sequence comprises ribonucleic acid (RNA) such as messenger RNA (mRNA) .
  • RNA ribonucleic acid
  • the present disclosure provides an expression cassette comprises any of the peptide as described herein and the peptide is operably linked to a promoter.
  • the present disclosure provides a vector comprising the expression cassette comprising any of the peptides as described herein and the peptide is operably linked to a promoter.
  • the present disclosure provides a host cell comprising the vector comprising the expression cassette.
  • the present disclosure provides a composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any of the peptides, the vector comprising the expression cassette, or the host cell; and (2) a pharmaceutically acceptable excipient.
  • the composition further comprises an adjuvant.
  • the composition comprises a plurality of peptides each comprising a T cell epitope set forth in Table 3 or Tables 15-47.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
  • the present disclosure provides a method of eliciting an immune response in a subject in need thereof, the method comprising administering to the subject an effective amount of a composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any
  • the method comprises administering the composition to the subject by a route selected from the group consisting of, subcutaneous, intramuscular, and oral.
  • the subject is at risk of exposure to SARS-CoV, SARS-CoV-2 or other coronavirus infections.
  • the subject is at risk of exposure to SARS-CoV-2 infections.
  • the subject is at risk of developing severe illnesses when contracted with SARS-CoV-2 (e.g., COVID-19) .
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
  • the present disclosure provides a kit for eliciting an immune response in a subject in need thereof, the kit comprising a first container containing the composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any of the peptides, or the vector comprising the expression cassette.
  • the kit optionally contains an additional container containing a therapeutic agent against SARS-CoV-2. In some embodiments, the kit further comprises at least a second container each containing at least one different composition.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids.
  • the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
  • the present disclosure provides a method for detecting T cell immunity against SARS-CoV-2 in a subject, comprising: (1) contacting T cells obtained from the subject with a T cell epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having an HLA allele associated with the epitope; and (2) detecting activation of the T cells, thereby detecting presence of T cell immunity against SARS-CoV-2 in the subject.
  • step (2) of the method comprises detection of T cell proliferation or T cell secretion of one or more cytokines.
  • step (2) of the method comprises T cell proliferation assay, flow cytometry, ELISPOT, or ELISA.
  • step (1) of the method comprises contacting T cells obtained from the subject with a composition comprising a plurality of peptides each comprising a T cell epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having HLA alleles associated with each of the plurality of epitopes.
  • FIGS. 1A-B show comparison of the similarity of structural proteins of SARS-CoV-2 with the corresponding proteins of SARS-CoV and MERS (Middle East Respiratory Syndrome) -CoV.
  • FIG. 1A shows the percentage genetic similarity of the individual structural proteins of SARS-CoV-2 with those of SARS-CoV and MERS-CoV. The reference sequence of each coronavirus (Materials and Methods) was used to calculate the percentage genetic similarity.
  • FIG. 1B is a circular phylogram of the phylogenetic trees of the four structural proteins. All trees were constructed based on the available unique sequences using PASTA and rooted with the outgroup Zaria Bat CoV strain (accession ID: HQ166910.1) .
  • FIGS. 2A-C depict location of SARS-CoV S protein subunits and SARS-CoV-derived B cell epitopes on the protein structure (PDB ID: 5XLR) .
  • FIG. 2A shows subunits S1 and S2 are indicated in medium and light grey color, respectively. The receptor binding motif lies within the S1 subunit and is indicated in dark grey color.
  • FIG. 2B shows residues of the linear B cell epitopes, that were identical in SARS-CoV-2 (Table 4) , are shown as shaded. The black and grey colors reflect the surface and buried residues, respectively.
  • FIG. 2C shows locations of discontinuous B cell epitopes that share at least one identical residue with corresponding SARS-CoV-2 sites (Table 5) . Identical epitope residues are shown in dark grey color, while the remaining epitope residues are shown in medium grey color. Both the side view (left panel) and the top view (right panel) of the structure are shown.
  • FIG. 3 shows fraction of mutations in the observed sequences of the structural proteins of the three coronaviruses. Mutation is defined here as an amino acid difference from the reference sequence of the respective coronavirus; accession IDs: NC_045512.2 (2019-nCoV) , NC_004718.3 (SARS-CoV) , and NC_019843.3 (MERS-CoV) .
  • FIG. 4 illustrates location of identified T cell epitopes on the SARS-CoV S protein structure (PDB ID: 5XLR) . Residues of the SARS-CoV-derived T cell epitopes (determined using positive MHC binding assays and that were identical in SARS-CoV-2) are shown with filled color. The dark and light shade reflect the surface and buried residues, respectively.
  • FIG. 5 shows pairwise sequence alignment of the reference sequences of the S proteins of SARS-CoV and SARS-CoV-2 (accession ID: NP_828851.1 and YP_009724390.1, respectively) . Identical residues are indicated by *.
  • FIGS. 6A-C show global diversity of HLA class I alleles and the corresponding T cell responses, and summary of the experimentally-determined SARS-CoV-2 CD8+ T cell epitope data.
  • FIG. 6A shows different HLA class I alleles are prevalent in different geographical regions (left panel) ; heatmaps depicting the diverse HLA allele distribution in North America, North East Asia, and Australia are shown as examples (middle panel) .
  • Each square in the heatmap represents a distinct HLA class I allele, and its color shade represents the frequency of the allele in the geographical region with a dark shade representing high frequency and vice versa; and different HLA alleles present different SARS-CoV-2 peptides, and consequently the T cell responses elicited against SARS-CoV-2 are expected to differ among geographical regions (right panel) .
  • peptide pools to test for SARS-CoV-2 CD8+ T cell responses need to be designed specific to the HLA alleles prevalent in the population being tested. This figure was created with BioRender. com.
  • FIG. 6B shows the number of HLA alleles in each geographical region for which at least one SARS-CoV-2 experimentally-determined CD8+ T cell epitope is known.
  • FIG. 6C shows the number of experimentally-determined epitopes within spike and other SARS-CoV-2 proteins determined so far for class I HLA alleles (left panel) and the frequencies of these HLA alleles in different geographical regions (right panel) (for details of experimentally-determined epitope data, see Methods) .
  • FIGS . 7A-B show an optimized in silico strategy to predict SARS-CoV-2 CD8 + T cell epitopes.
  • FIG. 7A shows a performance comparison of in silico HLA class I epitope prediction methods based on predicting experimentally-determined SARS-CoV-2 CD8 + T cell epitopes. Intersection of the top predictions of the analysed 12 in silico methods showed that MHCflurry2.0P predicts the highest number
  • FIGS. 8A-B show a schematic of SARS2TPools framework and snapshot of the platform’s interface for a use case.
  • FIG. 8A is a schematic showing how the SARS2TPools, based on user-selected options (protein, geographical region, specific HLA alleles) , combines experimentally-determined epitope data, in silico predictions, and information of prevalent HLA alleles across regions to obtain optimized peptide pools for assessing vaccine-induced T cell responses. This figure was created with BioRender. com.
  • 8B is a snapshot of the SARS2TPools interface for the “Region-specific” tab with selected options (Region: Oceania; Protein: NSP12; Pool-size of peptides: small; Threshold for in silico predictions: Default; and Length of peptides: 9) .
  • FIGS. 9A-C are summary of region-specific CD8+ T cell peptide pools provided by SARS2TPools.
  • FIG. 9A shows the number of peptides in region specific pools from experimental studies and in silico predictions.
  • FIG. 9B shows overlap between pairs of region-specific pools.
  • FIG. 9C shows a source of peptides for each region-specific pool.
  • Each circle represents one of the top 30 prevalent HLAs for a specific region and the color of the circle indicates whether the peptides were obtained from only experimental studies, or only from in silico predictions, or from both.
  • FIG. 10 shows histograms of ranks assigned by different in silico epitope prediction methods to the experimentally-determined SARS-CoV-2 CD8 + T cell epitopes associated with the 10 HLA alleles with the most experimental data (see Methods) .
  • Each bin here represents a set of peptides within the specified rank range of all 10 HLA alleles.
  • FIGS. 11A-D show the observation that predictions of MHCflurry2.0P are distinct from those of other in silico methods (FIG. 12) was robust to the size of the set of top-ranked predictions compared. Distinct sets of predicted epitopes ranked by the considered 12 in silico methods in their (FIG. 11A) top 10, (FIG. 11B) top 15, (FIG. 11C) top 20, and (FIG. 11D) top 25 predictions corresponding to each of the 10 HLA alleles having the most experimental data (see Methods) .
  • FIG. 12 shows the strategy of combining top-ranked predictions of MHCflurry2.0P with any of the other 11 in silico methods considered in this study had a higher hit-rate on experimentally-determined SARS-CoV-2 CD8 + epitope data than either of the individual methods.
  • hit-rate represents the fraction of experimentally known epitopes present in the set of top-ranked predicted peptides.
  • FIG. 13 shows hit-rates of union methods comprising of MHCflurry2.0P and one of the other 11 in silico epitope prediction methods. While none of these clearly outperformed the rest, the 4 union methods comprising of MHCflurry2.0P and either NetMHCpan4.1BA, NetMHCpan4.0BA, NetMHC4.0, or NetMHCpan4.1EL were among the top 4 performing methods (also see Table 13) .
  • hit-rate represents the fraction of experimentally known epitopes present in the set of top-ranked predicted peptides.
  • FIGS. 14A-B illustrate COVIDep providing an up-to-date set of B-cell and T-cell epitopes that can serve as potential vaccine targets for SARS-CoV-2.
  • FIG. 14A shows the identified epitopes are experimentally-derived from SARS-CoV and have a close genetic match with the available SARS-CoV-2 sequences.
  • FIG. 14B is an example of the T-cell epitopes reported by COVIDep (as of 20 May 2020) for the spike protein of SARS-CoV-2.
  • the Search box in the top right was used to select only the HLA-A*02: 01-restricted epitopes.
  • FIG. 15 is an epitope screening protocol used by COVIDep for providing vaccine target recommendations for SARS-CoV-2.
  • COVIDep periodically pools SARS-CoV-2 sequence data from the GISAID database (www. gisaid. org) and compares with experimentally-determined T cell and B cell epitopes of SARS-CoV, obtained from the ViPR database (www. viprbrc. org) .
  • the T cell epitopes were determined based on either positive T cell assays or positive MHC binding assays for SARS-CoV.
  • For the B cell epitopes both linear and discontinuous epitopes were considered.
  • the system outputs those epitopes that are genetically similar in SARS-CoV-2, based on an epitope screening parameter.
  • This user-defined parameter allows the user to select epitopes based on their conservation in the SARS-CoV-2 sequence data, where conservation is defined as the fraction of SARS-CoV-2 sequences with the exact epitope sequence.
  • the value of this parameter is set to 0.95 as default; however, the user may change this value to adjust the stringency of the screening criterion. For example, reducing the value of the parameter will allow for the consideration of epitopes with greater genetic variation, potentially increasing the set of recommended SARS-CoV-2 vaccine targets.
  • the population coverage analysis tool available at IEDB www. iedb. org
  • IEDB the population coverage analysis tool available at IEDB (www. iedb. org) is used to estimate the percentage of a specified population that can elicit a response against them.
  • nucleic acid or protein when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It is preferably in a homogeneous state although it can be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified. In particular, an isolated gene is separated from open reading frames that flank the gene and encode a protein other than the gene of interest.
  • purified denotes that a nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. Particularly, it means that the nucleic acid or protein is at least 60%, 70%, 80%, 85%, 90%, 95%, or 99%pure, more preferably at least 95%pure, and most preferably at least 99%pure.
  • amino acid refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids.
  • Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, ⁇ -carboxyglutamate, and O-phosphoserine.
  • Unnatural (non-naturally occurring) amino acids include, without limitation, amino acid analogs, amino acid mimetics, synthetic amino acids, N-substituted glycines, and N-methyl amino acids in either the L-or D-configuration that function in a manner similar to the naturally-occurring amino acids.
  • Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an ⁇ carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid.
  • Amino acid mimetics refers to chemical compounds that have a structure different from the general chemical structure of an amino acid, but capable of functioning in a manner similar to a naturally occurring amino acid.
  • nucleic acid or “polynucleotide” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single-or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences as well as the sequence explicitly indicated.
  • degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19: 5081 (1991) ; Ohtsuka et al., J. Biol. Chem. 260: 2605-2608 (1985) ; and Rossolini et al., Mol. Cell. Probes 8: 91-98 (1994) ) .
  • the term nucleic acid is used interchangeably with gene, cDNA, or mRNA encoded by a gene.
  • a downstream location is one at the 3' side of a reference point
  • an "upstream” location is one at the 5' side of a reference point.
  • polypeptide, ” “peptide, ” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues.
  • the terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers.
  • the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens) , wherein the amino acid residues are linked by covalent peptide bonds.
  • the amino acid sequence of a polypeptide is presented from the N-terminus to the C-terminus.
  • the first amino acid from the N-terminus is referred to as the “first amino acid. ”
  • heterologous refers to the relationship of one peptide fusion partner to the another peptide fusion partner: the manner in which the fusion partners are present in the fusion peptide is not one that can be found a naturally occurring protein.
  • a "heterologous polypeptide" fused with a T cell or B cell epitope to form a fusion peptide may be one that is originated from a protein other than the antigen from which the T cell or B cell epitope is derived, such as a granulocyte-macrophage colony-stimulating factor (GM-CSF) .
  • GM-CSF granulocyte-macrophage colony-stimulating factor
  • a “heterologous polypeptide” may be one derived from another portion of the T cell or B cell protein that is not immediately contiguous to the T cell or B cell epitope.
  • a “heterologous polypeptide” may contain modifications of a naturally occurring protein sequence or a portion thereof, such as deletions, additions, or substitutions of one or more amino acid residues.
  • the fusion peptide should not contain a subsequence of the human T cell or B cell that encompasses the amino acid sequences in Table 3 and Tables 15-47 and have more than 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more amino acids in length.
  • the fusion peptide should not contain a subsequence of the human T cell or B cell that encompasses the amino acid sequences in Table 3 and Tables 15-47 and have more than 500 amino acids in length.
  • a "heterologous polypeptide" for use in the present disclosure has no more than 15-20 amino acids in length; in other embodiments, a “heterologous polypeptide” has at least 100 amino acids in length.
  • a "heterologous polypeptide" is a recombinant polynucleotide (or a copy or complement of a recombinant polynucleotide) that has been manipulated using well known methods.
  • a "heterologous polypeptide” can comprise a recombinant expression cassette comprising a promoter operably linked to a second polynucleotide (e.g., a coding sequence of the antigen from which the T cell or B cell epitope is derived) .
  • the promoter is heterologous to the second polynucleotide as the result of human manipulation (e.g., by methods described in Sambrook et al., Molecular Cloning -A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, (1989) or Current Protocols in Molecular Biology Volumes 1-3, John Wiley &Sons, Inc. (1994-1998) ) .
  • the "heterologous polypeptide” may comprise a promoter that is heterologous to a second polynucleotide encoding the polypeptide of interest (e.g., a T cell and/or a B cell epitope) or a fragment thereof, and a third polynucleotide encoding a detecting (e.g., a tag or identifier) molecule.
  • the detecting molecule can be a fluorescent protein, e.g., a green fluorescent protein (GFP) or a variant or the like, such as DsRed and other red fluorescent protein.
  • fuse or “fused, " as used in the context of describing a peptide of this disclosure that comprises a T cell or B cell epitope joined with a heterologous polypeptide, refers to a connection between the epitope and the heterologous polypeptide by any covalent bond, including a peptide bond.
  • a nucleic acid sequence encoding refers to a nucleic acid which contains sequence information for a structural RNA such as rRNA, a tRNA, or the primary amino acid sequence of a specific protein or peptide, or a binding site for a trans-acting regulatory agent. This phrase specifically encompasses degenerate codons (i.e., different codons which encode a single amino acid) of the native sequence or sequences that may be introduced to conform to codon preference in a specific host cell.
  • An “expression cassette” is a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular polynucleotide sequence in a host cell.
  • An expression cassette may be part of a plasmid, viral genome, or nucleic acid fragment.
  • an expression cassette includes a polynucleotide to be transcribed, operably linked to a promoter.
  • recombinant when used with reference, e.g., to a cell, or nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a nucleic acid or protein from an outside source or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified.
  • recombinant cells express genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under-expressed or not expressed at all.
  • administration refers to various methods of contacting a substance with a mammal, especially a human.
  • Modes of administration may include, but are not limited to, methods that involve contacting the substance intravenously, intraperitoneally, intranasally, transdermally, topically, subcutaneously, parentally, intramuscularly, orally, or systemically, and via injection, ingestion, inhalation, implantation, or adsorption by any other means.
  • T cell and/or B cell peptide of this disclosure or a fusion peptide comprising a T cell and/or B cell epitope e.g., a T cell or B cell epitope derived from the S protein, N protein, or full length of SARS-CoV and/or SARS-CoV-2) and a heterologous polypeptide is via intramuscular delivery, where the peptide or fusion peptide can be formulated as a pharmaceutical composition in the form suitable for intramuscular injection, such as an aqueous solution, a suspension, or an emulsion, etc.
  • Other means for delivering a T cell and/or B cell epitope or a fusion peptide of this disclosure includes intradermal injection, subcutaneous injection, intravenous injection, or transdermal application as with a patch.
  • an "effective amount" of a certain substance refers to an amount of the substance that is sufficient to effectuate a desired result.
  • an effective amount of a composition comprising a peptide of this disclosure that is intended to induce an anti-SARS-CoV or SARS-CoV-2 immunity is an amount sufficient to achieve the goal of inducing the immunity when administered to a subject.
  • the effect to be achieved may include the prevention, correction, or inhibition of progression of the symptoms of a disease/condition and related complications to any detectable extent.
  • the exact quantity of an "effective amount” will depend on the purpose of the administration, and can be ascertainable by one skilled in the art using known techniques (see, e.g., Lieberman, Pharmaceutical Dosage Forms (vols. 1-3, 1992) ; Lloyd, The Art, Science and Technology of Pharmaceutical Compounding (1999) ; and Pickar, Dosage Calculations (1999) ) .
  • a “therapeutically effective amount” of a substance/molecule of the disclosure may vary depending on factors such as: the disease state, age, sex and weight of the individual, and the ability of the substance/molecule to induce a desired response in the individual.
  • a therapeutically effective amount is also an amount that has a therapeutically beneficial effect over any toxic or detrimental effect of the substance/molecule.
  • prophylactically effective amount refers to an amount of dosage and time necessary to effectively achieve the desired prophylactic result. Since prophylactic doses are used in individuals prior to or early in the disease, the prophylactically effective amount is typically (but not necessarily) less than the therapeutically effective amount.
  • a “physiologically or pharmaceutically acceptable excipient” is an inert ingredient used in the formulation of a composition of this disclosure, which contains the active ingredient (s) of a T cell and/or B cell peptide or a fusion peptide comprising a T cell and/or B cell peptide and a heterologous polypeptide and is suitable for use, e.g., by injection into a patient in need thereof.
  • This inert ingredient may be a substance that, when included in a composition of this disclosure, provides a desired pH, consistency, color, smell, or flavor of the composition.
  • T cell immune response refers to activation of antigen specific T cells as measured by proliferation or expression of molecules on the cell surface or secretion of proteins such as cytokines.
  • B cell immune response refers to either a T cell-independent immune response or a T cell-dependent immune response.
  • B cells respond directly to the antigen.
  • B cells rely on the assistance from T cells to respond.
  • Activated B cells may express IgA, IgE, IgG or retain IgM expression.
  • Cytokines produced by T cells and others may determine what isotype the B cells express.
  • Several general techniques are commonly used to identify antigen-specific B cells. Non-limiting examples are B cell enzyme linked immunospot (ELISPOT) , limiting dilution, flow cytometry, adoptive transfer, microscopy, and B cell receptor (BCR) transgenic mice.
  • ELISPOT B cell enzyme linked immunospot
  • BCR B cell receptor
  • the term “immune response” generally refers to the immune system recognizing the antigens, e.g., proteins, on the surface of substances or microorganisms, such as bacteria or virus, and attacks and destroys, or tries to destroy, them.
  • the term “immune response” refers to a cell-mediated (T-cell) immune response and/or an antibody (B-cell) immune response.
  • immunogenic composition As described herein, the terms “immunogenic composition, ” or “vaccine” are used interchangeably and refer to a composition that elicits an immune response in a subject, especially a human.
  • An immunogenic composition or vaccine can be used prophylactically to prevent COVID-19 (SARS-CoV-2) or other coronavirus diseases (e.g., SARS-CoV) .
  • the term “immunity” refers to protection from an infectious disease. For instance, if a subject is immune to a disease, the object can be exposed to it without becoming infected.
  • the term “vaccine” generally refers to a preparation that is used to stimulate the body’s immune system to generate immunity for a disease.
  • it refers to an antigen-containing formulation consisting of whole pathogenic organisms (killed or attenuated) or components of these organisms (such as proteins, peptides or polysaccharides) for conferring immunity against disease caused by these organisms.
  • Vaccine formulations may be natural, synthetic or obtained by recombinant DNA techniques. A vaccine may be administered in any route known in the field.
  • vaccines are usually administered intramuscularly through needle injections.
  • vaccines can be administered by mouth (oral) or sprayed into the nose (intranasal) .
  • immunogenic refers to the ability of an immunogen, antigen or vaccine to stimulate an immune response.
  • an “antigen” is defined as any substance capable of eliciting an immune response.
  • an “antigen” may be a small molecule or a macromolecule such as a protein, a peptide, a polysaccharide, a nucleic acid, a lipid, or a biomolecule.
  • antigen-specific refers to a property of a population of cells such that the provision of a particular antigen or antigen fragment causes the proliferation of a specific cell.
  • the specified antibodies bind to a particular protein at least two times the background and do not substantially bind in a significant amount to other proteins present in the sample.
  • Specific binding to an antibody under such conditions may require an antibody that is selected for its specificity for a particular protein.
  • polyclonal antibodies raised to fusion proteins can be selected to obtain only those polyclonal antibodies that are specifically immunoreactive with fusion protein and not with individual components of the fusion proteins. This selection may be achieved by subtracting out antibodies that cross-react with the individual antigens.
  • a variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein.
  • solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow &Lane, Antibodies, A Laboratory Manual (1988) , for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity) .
  • a specific or selective reaction will be at least twice background signal or noise and more typically more than 10 to 100 times background.
  • antibody and “immunoglobulin” are used interchangeably in a broad sense and include monoclonal antibodies (e.g., full-length or intact monoclonal antibodies) , polyclonal antibodies, multivalent antibodies, multispecific antibodies (e.g., bispecific antibodies, so long as they exhibit the desired biological activity) , and may also include certain antibody fragments (as described in more detail herein) .
  • the antibody may be a chimeric antibody, a human antibody, a humanized antibody, and/or an affinity matured antibody.
  • the term “about” denotes a range of ⁇ 10%of a specified value. For instance, “about 10” denotes a range of 9-11.
  • subject or “subject in need of treatment” refers to an individual who seeks medical attention due to risk of, or actual sufferance from, a condition involving undesirable inflammation (e.g., pneumonia or an infection that inflames air sacs) or a condition involving infection of the respiratory tract (e.g., the upper respiratory tract, the lungs) .
  • Subjects or individuals in need of treatment include those that demonstrate symptoms of infection of the respiratory tract or those are at risk of later developing the disease or disorder and/or its symptoms.
  • the subject may experience or is at risk of sufferance from a SARS-CoV infection.
  • the subject may have one or more symptoms of SARS-CoV infection as defined by the Centers for Disease Control and Prevention (www. cdc.
  • the subject may experience or have a wide range of symptoms ranging from mild to severe illness. Exemplary symptoms are fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or running nose, nausea or vomiting, or diarrhea, trouble breathing, persistent pain or pressure in the chest, new confusion, inability to wake or stay awake, and/or pale, gray, or blue-colored skin, lips, or nail beds, depending on skin tone, or combinations thereof. Symptoms may appear 0-60 days, 1-30 days, or 2-14 days after exposure to the source of infection (e.g., SARS-CoV or SARS-CoV-2 viruses) .
  • the term subject can include both animals, especially mammals, and humans.
  • peptides comprise at least one T cell epitope and/or at least one B cell epitope for eliciting an immune response in a subject against SARS-CoV-2. Epitopes that provide broad coverage broadly and region-specific are also explored.
  • SARS-CoV-2 is referred to the original SARS-CoV-2 (GenBank accession no. NC_045512) or any variants thereof.
  • a variant has one or more mutations that differentiate it from other variants of the SARS-CoV-2 viruses.
  • Current known variants include, but are not limited to, Alpha (B. 1.1.7 and Q lineages) , Beta (B. 1.351 and descendent lineages) , Gamma (P. 1 and descendent lineages) , Epsilon (B. 1.427 and B. 1.429) , Eta (B. 1.525) , Iota (B. 1.526) , Kappa (B.
  • a polynucleotide of the full length SARS-CoV-2 (GenBank accession no. NC_045512) , a variant, or a fragment thereof may activates a T cell and/or B cell response.
  • the polynucleotides may have at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more sequence identity with the full length SARS-CoV-2 (GenBank accession no. NC_045512) , a variant, or a fragment thereof.
  • peptides of the present disclosure may be synthesized chemically using conventional peptide synthesis or other protocols well known in the art.
  • the peptides are of no more than about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids, or in the range of from about 10 to about 50, 100, 200, 300, 400 or 500 amino acids, from about 20 to about 50, 100, 200, 300, 400 or 500 amino acids, from about 50 to about 100, 150, 200, 300, 400 or 500 amino acids, from about 100 to about 150, 200, 250, 300, 400 or 500 amino acids, or from about 200 to about 250, 300, 350, 400 or 500 amino acids.
  • Peptides may be synthesized by solid-phase peptide synthesis methods using procedures similar to those described by Merrifield et al., J. Am. Chem. Soc., 85: 2149-2156 (1963) ; Barany and Merrifield, Solid-Phase Peptide Synthesis, in The Peptides: Analysis, Synthesis, Biology Gross and Meienhofer (eds. ) , Academic Press, N.Y., vol. 2, pp. 3-284 (1980) ; and Stewart et al., Solid Phase Peptide Synthesis 2nd ed., Pierce Chem. Co., Rockford, Ill. (1984) .
  • N- ⁇ - protected amino acids having protected side chains are added stepwise to a growing polypeptide chain linked by its C-terminal and to a solid support, i.e., polystyrene beads.
  • the peptides are synthesized by linking an amino group of an N- ⁇ -deprotected amino acid to an ⁇ -carboxy group of an N- ⁇ -protected amino acid that has been activated by reacting it with a reagent such as dicyclohexylcarbodiimide. The attachment of a free amino group to the activated carboxyl leads to peptide bond formation.
  • the most commonly used N- ⁇ -protecting groups include Boc, which is acid labile, and Fmoc, which is base labile.
  • halomethyl resins such as chloromethyl resin or bromomethyl resin
  • hydroxymethyl resins such as phenol resins, such as 4- ( ⁇ - [2, 4-dimethoxyphenyl] -Fmoc-aminomethyl) phenoxy resin
  • tert-alkyloxycarbonyl-hydrazidated resins such as 4- ( ⁇ - [2, 4-dimethoxyphenyl] -Fmoc-aminomethyl) phenoxy resin
  • tert-alkyloxycarbonyl-hydrazidated resins and the like.
  • the C-terminal N- ⁇ -protected amino acid is first attached to the solid support.
  • the N- ⁇ -protecting group is then removed.
  • the deprotected ⁇ -amino group is coupled to the activated ⁇ -carboxylate group of the next N- ⁇ -protected amino acid.
  • the process is repeated until the desired peptide is synthesized.
  • the resulting peptides are then cleaved from the insoluble polymer support and the amino acid side chains deprotected. Longer peptides can be derived by condensation of protected peptide fragments.
  • nucleic acids sizes are given in either kilobases (kb) or base pairs (bp) . These are estimates derived from agarose or acrylamide gel electrophoresis, from sequenced nucleic acids, or from published DNA sequences.
  • kb kilobases
  • bp base pairs
  • proteins sizes are given in kilodaltons (kDa) or amino acid residue numbers. Proteins sizes are estimated from gel electrophoresis, from sequenced proteins, from derived amino acid sequences, or from published protein sequences.
  • Oligonucleotides that are not commercially available can be chemically synthesized, e.g., according to the solid phase phosphoramidite triester method first described by Beaucage &Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981) , using an automated synthesizer, as described in Van Devanter et. al., Nucleic Acids Res. 12: 6159-6168 (1984) . Purification of oligonucleotides is performed using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange HPLC as described in Pearson &Reanier, J. Chrom. 255: 137-149 (1983) .
  • Recombinant production is an effective means to obtain peptides of this disclosure, particularly those of relatively large molecular weight, for example, a fusion peptide of a HER-2/Neu epitope and a GM-CSF.
  • the sequence of a polynucleotide encoding a peptide of this disclosure, and synthetic oligonucleotides can be verified after cloning or subcloning using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981) .
  • a polynucleotide sequence encoding a peptide of this disclosure can be obtained by chemical synthesis, or can be purchased from a commercial supplier, which may then be further manipulated using standard techniques of molecular cloning.
  • the polynucleotide sequence encoding a peptide of this disclosure can be optionally altered to coincide with the preferred codon usage of a particular host.
  • the preferred codon usage of one strain of bacterial cells can be used to derive a polynucleotide that encodes a peptide of the disclosure and includes the codons favored by this strain.
  • the frequency of preferred codon usage exhibited by a host cell can be calculated by averaging frequency of preferred codon usage in a large number of genes expressed by the host cell (e.g., calculation service is available from web site of the Kazusa DNA Research Institute, Japan) . This analysis is preferably limited to genes that are highly expressed by the host cell.
  • the coding sequences are verified by sequencing and are then subcloned into an appropriate expression vector for recombinant production of the peptides of this disclosure.
  • the peptide of the present disclosure can be produced using routine techniques in the field of recombinant genetics.
  • a strong promoter to direct transcription e.g., in Sambrook and Russell, supra, and Ausubel et al., supra.
  • Bacterial expression systems for expressing a peptide of this disclosure are available in, e.g., E. coli, Bacillus sp., Salmonella, and Caulobacter. Kits for such expression systems are commercially available.
  • Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available.
  • the eukaryotic expression vector is an adenoviral vector, an adeno-associated vector, a retroviral vector, or a yeast artificial chromosome (YAC) .
  • the eukaryotic (e.g., mammalian) expression vector is a pEAK10, pEAK12, pEAK13, pCDNA3.0, pCDNA4.0, pCDM7, pCDM8, pCDM10, pCDM12.
  • the promoter is CMV, EF1 ⁇ , or SV40.
  • the expression system comprises a promoter operably linked to the peptide nucleic acid (e.g., cDNA) sequence such that it is under control of a promoter.
  • promoter describes the combination of the promoter (RNA polymerase binding site) and operators.
  • the promoter may function in vivo or in a cell-free system. Any number of promoters may be used depending on the needs and preferences of the practitioner. Promoters for controlling recombinant protein, (e.g., peptide) expression can be any promoter for any DNA dependent RNA polymerase, for example.
  • a promoter e.g., T7, T3, and SP6 RNA promoters and compatible RNA polymerases
  • a promoter is selected for in vitro expression for producing recombinant protein in a bacterial system, e.g., E. coli, which usually requires the molecular inducer isopropyl- ⁇ -D-thiogalactoside (IPTG) for regulating the promoter’s transcriptional activity.
  • IPTG molecular inducer isopropyl- ⁇ -D-thiogalactoside
  • a modified Self-Inducible Expression system that utilizes lactose as an inducer may be used, as described in Briand et al, A self-inducible heterologous protein expression system in Escherichia coli. Sci Rep 6, 33037 (2016) .
  • the promoter used to direct expression of a heterologous nucleic acid depends on the particular application.
  • the promoter is optionally positioned about the same distance from the heterologous transcription start site as it is from the transcription start site in its natural setting. As is known in the art, however, some variation in this distance can be accommodated without loss of promoter function.
  • the expression vector typically includes a transcription unit or expression cassette that contains all the additional elements required for the expression of a peptide of this disclosure in host cells.
  • a typical expression cassette thus contains a promoter operably linked to the polynucleotide sequence encoding the peptide and signals required for efficient polyadenylation of the transcript, ribosome binding sites, and translation termination.
  • the nucleic acid sequence encoding the peptide is typically linked to a cleavable signal peptide sequence to promote secretion of the peptide by the transformed cell.
  • signal peptides include, among others, the signal peptides from tissue plasminogen activator, insulin, and neuron growth factor, and juvenile hormone esterase of Heliothis virescens.
  • Additional elements of the cassette may include enhancers and, if genomic DNA is used as the structural gene (e.g., encoding the heterologous polypeptide) , introns with functional splice donor and acceptor sites.
  • the expression cassette should also contain a transcription termination region downstream of the structural gene to provide for efficient termination.
  • the termination region may be obtained from the same gene as the promoter sequence or may be obtained from different genes.
  • the particular expression vector used to transport the genetic information into the cell is not particularly critical. Any of the conventional vectors used for expression in eukaryotic or prokaryotic cells may be used. Standard bacterial expression vectors include plasmids such as pBR322 based plasmids, pSKF, pET23D, and fusion expression systems such as GST and LacZ. Epitope tags can also be added to recombinant proteins to provide convenient methods of isolation, e.g., c-myc.
  • Expression vectors containing regulatory elements from eukaryotic viruses are typically used in eukaryotic expression vectors, e.g., SV40 vectors, papilloma virus vectors, and vectors derived from Epstein-Barr virus.
  • eukaryotic vectors include pMSG, pAV009/A + , pMTO10/A + , pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV40 early promoter, SV40 later promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.
  • Some expression systems have markers that provide gene amplification such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase.
  • markers that provide gene amplification such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase.
  • high yield expression systems not involving gene amplification are also suitable, such as a baculovirus vector in insect cells, with a polynucleotide sequence encoding the peptide of this disclosure under the direction of the polyhedrin promoter or other strong baculovirus promoters.
  • the elements that are typically included in expression vectors also include a replicon that functions in E. coli, a gene encoding antibiotic resistance to permit selection of bacteria that harbor recombinant plasmids, and unique restriction sites in nonessential regions of the plasmid to allow insertion of eukaryotic sequences.
  • the particular antibiotic resistance gene chosen is not critical, any of the many resistance genes known in the art are suitable.
  • the prokaryotic sequences are optionally chosen such that they do not interfere with the replication of the DNA in eukaryotic cells, if necessary. Similar to antibiotic resistance selection markers, metabolic selection markers based on known metabolic pathways may also be used as a means for selecting transformed host cells.
  • the expression vector further comprises a sequence encoding a secretion signal, such as the E. coli OppA (Periplasmic Oligopeptide Binding Protein) secretion signal or a modified version thereof, which is directly connected to 5'of the coding sequence of the protein to be expressed.
  • a secretion signal such as the E. coli OppA (Periplasmic Oligopeptide Binding Protein) secretion signal or a modified version thereof, which is directly connected to 5'of the coding sequence of the protein to be expressed.
  • This signal sequence directs the recombinant protein produced in cytoplasm through the cell membrane into the periplasmic space.
  • the expression vector may further comprise a coding sequence for signal peptidase 1, which is capable of enzymatically cleaving the signal sequence when the recombinant protein is entering the periplasmic space.
  • Standard transfection methods are used to produce bacterial, mammalian, yeast, insect, or plant cell lines that express large quantities of a peptide of this disclosure, which are then purified using standard techniques (see, e.g., Colley et al., J. Biol. Chem. 264: 17619-17622 (1989) ; Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990) ) . Transformation of eukaryotic and prokaryotic cells are performed according to standard techniques (see, e.g., Morrison, J. Bact. 132: 349-351 (1977) ; Clark-Curtiss &Curtiss, Methods in Enzymology 101: 347-362 (Wu et al., eds, 1983) .
  • Any of the well-known procedures for introducing foreign nucleotide sequences into host cells may be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, liposomes, microinjection, plasma vectors, viral vectors and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into a host cell (see, e.g., Sambrook and Russell, supra) . It is only necessary that the particular genetic engineering procedure used be capable of successfully introducing at least one gene into the host cell capable of expressing the peptide of this disclosure.
  • the transfected cells are cultured under conditions favoring expression of the peptide of this disclosure.
  • the cells are then screened for the expression of the recombinant peptide, which is subsequently recovered from the culture using standard techniques (see, e.g., Scopes, Protein Purification: Principles and Practice (1982) ; U.S. Patent No. 4,673,641; Ausubel et al., supra; and Sambrook and Russell, supra) .
  • gene expression can be detected at the nucleic acid level.
  • a variety of methods of specific DNA and RNA measurement using nucleic acid hybridization techniques are commonly used (e.g., Sambrook and Russell, supra) .
  • Some methods involve an electrophoretic separation (e.g., Southern blot for detecting DNA and Northern blot for detecting RNA) , but detection of DNA or RNA can be carried out without electrophoresis as well (such as by dot blot) .
  • the presence of nucleic acid encoding a peptide of this disclosure in transfected cells can also be detected by PCR or RT-PCR using sequence-specific primers.
  • gene expression can be detected at the polypeptide level.
  • Various immunological assays are routinely used by those skilled in the art to measure the level of a gene product, particularly using polyclonal or monoclonal antibodies that react specifically with a peptide of the present disclosure, particularly one containing a sufficiently large heterolougs polypeptide (e.g., Harlow and Lane, Antibodies, A Laboratory Manual, Chapter 14, Cold Spring Harbor, 1988; Kohler and Milstein, Nature, 256: 495-497 (1975) ) .
  • Such techniques require antibody preparation by selecting antibodies with high specificity against the peptide or an antigenic portion thereof.
  • an initial salt fractionation can separate many of the unwanted host cell proteins (or proteins derived from the cell culture media) from the recombinant protein of interest, e.g., a peptide of the present disclosure.
  • the preferred salt is ammonium sulfate.
  • Ammonium sulfate precipitates proteins by effectively reducing the amount of water in the protein mixture. Proteins then precipitate on the basis of their solubility. The more hydrophobic a protein is, the more likely it is to precipitate at lower ammonium sulfate concentrations.
  • a typical protocol is to add saturated ammonium sulfate to a protein solution so that the resultant ammonium sulfate concentration is between 20-30%.
  • a protein of greater and lesser size can be isolated using ultrafiltration through membranes of different pore sizes (for example, Amicon or Millipore membranes) .
  • the protein mixture is ultrafiltered through a membrane with a pore size that has a lower molecular weight cut-off than the molecular weight of a protein of interest, e.g., a peptide of the present disclosure.
  • the retentate of the ultrafiltration is then ultrafiltered against a membrane with a molecular cut off greater than the molecular weight of the peptide of interest.
  • the recombinant protein will pass through the membrane into the filtrate.
  • the filtrate can then be chromatographed as described below.
  • a protein of interest (such as a peptide of the present disclosure) can also be separated from other proteins on the basis of its size, net surface charge, hydrophobicity, or affinity for ligands.
  • antibodies raised against a peptide of this disclosure can be conjugated to column matrices and the peptide immunopurified. All of these methods are well known in the art.
  • inclusion bodies typically involves the extraction, separation and/or purification of inclusion bodies by disruption of bacterial cells, e.g., by incubation in a buffer of about 100-150 ⁇ g/ml lysozyme and 0.1%Nonidet P40, a non-ionic detergent.
  • the cell suspension can be ground using a Polytron grinder (Brinkman Instruments, Westbury, NY) .
  • the cells can be sonicated on ice. Alternate methods of lysing bacteria are described in Ausubel et al. and Sambrook and Russell, both supra, and will be apparent to those of skill in the art.
  • the cell suspension is generally centrifuged and the pellet containing the inclusion bodies resuspended in buffer which does not dissolve but washes the inclusion bodies, e.g., 20 mM Tris-HCl (pH 7.2) , l mM EDTA, 150 mM NaCl and 2%Triton-X 100, a non-ionic detergent. It may be necessary to repeat the wash step to remove as much cellular debris as possible.
  • the remaining pellet of inclusion bodies may be resuspended in an appropriate buffer (e.g., 20 mM sodium phosphate, pH 6.8, 150 mM NaCl) .
  • an appropriate buffer e.g., 20 mM sodium phosphate, pH 6.8, 150 mM NaCl
  • Other appropriate buffers will be apparent to those of skill in the art.
  • the inclusion bodies are solubilized by the addition of a solvent that is both a strong hydrogen acceptor and a strong hydrogen donor (or a combination of solvents each having one of these properties) .
  • a solvent that is both a strong hydrogen acceptor and a strong hydrogen donor or a combination of solvents each having one of these properties.
  • the proteins that formed the inclusion bodies may then be renatured by dilution or dialysis with a compatible buffer.
  • Suitable solvents include, but are not limited to, urea (from about 4 M to about 8 M) , formamide (at least about 80%, volume/volume basis) , and guanidine hydrochloride (from about 4 M to about 8 M) .
  • Some solvents that are capable of solubilizing aggregate-forming proteins may be inappropriate for use in this procedure due to the possibility of irreversible denaturation of the proteins, accompanied by a lack of immunogenicity and/or activity.
  • SDS sodium dodecyl sulfate
  • 70%formic acid Some solvents that are capable of solubilizing aggregate-forming proteins, such as SDS (sodium dodecyl sulfate) and 70%formic acid, may be inappropriate for use in this procedure due to the possibility of irreversible denaturation of the proteins, accompanied by a lack of immunogenicity and/or activity.
  • guanidine hydrochloride and similar agents are denaturants, this denaturation is not irreversible and renaturation may occur upon removal (by dialysis, for example) or dilution of the denaturant, allowing re-formation of the immunologically and/or biologically active protein of interest.
  • the protein can be separated from other bacterial proteins by standard separation techniques.
  • recombinant polypeptides e.g., a peptide of this disclosure
  • the periplasmic fraction of the bacteria can be isolated by cold osmotic shock in addition to other methods known to those of skill in the art (see e.g., Ausubel et al., supra) .
  • the bacterial cells are centrifuged to form a pellet. The pellet is resuspended in a buffer containing 20%sucrose.
  • the bacteria are centrifuged and the pellet is resuspended in ice-cold 5 mM MgSO 4 and kept in an ice bath for approximately 10 minutes.
  • the cell suspension is centrifuged and the supernatant decanted and saved.
  • the recombinant peptides present in the supernatant can be separated from the host proteins by standard separation techniques well known to those of skill in the art.
  • a recombinant polypeptide e.g., a peptide of the present disclosure
  • its purification can follow the standard protein purification procedure described below. This standard purification procedure is also suitable for purifying peptides obtained from chemical synthesis.
  • the amino acid sequence of a peptide of this disclosure can be confirmed by a number of well established methods.
  • the conventional method of Edman degradation can be used to determine the amino acid sequence of a peptide.
  • sequencing methods based on Edman degradation including microsequencing, and methods based on mass spectrometry are also frequently used for this purpose.
  • the peptides of the present disclosure can be modified to achieve more desirable properties.
  • the design of chemically modified peptides and peptide mimics that are resistant to degradation by proteolytic enzymes or have improved solubility or binding ability is well known.
  • Modified amino acids or chemical derivatives of the T cell or B cell peptides or fusion peptides of this disclosure may contain additional chemical moieties of modified amino acids not normally a part of the T cell or B cell protein.
  • Covalent modifications of the peptides are within the scope of the present disclosure. Such modifications may be introduced into a peptide by reacting targeted amino acid residues of the peptide with an organic derivatizing agent that is capable of reacting with selected side chains or terminal residues.
  • the following examples of chemical derivatives are provided by way of illustration and not by way of limitation.
  • modifications include substitution of a natural amino acid with an unnatural hydroxylated amino acid, substitution of the carboxy groups in acidic amino acids with nitrile derivatives, substitution of the hydroxyl groups in basic amino acids with alkyl groups, or substitution of methionine with methionine sulfoxide.
  • an amino acid of a HER-2/Neu peptide or a fusion peptide of this disclosure can be replaced by the same amino acid but of the opposite chirality, i.e., a naturally-occurring L-amino acid may be replaced by its D-configuration.
  • Full length recombinant whole or variants thereof of SARS-CoV or SARS-CoV-2 virus and/or other coronavirous, or recombinant functional proteins, e.g., S protein, N protein, M protein, from SARS-CoV or SARS-CoV-2, and/or other coronavirous may be generated and purified using technologies known in the field.
  • full length SARS-CoV-2 can be generated based on “transformation-associated recombination” (TAR) in yeast (Thao et al. (2020) , “Rapid reconstruction of SARS-CoV-2 using a synthetic genomics platform, ” . bioRxiv, 2020.02.21.959817) .
  • synthetic proteins e.g., defined as >35 amino acids
  • peptides from the SARS-CoV or SARS-CoV-2 recombinant binding domain (RBD) may be synthesized in the Peptide Core at Los Alamos National Laboratory.
  • Positive control rabbit serum (polyclonal, against SARS/SARS-CoV-2 Coronavirus spike protein subunit 1) is Invitrogen PA5-81795. See e.g., Schein et al., 2021. “Synthetic proteins for COVID-19 diagnostics. ” Peptides. 143: 170583. doi: 10.1016/j. peptides. 2021.170583.
  • any protein fragment meaning a polypeptide sequence at least one amino acid residue shorter than a reference polypeptide sequence but otherwise identical
  • a reference protein having a length of 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500 or longer than 500 amino acids.
  • any protein that includes a stretch of 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 (contiguous) amino acids that are 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%identical to any of the sequences described herein can be utilized in accordance with the disclosure.
  • a polypeptide includes 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutations as shown in any of the sequences provided herein or referenced herein.
  • a peptide corresponding to a SARS-CoV or SARS-CoV-2 promiscuous T cell epitope or B cell epitope is attached to a heterologous polypeptide via a covalent bond to form a fusion peptide, such that the ability of the SARS-CoV or SARS-CoV-2 epitope to induce a T cell or B cell response is enhanced.
  • this covalent bond is a peptide bond and the SARS-CoV and/or SARS-CoV-2 epitope and the heterologous polypeptide form a new polypeptide.
  • This peptide bond may be a direct peptide bond between the SARS-CoV and/or SARS-CoV-2 epitope and the heterologous polypeptide, or it may be an indirect peptide bond provided by way of a peptide linker between the SARS-CoV or SARS-CoV-2 epitope and the heterologous polypeptide.
  • the SARS-CoV and/or SARS-CoV-2 epitope may be linked to a detecting molecule such as a GFP, RFP, dsRed, or the like.
  • the SARS-CoV and/or SARS-CoV-2 epitope may be linked to an affinity tag such as a six histidine (6His) on the N-terminal or the C-terminal, or the like.
  • covalent bonds are also suitable for the purpose of fusing the SARS-CoV or SARS-CoV-2 peptide with the heterologous polypeptide.
  • a functional group such as a non- terminal amine group, a non-terminal carboxylic acid group, a hydroxyl group, and a sulfhydryl group
  • one peptide may easily react with a functional group of the other peptide and establish a covalent bond, other than a peptide bond, that conjugates the two peptides.
  • a covalent connection between a peptide of a SARS-CoV or SARS-CoV-2 epitope and a heterologous polypeptide can also be provided by way of a linker molecule with suitable functional group (s) .
  • a linker molecule can be a peptide linker or a non-peptide linker.
  • a linker may be derivatized to expose or to attach additional reactive functional groups prior to conjugation. The derivatization may involve attachment of any of a number of molecules such as those available from Pierce Chemical Company, Rockford, Illinois.
  • T lymphocytes express T-cell receptors (TCR) and detect short peptides produced by a proteolytic machinery and presented by major histocompatibility complex (MHC) molecules at the cell surface. (Bercovici et al., 2000) .
  • T lymphocytes are able to detect foreign peptides synthesized by infected cells.
  • T cells which express CD8 molecules have the capacity to lyse directly the target cells.
  • a subset of CD4 + T lymphocytes is specialized in regulating the immune response via cytokine secretion and activation of the antigen-presenting cells (APC) .
  • APC antigen-presenting cells
  • Chromium release assays and limiting-dilution analyses have been commonly used to measure specific T-cell responses. Additionally, various assays are available for immune monitoring of specific T-cell responses. These various assays are schematically divided into functional assays, which measure the secretion of a particular cytokine (ELISPOT and intracellular cytokines) ; assays which assess the specificity of the T cells irrespective of their functionality and which are based on structural features of the TCR (tetramers and immunoscope) ; and assays aimed at detecting T-cell precursors by amplifying cells that proliferate in response to antigenic stimulation.
  • ELISPOT cytokine
  • intracellular cytokines assays which assess the specificity of the T cells irrespective of their functionality and which are based on structural features of the TCR (tetramers and immunoscope)
  • assays aimed at detecting T-cell precursors by amplifying cells that proliferate in response to antigenic stimulation aimed at detecting T-cell precursors
  • a SARS-CoV or SARS-CoV-2 epitope of this disclosure (or a fusion peptide comprising a SARS-CoV or SARS-CoV-2 epitope and a heterologous polypeptide) is useful for its capability to induce a T cell immune response specific to a SARS-CoV or SARS-CoV-2 protein, when the epitope is presented by an APC that may have one of a HLA-A, HLA-B, and/or HLA-DR allele.
  • Various functional assays can be used to confirm the ability of a SARS-CoV or SARS-CoV-2 epitope to induce such a SARS-CoV or SARS-CoV-2 specific T cell immune response in a promiscuous manner with regard to antigen presenting cells of different HLA alleles, including proliferation assay and flow cytometry assays detecting the binding between a T cell receptor and a peptide epitope or the production of cytokines by T cells.
  • the function assay ELISPOT may be used for this purpose.
  • the ELISPOT (enzyme-linked immunospot) technique detects T cells that secrete a given cytokine (e.g., gamma interferon [IFN- ⁇ ] ) in response to an antigenic stimulation (e.g., an antigen presented by SARS-CoV or SARS-CoV-2) .
  • a given cytokine e.g., gamma interferon [IFN- ⁇ ]
  • an antigenic stimulation e.g., an antigen presented by SARS-CoV or SARS-CoV-2
  • T cells are cultured with antigen-presenting cells in wells which have been coated with anti-IFN- ⁇ antibodies.
  • the secreted IFN- ⁇ is captured by the coated antibody and then revealed with a second antibody coupled to a chromogenic substrate.
  • cytokine molecules form spots, with each spot corresponding to one IFN- ⁇ -secreting cell.
  • the number of spots allows one to determine the frequency of IFN- ⁇ -secreting cells specific for a given antigen in the analyzed sample.
  • the ELISPOT assay has also been described for the detection of tumor necrosis factor alpha, interleukin-4 (IL-4) , IL-5, IL-6, IL-10, IL-12, granulocyte-macrophage colony-stimulating factor, and granzyme B-secreting lymphocytes.
  • T cells recognize short peptides presented by MHC molecules through their clonotypic TCR.
  • Tetramers of MHC class I-peptide complexes have been used in cytometry to enumerate, characterize, and purify peptide-specific CD8 cells.
  • the heavy and light chains of the MHC are produced in Escherichia coli, solubilized in urea, and refolded in vitro in the presence of high concentrations of the antigenic peptide.
  • the refolded complexes are purified by gel filtration, and a single biotin is added at the C-terminal end of the heavy chain using the bacterial BirA enzyme. Incubation with fluorescent streptavidin yields tetramers which can be used like any clonotypic antibody.
  • Tetramers of MHC class II molecules may also be produced and used to analyze CD4 + T-cell responses.
  • B lymphocytes recognize intact proteins and produce immunoglobulins (Ig) .
  • Ig immunoglobulins
  • B cell response is activated in two pathways: T cell-dependent and T cell-independent.
  • T cell-dependent manner antigens that activate T cells as well as B cells establish Ig responses in which T cells provide ‘help’ for the B cells to mature.
  • T cell-independent B cell activation occurs without the assistance of T cell co-stimulatory proteins.
  • monomeric antigens are unable to activate B cells.
  • Polymeric antigens with a repeating structure are able to activate B cells, probably because they can crosslink and cluster Ig molecules on the B cell surface.
  • T cell-independent antigens include bacterial lipopolysaccharide (LPS) , certain other polymeric polysaccharides, and certain polymeric proteins.
  • LPS lipopolysaccharide
  • BCR B cell receptor
  • Another example for confirming B cell immune response specific to a SARS-CoV or SARS-CoV-2 protein is flow cytometry-based analysis of antigen-specific B cells.
  • This technique is dependent on labeling antigen with a fluorescent tag to allow detection.
  • Fluorochromes can either be attached covalently via chemical conjugation to the antigen, expressed as a recombinant fusion protein, or attached non-covalently by biotinylating the antigen. After biotinylation, fluorochrome-conjugated streptavidin is added to generate a labeled tetramer of the antigen.
  • Biotinylation of the antigen may be set at a ratio ⁇ 1 biotin to 1 antigen.
  • site directed biotinylation can be accomplished by adding either an AviTag or BioEase tag to the recombinant antigen prior to expression.
  • the present disclosure further provides a method for eliciting an immune response in a subject in need thereof such as at risk of exposure to SARS-CoV or SARS-CoV-2 infection.
  • the term exposure refers to contact with infectious agents (bacteria or virus, e.g., SARS-CoV or SARS-CoV-2) in a manner that promotes transmission and increases the likihood of disease.
  • infectious agents bacteria or virus, e.g., SARS-CoV or SARS-CoV-2
  • close contact refers to within 6 feet of someone for a cumulative total of 15 minites or more over a 24-hour period.
  • Subjects who have underlying medical conditions e.g., immune-comprimised due to treatment of immunosuppressants, under chemotherapy, prior heart and cardio diseases, prior pulmonary diseases, cancer, diabetes, obesity, etc.
  • age e.g., 65 or over
  • geneteic predisposition e.g., geneteic predisposition
  • the subject may be at risk for severe illness when contracted with COVID-19 such that the subject may need hospitalization, intensive care, a ventilator to help them breathe or they may even die. Elicitation an immune response in a subject may provide prophylactic and therapeutic applications.
  • the method includes the following steps: first, lymphocytes including at least a T cell epitope, and/or at least a B cell epitope, and/or optionally an antigen-presenting cell are obtained from a patient. Suitable samples that yield such lymphocytes include blood, and lymph nodes or lymphatic fluids.
  • At least a T cell epitope, and/or at least a B cell epitope, and/or optionally an antigen-presenting cell are exposed to a SARS-CoV or SARS-CoV-2 peptide (or a fusion peptide comprising the SARS-CoV or SARS-CoV-2 peptide and a heterologous peptide) of this disclosure under conditions that would allow, e.g., proper presentation of a T cell epitope by the antigen-presenting cell to the T cell.
  • signs of a T cell response and/or B cell response is measured in vitro by means well known in the art such as ELISPOT, ELISA, proliferation assay, or flow cytometry. When a T cell response and/or B cell response is detected by any of these methods, it can be concluded that there exists a T cell and/or a B cell immune response specific to a SARS-CoV or SARS-CoV-2 protein in the patient.
  • vaccines are provided.
  • the vaccines will generally comprise a peptide comprising one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2, such as those discussed above, in combination with an immunostimulant.
  • An immunostimulant may be any substance that enhances or potentiates an immune response (antibody and/or cell-mediated) to an exogenous antigen.
  • immunostimulants include adjuvants, biodegradable microspheres (e.g., polylactic galactide) and liposomes (into which the compound is incorporated; see, e.g., U.S.
  • compositions and vaccines within the scope of the present disclosure may also contain other compounds, which may be biologically active or inactive.
  • one or more T cell and/or B cell epitopes derived from immunogenic portions of other coronavirus antigens e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV- NL63
  • MERS-CoV e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV- NL63
  • a pharmaceutical vaccine formulation may comprise the petide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2, one or more adjuvnts, pharmaceutically-acceptable carriers or other ingredients including immunological adjuvants routinely provided in vaccine formulation. Suitable adjuvants are described in detail below.
  • the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is biologically produced such as using an expression cassette in a host cell.
  • Short to medium length peptides, for example peptides that do no require specific folding, can be chemically synthesized.
  • Large scale production of chemically synthesized peptides may be used for manufacturing large quantities of peptide vaccines. See e.g., Bray, B.L., 2003. Large-scale manufacture of peptide therapeutics by chemical synthesis. Nature Reviews Drug Discovery, 2 (7) , pp. 587-593.
  • the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is a cationic peptide.
  • a “cationic peptide” refers to a peptide that is positively charged at a pH in the range of 5.0 to 8.0.
  • the net charge on the peptide or peptide cocktails is calculated by assigning a +1 charge for each lysine (K) , arginine (R) or histidine (H) , a -1 charge for each aspartic acid (D) or glutamic acid (E) and a charge of 0 for the other amino acid within the sequence.
  • the charge contributions from the N-terminal amine (+1) and C-terminal carboxylate (-1) end groups of each peptide effectively cancel each other when unsubstituted.
  • the charges are summed for each peptide and expressed as the net average charge.
  • a suitable peptide has a net average positive charge of +1.
  • the peptide has a net positive charge in the range that is larger than +2.
  • the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is an anionic peptide.
  • an “anionic molecule” refers to a molecule that is negatively charged at a pH in the range of 5.0-8.0. The net negative charge on the oligomer or polymer is calculated by assigning a -1 charge for each phosphodiester or phosphorothioate group in the oligomer.
  • Illustrative vaccines may contain polynucleotides encoding one or more of the polypeptides (e.g., T cell and/or B cell epitopes derived from one or more immunogenic portions of a SARS-CoV, SARS-CoV-2, and/or other coronaviruses) as described above, such that the polypeptide is generated in situ.
  • the polynucleotides may be present within any of a variety of delivery systems known to those of ordinary skill in the art, including nucleic acid expression systems, bacteria and viral expression systems. Numerous gene delivery techniques are well known in the art, such as those described by Rolland, Crit. Rev. Therap. Drug Carrier Systems 15: 143-198 (1998) , and references cited therein.
  • Appropriate nucleic acid expression systems contain the necessary polynucleotides sequences for expression in the patient (such as a suitable promoter and terminating signal) .
  • Bacterial delivery systems involve the administration of a bacterium (such as Bacillus-Calmette-Guerrin) that expresses an immunogenic portion of the polypeptide on its cell surface or secretes such an epitope.
  • the polynucleotides may be introduced using a viral expression system (e.g., vaccinia or other pox virus, retrovirus, or adenovirus) , which may involve the use of a non-pathogenic (defective) , replication competent virus.
  • polynucleotides may also be “naked, ” as described, for example, in Ulmer et al., Science 259: 1745-1749 (1993) and reviewed by Cohen, Science 259: 1691-1692 (1993) .
  • the uptake of naked polynucleotides may be increased by coating the polynucleotides onto biodegradable beads, which are efficiently transported into the cells.
  • a vaccine may comprise both a polynucleotide and a polypeptide component. Such vaccines may provide for an enhanced immune response.
  • the polynucleotides is formulated in lipid nanoparticles.
  • the lipid nanoparticle is a mucus penetrating lipid nanoparticle.
  • the lipid nanoparticle is a solid lipid nanoparticle. Formulations for polynucleotides in lipid nanoparticles are discussed in U.S. Patent No: 10272150, which is incorporated herein in its entirety.
  • a vaccine may contain pharmaceutically acceptable salts of the polynucleotides and polypeptides provided herein.
  • Such salts may be prepared from pharmaceutically acceptable non-toxic bases, including organic bases (e.g., salts of primary, secondary and tertiary amines and basic amino acids) and inorganic bases (e.g., sodium, potassium, lithium, ammonium, calcium and magnesium salts) .
  • the peptide vaccines and/or the polynucleotide vaccines of the disclosure are superior to conventional vaccines by a factor of at least 2 fold, 5 fold, 10 fold, 20 fold, 40 fold, 50 fold, 100 fold, 500 fold or 1,000 fold.
  • compositions of the present disclosure may be formulated for any appropriate manner of administration, including for example, topical, oral, nasal, intravenous, intracranial, intraperitoneal, subcutaneous or intramuscular administration.
  • parenteral administration such as subcutaneous injection
  • the carrier preferably comprises water, saline, alcohol, a fat, a wax or a buffer.
  • any of the above carriers or a solid carrier such as mannitol, lactose, starch, magnesium stearate, sodium saccharine, talcum, cellulose, glucose, sucrose, and magnesium carbonate, may be employed.
  • Biodegradable microspheres may also be employed as carriers for the pharmaceutical compositions of this disclosure.
  • Suitable biodegradable microspheres are disclosed, for example, in U.S. Patent Nos. 4,897,268; 5,075,109; 5,928,647; 5,811,128; 5,820,883; 5,853,763; 5,814,344 and 5,942,252.
  • One may also employ a carrier comprising the particulate-protein complexes described in U.S. Patent No. 5,928,647, which are capable of inducing a class I-restricted cytotoxic T lymphocyte responses in a host.
  • compositions may also comprise buffers (e.g., neutral buffered saline or phosphate buffered saline) , carbohydrates (e.g., glucose, mannose, sucrose or dextrans) , mannitol, proteins, polypeptides or amino acids such as glycine, antioxidants, bacteriostats, chelating agents such as EDTA or glutathione, adjuvants (e.g., aluminum hydroxide) , solutes that render the formulation isotonic, hypotonic or weakly hypertonic with the blood of a recipient, suspending agents, thickening agents and/or preservatives.
  • buffers e.g., neutral buffered saline or phosphate buffered saline
  • carbohydrates e.g., glucose, mannose, sucrose or dextrans
  • mannitol proteins
  • proteins polypeptides or amino acids
  • proteins e.glycine
  • antioxidants e.g., mannito
  • an adjuvant or an immune potentiator may be included.
  • An adjuvant may act as a co-signal to prime T-cells and/or B-cells and/or NK cells as to the existence of an infection.
  • Adjuvants useful in the present disclosure may include, but are not limited to, natural or synthetic. They may be organic or inorganic.
  • adjuvants useful in the present disclosure include adjuvants for vaccines (e.g., influenza vaccines) as shown in Table 48.
  • Adjuvants for DNA nucleic acid vaccines (DNA) have been disclosed in, for example, Kobiyama, et al Vaccines, 2013, 1 (3) , 278-292, the contents of which are incorporated herein by reference in their entirety.
  • adjuvants contain a substance designed to protect the antigen from rapid catabolism, such as aluminum hydroxide or mineral oil, and a stimulator of immune responses, such as lipid A, Bortadella pertussis or Mycobacterium species or Mycobacterium derived proteins.
  • a stimulator of immune responses such as lipid A, Bortadella pertussis or Mycobacterium species or Mycobacterium derived proteins.
  • delipidated, deglycolipidated M. vaccae “pVac”
  • pVac deglycolipidated M. vaccae
  • Suitable adjuvants are commercially available as, for example, Freund’s Incomplete Adjuvant and Complete Adjuvant (Difco Laboratories, Detroit, MI) ; Merck Adjuvant 65 (Merck and Company, Inc., Rahway, NJ) ; AS-2 and derivatives thereof (SmithKline Beecham, Philadelphia, PA) ; CWS, TDM, Leif, aluminum salts such as aluminum hydroxide gel (alum) or aluminum phosphate; salts of calcium, iron or zinc; an insoluble suspension of acylated tyrosine; acylated sugars; cationically or anionically derivatized polysaccharides; polyphosphazenes; biodegradable microspheres; monophosphoryl lipid A and quil A. Cytokines, such as GM-CSF or interleukin-2, -7, or -12, may also be used as adjuvants.
  • Cytokines such as GM-CSF or interleukin-2, -7, or -12, may also be
  • adjuvants may include, without limitation, cationic liposome-DNA complex JVRS-100, aluminum hydroxide vaccine adjuvant, aluminum phosphate vaccine adjuvant, aluminum potassium sulfate adjuvant, alhydrogel, ISCOM (s) TM , Freund's Complete Adjuvant, Freund's Incomplete Adjuvant, CpG DNA Vaccine Adjuvant, Cholera toxin, Cholera toxin B subunit, Liposomes, Saponin Vaccine Adjuvant, DDA Adjuvant, Squalene-based Adjuvants, Etx B subunit Adjuvant, IL-12 Vaccine Adjuvant, LTK63 Vaccine Mutant Adjuvant, TiterMax Gold Adjuvant, Ribi Vaccine Adjuvant, Corynebacterium-derived P40 Vaccine Adjuvant, MPL TM Adjuvant, AS04, AS02, Lipopolysaccharide Vaccine
  • adjuvants which may be co-administered with the polypeptides of the disclosure include, but are not limited to interferons, TNF-alpha, TNF-beta, chemokines such as CCL21, eotaxin, HMGB1, SA100-8alpha, GCSF, GMCSF, granulysin, lactoferrin, ovalbumin, CD-40L, CD28 agonists, PD-1, soluble PD1, L1 or L2, or interleukins such as IL-1, IL-2, IL-4, IL-6, IL-7, IL-10, IL-12, IL-13, IL-21, IL-23, IL-15, IL-17, and IL-18.
  • interferons such as CCL21, eotaxin, HMGB1, SA100-8alpha
  • GCSF GMCSF
  • granulysin lactoferrin
  • ovalbumin CD-40L
  • CD28 agonists CD28 agonists
  • the adjuvant composition is preferably designed to induce an immune response predominantly of the Th1 type.
  • High levels of Th1-type cytokines e.g., IFN- ⁇ , TNF ⁇ , IL-2 and IL-12
  • Th2-type cytokines e.g., IL-4, IL-5, IL-6 and IL-10
  • a patient will support an immune response that includes Th1-and Th2-type responses.
  • Th1-type cytokines will increase to a greater extent than the level of Th2-type cytokines.
  • the levels of these cytokines may be readily assessed using standard assays. For a review of the families of cytokines, see Mosmann &Coffman, Ann. Rev. Immunol. 7: 145-173 (1989) .
  • Preferred adjuvants for use in eliciting a predominantly Th1-type response include, for example, a combination of monophosphoryl lipid A, preferably 3-de-O-acylated monophosphoryl lipid A (3D-MPL) , together with an aluminum salt.
  • MPL adjuvants are available from Corixa Corporation (Seattle, WA; see US Patent Nos. 4,436,727; 4,877,611; 4,866,034 and 4,912,094) .
  • CpG-containing oligonucleotides in which the CpG dinucleotide is unmethylated also induce a predominantly Th1 response.
  • oligonucleotides are well known and are described, for example, in WO 96/02555, WO 99/33488 and U.S. Patent Nos. 6,008,200 and 5,856,462. Immunostimulatory DNA sequences are also described, for example, by Sato et al., Science 273: 352 (1996) .
  • Another preferred adjuvant comprises a saponin, such as Quil A, or derivatives thereof, including QS21 and QS7 (Aquila Biopharmaceuticals Inc., Framingham, MA) ; Escin; Digitonin; or Gypsophila or Chenopodium quinoa saponins .
  • Other preferred formulations include more than one saponin in the adjuvant combinations of the present disclosure, for example combinations of at least two of the following group comprising QS21, QS7, Quil A, ⁇ -escin, or digitonin.
  • the saponin formulations may be combined with vaccine vehicles composed of chitosan or other polycationic polymers, polylactide and polylactide-co-glycolide particles, poly-N-acetyl glucosamine-based polymer matrix, particles composed of polysaccharides or chemically modified polysaccharides, liposomes and lipid-based particles, particles composed of glycerol monoesters, etc.
  • vaccine vehicles composed of chitosan or other polycationic polymers, polylactide and polylactide-co-glycolide particles, poly-N-acetyl glucosamine-based polymer matrix, particles composed of polysaccharides or chemically modified polysaccharides, liposomes and lipid-based particles, particles composed of glycerol monoesters, etc.
  • the saponins may also be formulated in the presence of cholesterol to form particulate structures such as liposomes or ISCOMs.
  • the saponins may be formulated together with a polyoxyethylene ether or ester, in either a non-particulate solution or suspension, or in a particulate structure such as a paucilamelar liposome or ISCOM.
  • the saponins may also be formulated with excipients such as Carbopol R to increase viscosity, or may be formulated in a dry powder form with a powder excipient such as lactose.
  • the adjuvant system includes the combination of a monophosphoryl lipid A and a saponin derivative, such as the combination of QS21 and adjuvant, as described in WO 94/00153, or a less reactogenic composition where the QS21 is quenched with cholesterol, as described in WO 96/33739.
  • a monophosphoryl lipid A and a saponin derivative such as the combination of QS21 and adjuvant, as described in WO 94/00153
  • a less reactogenic composition where the QS21 is quenched with cholesterol
  • Other preferred formulations comprise an oil-in-water emulsion and tocopherol.
  • Another particularly preferred adjuvant formulation employing QS21, adjuvant and tocopherol in an oil-in-water emulsion is described in WO 95/17210.
  • Another enhanced adjuvant system involves the combination of a CpG-containing oligonucleotide and a saponin derivative particularly the combination of CpG and QS21 as disclosed in WO 00/09159.
  • the formulation additionally comprises an oil in water emulsion and tocopherol.
  • adjuvants include Montanide ISA 720 (Seppic, France) , SAF-1 (Chiron, California, United States) , ISCOMS (CSL) , MF-59 (Chiron) , the SBAS series of adjuvants (e.g., SBAS-2, AS2’, AS2, ” SBAS-4, or SBAS6, available from SmithKline Beecham, Rixensart, Belgium) , Detox (Corixa, Hamilton, MT) , RC-529 (Corixa, Hamilton, MT) and other aminoalkyl glucosaminide 4-phosphates (AGPs) , such as those described in pending U.S. Patent Application Serial Nos. 08/853,826 and 09/074, 720, the disclosures of which are incorporated herein by reference in their entireties, and polyoxyethylene ether adjuvants such as those described in WO 99/52549A1.
  • adjuvants include adjuvant molecules of the general formula (I) : HO (CH 2 CH 2 O) n -A-R, wherein, n is 1-50, A is a bond or –C (O) -, R is C 1-50 alkyl or Phenyl C 1-50 alkyl.
  • One embodiment of the present disclosure consists of a vaccine formulation comprising a polyoxyethylene ether of general formula (I) , wherein n is between 1 and 50, preferably 4-24, most preferably 9; the R component is C 1-50 , preferably C 4 -C 20 alkyl and most preferably C 12 alkyl, and A is a bond.
  • the concentration of the polyoxyethylene ethers should be in the range 0.1-20%, preferably from 0.1-10%, and most preferably in the range 0.1-1%.
  • Preferred polyoxyethylene ethers are selected from the following group: polyoxyethylene-9-lauryl ether, polyoxyethylene-9-steoryl ether, polyoxyethylene-8-steoryl ether, polyoxyethylene-4-lauryl ether, polyoxyethylene-35-lauryl ether, and polyoxyethylene-23-lauryl ether.
  • Polyoxyethylene ethers such as polyoxyethylene lauryl ether are described in the Merck index (12 th edition: entry 7717) . These adjuvant molecules are described in WO 99/52549.
  • polyoxyethylene ether according to the general formula (I) above may, if desired, be combined with another adjuvant.
  • a preferred adjuvant combination is preferably with CpG as described in the pending UK patent application GB 9820956.2.
  • Any vaccine provided herein may be prepared using well known methods that result in a combination of antigen, immune response enhancer and a suitable carrier or excipient.
  • the compositions described herein may be administered as part of a sustained release formulation (i.e., a formulation such as a capsule, sponge or gel (composed of polysaccharides, for example) that effects a slow release of compound following administration) .
  • a sustained release formulation i.e., a formulation such as a capsule, sponge or gel (composed of polysaccharides, for example) that effects a slow release of compound following administration
  • Such formulations may generally be prepared using well known technology (see, e.g., Coombes et al., Vaccine 14: 1429-1438 (1996) ) and administered by, for example, oral, rectal or subcutaneous implantation, or by implantation at the desired target site.
  • Sustained-release formulations may contain a polypeptide, polynucleotide or antibody dispersed in a carrier matrix and/or contained within a reservoir surrounded by a rate controlling membrane.
  • the vaccine is formulated for induction of systemic or localized mucosal immunity through immunogen entrapment and co-administration with microparticles.
  • Carriers for use within such formulations are biocompatible, and may also be biodegradable; preferably the formulation provides a relatively constant level of active component release.
  • Such carriers include microparticles of poly (lactide-co-glycolide) , polyacrylate, latex, starch, cellulose, dextran and the like.
  • Other delayed-release carriers include supramolecular biovectors, which comprise a non-liquid hydrophilic core (e.g., a cross-linked polysaccharide or oligosaccharide) and, optionally, an external layer comprising an amphiphilic compound, such as a phospholipid (see, e.g., U.S. Patent No.
  • Vaccines and pharmaceutical compositions may be presented in unit-dose or multi-dose containers, such as sealed ampoules or vials. Such containers are preferably hermetically sealed to preserve sterility of the formulation until use.
  • formulations may be stored as suspensions, solutions or emulsions in oily or aqueous vehicles.
  • a vaccine or pharmaceutical composition may be stored in a freeze-dried condition requiring only the addition of a sterile liquid carrier immediately prior to use.
  • kits for conveniently and/or effectively carrying out methods of the present disclosure.
  • kits will comprise sufficient amounts and/or numbers of components to allow a user to perform multiple treatments of a subject (s) .
  • the disclosure provides kits for eliciting an immune response in a subject according to the method of the present disclosure.
  • the kits typically include a container that contains a pharmaceutical composition having an effective amount of SARS-CoV-derived or SARS-CoV-2-derived T cell epitope polypeptides or a function fragment thereof, and/or B cell epitope polypeptides or a function fragment thereof) , and/or the polypeptides optionally further comprising at least one heterologous polypeptide sequence.
  • the kit may comprise nucleic acid sequences encoding the SARS-CoV-derived or SARS-CoV-2-derived T cell and/or B cell polypeptides, and/or the expression cassettes for expressing the T cell and/or B cell polypeptides, and/or the vectors for expressing the expression cassettes.
  • the kit may optionally comprise an additional container containing a therapeutic agent against SARS-CoV or variants thereof, SARS-CoV-2 or variants thereof, and/or other coronavirus or variants (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) .
  • kits comprising the SARS-CoV-and/or SARS-CoV-2-derived T cell and/or B cell epitopes (including any polypeptides, proteins or function fragments thereof) of the disclosure.
  • the T cell and/or B cell epitopes may be derived from one or more functional proteins (e.g., S protein, N protein, or M protein) or full length SARS-CoV or SARS-CoV-2.
  • the kit further comprises polypeptides comprising one or more T cell and/or B cell epitopes derived from one or more of a SARS-CoV and/or SARS-CoV-2 variants.
  • the kit further comprises polypeptides comprising one or more T cell and/or B cell epitopes derived from one or more coronaviruses or variants thereof (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) .
  • the kit may further comprise packaging and instructions and/or a delivery agent to form a formulation composition.
  • the delivery agent may comprise a saline, a buffered solution, a lipidoid or any delivery agent disclosed herein.
  • kits can be for protein or polypeptide production, comprising a polynucleotide comprising a translatable region of one or more T cell and/or B cell epitopes derived from a functional protein (e.g., S protein, N protein, or M protein) or full length SARS-CoV or SARS-CoV-2.
  • the kit further comprises a polynucleotide comprising a translatable region of one or more T cell and/or B cell epitopes derived from one or more coronaviruses (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) .
  • the kit may further comprise packaging and instructions and/or a delivery agent to form a formulation composition.
  • the delivery agent may comprise a saline, a buffered solution, a lipidoid or any delivery agent disclosed herein.
  • the kit further comprises devices for administering the polypeptides or a protein or polypeptide product descried herein.
  • the kit may comprise syringes, needles, and instructions on how to dispense the pharmaceutical composition, including description of the type of patients who may be treated (e.g., subjects who are at risk of having COVID-19 due to individual health condition, pre-existing condition of any kind, compromised immune system, living condition, age, or occupation; subjects who has close contact with someone who has COVID-19) .
  • Coronaviruses are positive-sense single-stranded RNA viruses belonging to the family Coronaviridae. These viruses mostly infect animals, including birds and mammals. In humans, they generally cause mild respiratory infections, such as those observed in the common cold. However, some recent human coronavirus infections have resulted in lethal endemics, which include the SARS (Severe Acute Respiratory Syndrome) and MERS (Middle East Respiratory Syndrome) endemics. Both of these are caused by zoonotic coronaviruses that belong to the genus Betacoronavirus within Coronaviridae. SARS-CoV originated from Southern China and caused an endemic in 2003.
  • the recent SARS-CoV-2 belongs to the Betacoronavirus genus [12] . It has a genome size of ⁇ 30 kilobases which, like other coronaviruses, encodes for multiple structural and non-structural proteins.
  • the structural proteins include the spike (S) protein, the envelope (E) protein, the membrane (M) protein, and the nucleocapsid (N) protein.
  • SARS-CoV-2 being discovered very recently, there is currently a lack of immunological information available about the virus (e.g., information about immunogenic epitopes eliciting antibody or T cell responses) .
  • SARS-CoV-2 is quite similar to SARS-CoV based on the full-length genome phylogenetic analysis [9, 12] , and the putatively similar cell entry mechanism and human cell receptor usage [9, 13, 14] . Due to this apparent similarity between the two viruses, previous research that has provided an understanding of protective immune responses against SARS-CoV may potentially be leveraged to aid vaccine development for SARS-CoV-2.
  • T cell responses have been shown to provide long-term protection [21–23] , even up to 11 years post-infection [24] , and thus have also attracted interest for a prospective vaccine against SARS-CoV [reviewed in [25] ] .
  • SARS-CoV proteins T cell responses against the structural proteins have been found to be the most immunogenic in peripheral blood mononuclear cells of convalescent SARS-CoV patients as compared to the non-structural proteins [26] .
  • T cell responses against the S and N proteins have been reported to be the most dominant and long-lasting [27] .
  • SARS-CoV-derived B cell and T cell epitopes were searched on the NIAID Virus Pathogen Database and Analysis Resource (ViPR) (www. viprbrc. org/; accessed 21 February 2020) [29] by querying for the virus species name: “Severe acute respiratory syndrome-related coronavirus” from “human” hosts.
  • ViPR NIAID Virus Pathogen Database and Analysis Resource
  • Positive B cell assays e.g., enzyme-linked immunosorbent assay (ELISA) -based qualitative binding
  • T cell assays such as enzyme-linked immune absorbent spot (ELISPOT) or intracellular cytokine staining (ICS) IFN- ⁇ release
  • ELISPOT enzyme-linked immune absorbent spot
  • ICS intracellular cytokine staining
  • MHC major histocompatibility complex
  • SARS-CoV Severe Acute Respiratory Syndrome Coronavirus
  • Population-Coverage-Based T Cell Epitope Selection Population coverages for sets of T cell epitopes were computed using the tool provided by the Immune Epitope Database (IEDB) (tools. iedb. org/population/; accessed 21 February 2020) [30] .
  • This tool uses the distribution of MHC alleles (with at least 4-digit resolution, e.g., A*02: 01) within a defined population (obtained from allelefrequencies. net/) to estimate the population coverage for a set of T cell epitopes.
  • the estimated population coverage represents the percentage of individuals within the population that are likely to elicit an immune response to at least one T cell epitope from the set.
  • Structural Proteins of SARS-CoV-2 are Genetically Similar to SARS-CoV, but Not to MERS-CoV. SARS-CoV-2 has been observed to be close to SARS-CoV-much more so than MERS-CoV-based on full-length genome phylogenetic analysis [9, 12] . We checked whether this is also true at the level of the individual structural proteins (S, E, M, and N) . A straightforward reference-sequence-based comparison indeed confirmed this, showing that the M, N, and E proteins of SARS-CoV-2 and SARS-CoV have over 90%genetic similarity, while that of the S protein was notably reduced (but still high) (FIG. 1A) .
  • SARS-CoV-Derived T Cell Epitopes that ARE identical in SARS-CoV-2, and Determining those with Greatest Estimated Population Coverage.
  • the SARS-CoV-derived T cell epitopes used in this study were experimentally-determined from two different types of assays [29] : (i) Positive T cell assays, which tested for a T cell response against epitopes, and (ii) positive MHC binding assays, which tested for epitope-MHC binding. We aligned these T cell epitopes across the SARS-CoV-2 protein sequences.
  • Table 7 List of all SARS-CoV-derived T cell epitopes determined using positive MHC binding assays (with associated MHC allele information available at 4-digit resolution) and found to be identical in SARS-CoV-2.
  • Table 8 Distribution of all SARS-CoV-derived T cell epitopes obtained using positive MHC binding assays (with associated MHC allele information available at 4-digit resolution) and that are identical in SARS-CoV-2.
  • T cell epitopes for which a positive immune response has been determined using T cell assays were presented by the globally most-prevalent MHC allele (shown in bold text in Table 3) .
  • the functionally important epitopes located in the SARS-CoV receptor binding motif were associated with the second and third most-prevalent MHC alleles (underlined in Table 3) .
  • the ordering of T cell epitopes in Table 3 is based on the estimated global population coverage of the associated MHC alleles, it is also a natural order in which these epitopes should be tested experimentally for determining their potential to induce a positive immune response against SARS-CoV-2.
  • Table 9 Set of SARS-CoV-derived S and N protein T cell epitopes (obtained using positive MHC binding assays) that are identical in SARS-CoV-2 and that maximize estimated population coverage in China (86 distinct epitopes) .
  • Table 3 Set of the SARS-CoV-derived spike (S) and nucleocapsid (N) protein T cell epitopes (obtained from positive MHC binding assays) that are identical in SARS-CoV-2 and that maximize estimated population coverage globally (87 distinct epitopes) .
  • Table 10 Estimated global and Chinese population coverages for the individual SARS-CoV-derived S or N protein T cell epitopes (obtained using positive MHC binding assays) that are identical in SARS-CoV-2.
  • SARS-CoV-Derived B cell Epitopes that are Identical in SARS-CoV-2 . Similar to T cell epitopes, we used in our study the SARS-CoV-derived B cell epitopes that have been experimentally-determined from positive B cell assays [29] . These epitopes were classified as: (i) Linear B cell epitopes (antigenic peptides) , and (ii) discontinuous B cell epitopes (conformational epitopes with resolved structural determinants) .
  • SARS-CoV-derived linear B cell epitopes excluding those in S and N proteins, that are identical in SARS-CoV-2.
  • SARS-CoV-derived discontinuous B cell epitopes (and associated known antibodies [39–41] ) that have at least one site with an identical amino acid to the corresponding site in SARS-CoV-2.
  • SARS-CoV-specific antibodies S230, m396, and 80R known to bind to these epitopes in SARS-CoV might not be able to bind to the same regions in SARS-CoV-2 S protein.
  • this paper was under review, this has been confirmed experimentally [47] .
  • Further studies are currently under way to identify other SARS-CoV antibodies that may bind to discontinuous epitopes of the SARS-CoV-2 S protein [48] .
  • Beta (B. 1.351) , Gamma (P. 1) and Delta (B1.617.2) , that first emerged in South Africa, Brazil, and India respectively, are currently under investigation due to the observed reduction in neutralizing antibody titres against them in sera of vaccinated individuals (Abdool Karim and de Oliveira 2021; Liu et al. 2021; Madhi et al. 2021; Planas et al. 2021; Wall et al. 2021; Zhou et al. 2021) .
  • SARS-CoV-2 vaccines in use or development also stimulate T cell responses.
  • T cells There is increasing evidence of the role of T cells in protection from severe disease in SARS-CoV-2 infected patients (Altmann and Boyton 2020; Chen and John Wherry 2020; Liao et al. 2020; Mazzoni et al. 2020; Reynolds et al. 2020; Rydyznski Moderbacher et al. 2020; Wyllie et al. 2020; Bergamaschi et al. 2021; Bertoletti et al. 2021; Cohen et al. 2021) , and of their robustness to mutations associated with SARS-CoV-2 variants (Quadeer et al.
  • T cell responses have been coarsely measured using immune assays that stimulate blood samples of vaccinated individuals using overlapping peptide pools (Anderson et al. 2020; Folegatti et al. 2020; Keech et al. 2020; Logunov et al. 2020; Ramasamy et al. 2020; Sahin et al. 2020; Zhu et al. 2020; Ella et al. 2021; Klasse et al. 2021; Sadoff et al. 2021) .
  • T cell responses due to peptide competition, where immunogenic peptides compete with a large number of irrelevant peptides in the pool that are not recognized by T cells (Pala et al. 1988; Sahin et al. 2021) .
  • assays based on optimized peptide pools would be more efficient at estimating T cell responses as these comprise of a selected set of most relevant peptides against which a T cell response is expected to be stimulated.
  • Such pools also enable identifying precise T cell epitopes in the context of the associated human leukocyte antigen (HLA) alleles presenting them.
  • HLA human leukocyte antigen
  • T cells recognize peptides restricted by an individual’s HLA alleles which are highly diverse across the global population (albeit with some commonalities in a given region) . Consequently, the peptides restricted by these HLA alleles are also different, and hence T cell responses are expected to differ between geographical regions, even for the same vaccine.
  • current vaccines employ different SARS-CoV-2 antigens (e.g., based on the spike (S) protein only or employing the whole inactivated virion) , and these are expected to elicit distinct T cell responses.
  • SARS2TPools that provides optimized peptide pools for assessing region-specific vaccine-induced SARS-CoV-2 T cell responses. These pools are designed by exploiting information of prevalent HLA alleles in a population, the experimentally-determined and in-silico-predicted SARS-CoV-2 T cell epitopes associated with these alleles, and the antigen employed in the vaccine.
  • the optimized pools provided by SARS2TPools in addition to characterizing the vaccine-induced T cell responses in detail, can be useful for designing T cell based diagnostics, monitoring durability of T cell response, and any change in T cell responses due to emerging SARS-CoV-2 variants.
  • HLA class I and class II restricted SARS-CoV-2 T cell epitope data CD8 + and CD4 + , respectively
  • IEDB immune epitope database
  • HLA class II restricted epitopes all the available epitopes were 15 resides long, and only 3 HLA alleles had more than 20 epitopes in the data.
  • NetMHC4.0 (Andreatta and Nielsen 2016)
  • other common prediction methods that have been employed for predicting SARS-CoV-2 epitopes (Sohail et al. 2021) such as NetMHCpan3.0 (Nielsen and Andreatta 2016) , SMM (Peters and Sette 2005) , SMMPMBEC (Kim et al. 2009) , and IEDB consensus (Moutaftsi et al. 2006) .
  • MHCflurry2.0BA binding affinity and the presentation score predictions of MHCflurry as two separate methods.
  • MHCflurry2.0P binding affinity and the presentation score predictions of MHCflurry.
  • top predictions of each method contained a large number of experimentally-determined epitopes, a good number of epitopes were also ranked quite low by each method (FIG. 10) .
  • Exploring the relationships among the set of top 20 ranked peptides per HLA allele predicted by these methods revealed that the predictions of MHCflurry2.0P were most distinct from those of other methods (FIGS. 11A-D) . Consistent results were obtained when this set was constructed by pooling the top 10 to top 25 ranked peptides restricted by each HLA allele (FIGS. 11A-D) .
  • Predictions of MHCflurry2.0P also contained a large number of experimentally-determined SARS-CoV-2 epitopes that were not present in the set of top-ranked peptides predicted by any other method (FIG. 7A) .
  • a strategy that combines the predictions of MHCflurry2.0P with any of the other 11 methods would work better than any individual method.
  • the union strategy combines the top x predictions of any two methods and provides a set of peptides whose size can vary between x and 2x depending on the number of common peptides predicted by each method.
  • SARS2TPools The union method implemented in SARS2TPools combines the predictions of MHCflurry2.0P with NetMHCpan4.1BA.
  • SARS2TPools provides optimized peptide pools by supplementing experimentally-determined epitopes with small or large sized group of in silico predicted epitopes corresponding respectively to top 10 and top 20 predictions of each method being combined.
  • the platform also provides a relaxed threshold which can be particularly useful for specific proteins with very limited number of predicted epitopes.
  • SARS2TPools provides optimized peptide pools for assessing T cell responses leveraging experimentally-determined SARS-CoV-2 epitope data.
  • the designed CD8+ T cell peptide pools also include in silico predictions obtained using a computational approach optimized for predicting SARS-CoV-2 CD8 + T cell epitopes. This approach is based on benchmarking predictions of state- of-the-art in silico methods against the ample experimentally-determined SARS-CoV-2 CD8+epitope data that is now available (Methods) .
  • HLA class I diversity and summary of experimentally-determined SARS-CoV-2 CD8+ T cell epitope data Each individual possesses three major HLA class I alleles, HLA-A, HLA-B, and HLA-C, which are among the most polymorphic loci of the human genome (Jin et al. 2018) . In fact, more than 6, 500 alleles for each of these three loci have been identified so far (Robinson et al. 2015) . In order to design a peptide pool for assessing T cell response in a specific geographical region, information of the set of most common HLA alleles in that population is required.
  • the allele frequency net database (AFND) has curated a list of ten most common HLA alleles per locus for 11 distinct geographical regions encompassing the global population (Gonzalez-Galarza et al. 2019) . These sets of HLA alleles have a population coverage of 96%or more in the respective regions. By combining alleles prevalent in all regions, we obtained a total of 136 distinct HLA class I alleles (FIG. 6A, Table 12) . Ideally, a peptide pool for assessing T cell responses in a specific region should include a set of experimentally-determined epitopes against all HLA alleles prevalent in that region.
  • Table 12 List of 136 HLA alleles that are ranked in the top 10 most-prevalent HLA alleles in at least one of the 13 geographical regions defined by AFND.
  • the platform SARS2TPools integrates experimentally-determined SARS-CoV-2 epitope data, in silico predictions, and information of prevalent HLA alleles across regions to provide optimized peptide pools for assessing vaccine-induced SARS-CoV-2 T cell responses (FIG. 8A) . It enables users to obtain region-specific, host-specific, and protein-specific optimized peptide pools through a simple and user-friendly interface (FIG. 8B) . Peptide pools optimized for each of the 11 regions, by taking into account the information of HLA alleles prevalent in these regions (FIG.
  • FIG. 6A can be obtained by selecting the ‘Region-specific’ tab on SARS2TPools interface (FIG. 8B) .
  • This optimization is important due to the heterogeneity of prevalent HLA alleles among regions that may result in presentation of different epitopes and consequently different T cell responses (FIG. 6A) .
  • Region-specific pools can be useful to contrast T cell responses induced by the same vaccine in different geographical regions, and to understand the role of population heterogeneity in mediating different disease outcomes.
  • Host-specific (or cohort-specific) peptide pools for a range of peptide lengths (8–11 residues) , optimized for HLA haplotypes of the host (or cohort) can be obtained by selecting the ‘Host-specific’ tab on SARS2TPools. This optimization can be useful for cases when host (or cohort) HLA typing has been performed.
  • SARS2TPools also provides the flexibility to select protein-specific pools derived from the entire SARS-CoV-2 proteome or any number of specific proteins for both region-specific and host-specific options. Moreover, the user can also select peptides belonging to a specific domain of a protein (e.g., receptor binding domain of the spike protein) by specifying a range of amino- acid positions. Protein-specific peptide pools are important for assessing T cell responses induced by vaccine comprising of different antigens, e.g., S only, S with other proteins, or whole-virion.
  • the CD8 + T cell peptide pools provided by SARS2TPools comprise of both experimentally-determined and in silico predicted epitopes, while those for measuring CD4 + T cell responses comprise only of experimentally-determined epitopes.
  • the platform indicates whether a peptide is an experimentally-determined epitope, predicted epitope, or both. It also provides users the flexibility to select peptide pools consisting exclusively of experimentally-determined CD8 + T cell epitopes.
  • SARS2TPools ranks peptides within an optimized pool based on HLA promiscuity, which can be used as an additional prioritization criterion among peptides. In summary, SARS2TPools allows users to obtain pools for assessing T cell responses by flexibly selecting various optimization criteria.
  • Region-specific CD8+ T cell pools from the whole SARS-CoV-2 proteome As an illustrative example, we used SARS2TPools to obtain region-specific optimized pools, comprising of peptides derived from the whole SARS-CoV-2 proteome, for measuring CD8 + T cell responses for each of the 11 regions defined by AFND. These pools were designed to include around 20 peptides corresponding to each HLA allele prevalent in a region as follows: (i) If an HLA allele prevalent in a region had 20 or more associated experimentally-determined SARS-CoV-2 epitopes, we selected 20 of these based on response frequency (proportion of responding donors (Quadeer et al.
  • each optimized pool can be grouped into 3 classes based on whether their associated peptides in the pool are all (i) experimentally-determined, (ii) in silico predicted, or (iii) a mix of both (FIG. 9B) .
  • in silico predictions help to fill the gap for alleles for which limited or no experimentally-determined epitopes are available at present. Comparing the 11 region-specific optimized pools, we found that these pools are largely distinct from each other (FIG. 9C) .
  • SARS2TPools provides optimized peptide pools for assessing CD8 + T cell responses against any of the 30 individual proteins of SARS-CoV-2 for each of the 11 regions.
  • a total of 31 x 11 341 region-specific optimized peptide pools are provided by the platform.
  • Table 15 List of all peptide sequences (with their IDs) for CD8 T cell pools.
  • Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide 1 STECSNLLLQY 51 AIVSTIQRK 101 VYSSANNCTF 2 FADDLNQLTGY 52 YMSALNHTK 102 VYRGTTTYKL 3 VTVKNGSIHLY 53 RIAGHHLGR 103 AYILFTRFF 4 SSANNCTFEY 54 QLRARSVSPK 104 EYHDVRVVL 5 FSAVGNICY 55 SVYAWNRKR 105 IFFITGNTL 6 VVDYGARFY 56 TLADAGFIK 106 DVFYKENSY 7 YKIEELFYSY 57 GVYYHKNNK 107 NTVKSVGKF 8 SSEAFLIGCNY 58 AIVSTIQRKYK 108 DTFCAGSTF 9 LADAGFIKQY 59 ALDPLSETK 109 FTISVTTEI 10 TDEMIAQY 60 RASANLA
  • Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide 1661 RLDKVEAEV 1711 AQVAKSHNI 1761 SIIIGGAKL 1662 TTAYANSVF 1712 KVTLVFLFV 1762 FFFLYENAF 1663 TGDSCNNYM 1713 VRIKIVQML 1763 MADLVYALR 1664 QESPFVMMS 1714 ARVECFDKF 1764 YINVFAFPF 1665 GYKSVNITF 1715 SYIVDSVTV 1765 SEAKCWTET 1666 EYVSQPFLM 1716 YQVNGYPNM 1766 CTSCCFSER 1667 KGLPWNVVR 1717 TVLCLTPVY 1767 YVPLKSATC 1668 KEKVNINIV 1718 KWGKARLYY 1768 QQGEVPVSI 1669 AAFHQECSL 1719 LEGNFYGPF 1769 VHTANKWDL 1670 LAYILFTRF 1720
  • Table 21 Region-specific peptide pools derived from NSP16 protein for CD8 T cell assays.
  • Table 23 Region-specific peptide pools derived from M protein for CD8 T cell assays.
  • Table 25 Region-specific peptide pools derived from NSP9 protein for CD8 T cell assays.
  • Table 28 Region-specific peptide pools derived from E protein for CD8 T cell assays.
  • Table 30 Region-specific peptide pools derived from ORF3a protein for CD8 T cell assays.
  • Table 33 Region-specific peptide pools derived from NSP2 protein for CD8 T cell assays.
  • Table 40 Region-specific peptide pools derived from ORF3d protein for CD8 T cell assays.
  • Table 41 Region-specific peptide pools derived from ORF10 protein for CD8 T cell assays.
  • Table 45 Region-specific peptide pools derived from NSP7 protein for CD8 T cell assays.
  • SARS2TPools is a first-of-its-kind software platform that provides optimized SARS-CoV-2 peptide pools for assessing vaccine-induced T cell responses in a geographical region. These pools can be used to explore key questions related to COVID-19 vaccines and the role of T cell immunity. For example, characterizing effects of mutations within T cell epitopes in emerging SARS-CoV-2 variants on vaccine-induced T cell responses, investigating association between targeting specific T cell epitopes and protection from severe disease, characterizing the breadth, diversity, and durability of vaccine-induced T cell responses, and contrasting the T cell responses elicited by different vaccines. These questions are important and relevant for both academic and industry research focused on development and assessment of COVID-19 vaccines. Answering them can help provide pre-emptive indicators of the potential for T-cell escape (e.g., due to variants) for specific vaccines.
  • HLA alleles commonly found in various regions are underrepresented in the available experimental data (FIG. 6B) .
  • T cell responses against SARS-CoV-2 proteins other than S are also understudied (FIG. 6C) . Characterizing these responses is particularly important for inactivated whole-virion based vaccines (already in use in more than 48 countries (Shrotri et al. 2021) ) that have been shown to elicit T cell responses against not only S but also against other SARS-CoV-2 proteins (Bueno et al. 2021) . This characterization will also be important for emerging SARS-CoV-2 vaccines (Hwang et al.
  • HLA-A*02: 01 and HLA-B*08: 01 HLA-A*02: 01 and HLA-B*08: 01
  • Table 14 List of epitopes having experimentally-determined association with at least 2 HLA alleles and in silico predicted association with at least 8 additional HLA alleles.
  • peptide pools optimized for a specific region comprise of a limited set of peptides in the context of cognate HLA alleles, and would thus be less susceptible to peptide competition (Pala et al. 1988) . This is of practical importance particularly for a virus like SARS-CoV-2 which has a large ( ⁇ 10k residues) proteome.
  • generalized peptide pools have also been proposed to assess SARS-CoV-2 T cell responses. Such pools comprise of peptides associated with a few globally prevalent HLA alleles (e.g., the pool proposed in (Grifoni et al.
  • peptides associated with 12 most-prevalent HLA-A and -B alleles comprises of peptides associated with 12 most-prevalent HLA-A and -B alleles.
  • a peptide pool optimized for the HLA alleles prevalent in a specific region would be expected to measure T cell responses in a population more comprehensively.
  • CD4 + T cell responses comprise only of experimentally-determined CD4 + T cell epitopes.
  • CD4 + pools provided by SARS2TPools may then be supplemented by in silico predictions following an approach similar to the one employed for CD8 + T cell epitopes. SARS2TPools will be periodically updated to incorporate more experimental CD4 + and CD8 + T cell epitope data as it becomes available.
  • HLA class I and class II restricted SARS-CoV-2 T cell epitope data CD8 + and CD4 + , respectively
  • IEDB immune epitope database
  • HLA class II restricted epitopes all the available epitopes were 15 resides long, and only 3 HLA alleles had more than 20 epitopes in the data.
  • NetMHC4.0 (Andreatta and Nielsen 2016)
  • other common prediction methods that have been employed for predicting SARS-CoV-2 epitopes (Sohail et al. 2021) such as NetMHCpan3.0 (Nielsen and Andreatta 2016) , SMM (Peters and Sette 2005) , SMMPMBEC (Kim et al. 2009) , and IEDB consensus (Moutaftsi et al. 2006) .
  • MHCflurry2.0BA binding affinity and the presentation score predictions of MHCflurry as two separate methods.
  • MHCflurry2.0P binding affinity and the presentation score predictions of MHCflurry.
  • top predictions of each method contained a large number of experimentally-determined epitopes, a good number of epitopes were also ranked quite low by each method (FIG. 10) .
  • Exploring the relationships among the set of top 20 ranked peptides per HLA allele predicted by these methods revealed that the predictions of MHCflurry2.0P were most distinct from those of other methods (FIGS. 11A-D) . Consistent results were obtained when this set was constructed by pooling the top 10 to top 25 ranked peptides restricted by each HLA allele (FIGS. 11A-D) .
  • Predictions of MHCflurry2.0P also contained a large number of experimentally-determined SARS-CoV-2 epitopes that were not present in the set of top-ranked peptides predicted by any other method (FIG 7A) .
  • FOG 7A the uniqueness of the predictions of MHCflurry2.0P
  • the union strategy combines the top x predictions of any two methods and provides a set of peptides whose size can vary between x and 2x depending on the number of common peptides predicted by each method.
  • R i Rank-sum metric R i is defined as where r i (x) is the i-th union method’s rank (assigned based on hit-rate) among the 11 union methods for the set of top x ranked predicted peptides (FIG. 13) .
  • Hit-rate represents the fraction of experimentally known epitopes present in the set of top x ranked predicted peptides.
  • SARS2TPools The union method implemented in SARS2TPools combines the predictions of MHCflurry2.0P with NetMHCpan4.1BA.
  • SARS2TPools provides optimized peptide pools by supplementing experimentally-determined epitopes with small or large sized group of in silico predicted epitopes corresponding respectively to top 10 and top 20 predictions of each method being combined.
  • the platform also provides a relaxed threshold which can be particularly useful for specific proteins with very limited number of predicted epitopes.
  • EXAMPLE 3 –COVIDEP A WEB-BASED PLATFORM FOR REAL-TIME RESPORTING OF VACCINE TARGET RECOMMENDATIONS FOR SARS-COV-2
  • COVIDep COVIDep. ust. hk
  • COVIDep a first-of-its-kind web-based platform that pools genetic data for SARS-CoV-2 and immunological data for the 2003 SARS virus, SARS-CoV, to identify B-cell and T-cell epitopes to serve as vaccine target recommendations for SARS-CoV-2 (FIG. 14A) .
  • the identified epitopes are experimentally-derived from SARS-CoV and have a close genetic match with the available SARS-CoV-2 sequences (see FIG. 15 for a detailed protocol description) .
  • COVIDep periodically pools SARS-CoV-2 sequence data from the GISAID database (gisaid.
  • T cell epitopes were determined based on either positive T cell assays or positive MHC binding assays for SARS-CoV.
  • B cell epitopes both linear and discontinuous epitopes were considered.
  • the system outputs those epitopes that are genetically similar in SARS-CoV-2, based on an epitope screening parameter. This user-defined parameter allows the user to select epitopes based on their conservation in the SARS-CoV-2 sequence data, where conservation is defined as the fraction of SARS-CoV-2 sequences with the exact epitope sequence.
  • This parameter is set to 0.95 as default; however, the user may change this value to adjust the stringency of the screening criterion. For example, reducing the value of the parameter will allow for the consideration of epitopes with greater genetic variation, potentially increasing the set of recommended SARS-CoV-2 vaccine targets.
  • the population coverage analysis tool available at IEDB www. iedb. org
  • IEDB www. iedb. org
  • COVIDep is flexible and user-friendly, comprising an intuitive graphical interface and interactive visualizations.
  • COVIDep includes displays for each of the SARS-CoV-2 proteins, showing the locations of the identified epitopes on the primary structure. Further graphical displays are provided to aid interpretation of the data, including a temporal and geographical breakdown of the analysed sequences, and a display of the observed genetic variation (amino acid mutation frequencies) for each of the SARS-CoV-2 proteins.
  • the platform is updated daily, based on the latest SARS-CoV-2 sequence data in the GISAID database (gisaid. org) . Periodic updates are important since SARS-CoV-2 sequences are being made available at an increasing rate through international data sharing efforts, and the identification of vaccine targets is influenced by newly observed genetic variation.
  • the vaccine targets recommended by COVIDep exploit the genetic similarities between SARS-CoV-2 and SARS-CoV, along with known immune targets for SARS-CoV that have been determined experimentally (available at the ViPR database; viprbrc. org) .
  • the system implements a protocol that identifies, from among the SARS epitopes that can induce a human immune response, those that are genetically similar in SARS-CoV-2.
  • This approach proposed and tested in Example 1 [see also 56] based on limited early data, identified known SARS-CoV epitopes that had an identical genetic match in SARS-CoV-2. These epitopes presented initial vaccine target recommendations for potentially eliciting a protective, cross-reactive immune response against SARS-CoV-2. Similar results were reported in a subsequent independent study [57] , where a related approach exploiting genetic similarity between SARS-CoV and SARS-CoV-2 was used to identify potential SARS-CoV-2 vaccine targets.
  • FIG. 14B illustrates the T-cell epitopes reported by COVIDep (as of 20 May 2020) for the spike protein of SARS-CoV-2.
  • the Search box in the top right was used to select only the HLA-A*02: 01-restricted epitopes.
  • COVIDep may be used to broadly guide vaccine designs and associated experimental studies, and may help to expedite the discovery of an effective vaccine for COVID-19.
  • the SARS-CoV-2 full genome sequence data was periodically downloaded from the Global Initiative on Sharing Avian Influenza Database (GISAID; www. gisaid. org) .
  • the SARS-CoV epitope sequence data was downloaded from the Virus Pathogen Database and Analysis Resource (ViPR; viprbrc. org) .
  • the population coverage statistics of HLA alleles were obtained from the Immune Epitope Database and Analysis Resource (IEDB; iedb. org) .
  • the source code for the developed platform is available at the COVIDep GitHub repository github. com/COVIDep) .
  • the novel coronavirus 2019 uses the SARS-coronavirus receptor ACE2 and the cellular protease TMPRSS2 for entry into target cells. bioRxiv 2020.01.31.929042 2020.
  • MHCflurry 2.0 Improved pan-allele prediction of MHC class I-presented peptides by incorporating antigen processing.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Virology (AREA)
  • Organic Chemistry (AREA)
  • Molecular Biology (AREA)
  • Immunology (AREA)
  • General Health & Medical Sciences (AREA)
  • Medicinal Chemistry (AREA)
  • Biochemistry (AREA)
  • Cell Biology (AREA)
  • Engineering & Computer Science (AREA)
  • Hematology (AREA)
  • Genetics & Genomics (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Urology & Nephrology (AREA)
  • Microbiology (AREA)
  • Communicable Diseases (AREA)
  • Veterinary Medicine (AREA)
  • Public Health (AREA)
  • Animal Behavior & Ethology (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Analytical Chemistry (AREA)
  • Physics & Mathematics (AREA)
  • Pathology (AREA)
  • Gastroenterology & Hepatology (AREA)
  • Biotechnology (AREA)
  • General Physics & Mathematics (AREA)
  • Tropical Medicine & Parasitology (AREA)
  • Zoology (AREA)
  • Food Science & Technology (AREA)
  • Oncology (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • General Chemical & Material Sciences (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • Pulmonology (AREA)
  • Mycology (AREA)
  • Epidemiology (AREA)

Abstract

The present disclosure identifies a set of T cell and B cell epitopes derived from Severe Acute Respiratory Syndrome Coronavirus 1 (SARS-CoV-1) and their use in designing vaccines against SARS-CoV-2. A method for eliciting an immune response for prophylactic or therapeutic application using the T cell and/or B cell epitopes is described. Further disclosed are SARS-CoV-2-specific polypeptides, nucleic acids, host cells, and corresponding compositions for eliciting or measuring an immune response to SARS-CoV-2 and/or other coronaviruses.

Description

IDENTIFICATION AND USES OF PEPTIDE SEQUENCES OF SARS-COV-2 T CELL AND B CELL EPITOPES
CROSS-REFERENCES TO RELATED APPLICATIONS
This application claims priority to U.S. Provisional Patent Application No. 63/229,063, filed August 3, 2021, and U.S. Provisional Patent Application No. 63/136,175, filed January 11, 2021, the contents of both are hereby incorporated by reference in the entirety for all purposes.
BACKGROUND
COVID-19 is caused by a novel coronavirus, Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) . Worldwide collaborative efforts from scientists are working on this disease and SARS-CoV-2 to develop effective interventions for controlling and preventing it [6–9] .
Currently, treatment options for COVID-19, which is caused by a novel coronavirus, Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) , are limited. The Food and Drug Authorization (FDA) has approved the antiviral drug Veklury (remdesivir) for adults and certain pediatric patients with COVID-19 who are sick enough to need hospitalization. By Emergency Use Authorization (EUA) , the FDA may authorize the use of unapproved drugs or unapproved uses of approved drugs under certain conditions. For example, the FDA has issued EUAs for several monoclonal antibody treatments for COVID-19 for the treatment, and in some cases prevention (prophylaxis) , of COVID-19 in adults and pediatric patients. Monoclonal antibodies are laboratory-made molecules that act as substitute antibodies. They can help the immune system recognize and respond more effectively to the virus, making it more difficult for the virus to reproduce and cause harm. Like other infectious organisms, SARS-CoV-2 can mutate over time, resulting in genetic variation in the population of circulating viral strains. The risk-benefit assessment for using certain monoclonal antibodies may not be favorable due to the increased frequency of resistant variants. Global efforts to combat COVID-19 have led to the rapid development of multiple vaccines. As the virus continues to circulate worldwide, virus variants have emerged in several regions, raising concerns about their potential to escape vaccine-induced antibody responses.
As such, there is an imminent need to improve efficacy of COVID-19 treatments and provide protection against severe disease and hospitalization. While vaccines have been pursued with success, there is an unmet need for developing vaccines that may induce a robust neutralizing antibody response to combat the COVID-19 pandemic.
SUMMARY
The summary is provide to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter.
This present disclosure provides new compositions and methods useful for eliciting an immune response to SARS-CoV-2 and/or other coronavirus infections, based on the discovery that SARS-CoV-2 and/or other coronavirus infections may elicit T-cell and/or B cell response in a subject. Thus, in one aspect, the present disclosure is related to a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence.
In some embodiments, the peptide comprises at least one T cell epitope set forth in Table 3 and at least one T cell epitope set forth in Tables 15-47. In some embodiments, the peptide comprises at least one T cell epitope set forth in Table 3 and at least one B cell epitope set forth in Table 4. In some embodiments, the peptide comprises at least one T cell epitope set forth in Tables 15-47 and at least one B cell epitope set forth in Table 4. In some embodiments, the peptide comprises at least one T cell epitope set forth in Table 3 and at least one T cell epitope set forth in Tables 15-47 and at least one B cell epitope set forth in Table 4. In any cases, the peptide may optionally further comprise at least one heterologous amino acid sequence. In some embodiments, the peptide comprises or consists of at least one of the T cell epitopes. In some embodiments, the peptide comprises or consists of at least one of the B cell epitopes. In some embodiments, the peptide comprises or consists of at least one of the T cell epitopes and at least one heterologous amino acid sequence. In some embodiments, the peptide comprises or consists of at least one of the B cell epitopes and at least one heterologous amino acid sequence.
In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids. In any cases, the peptide may further comprise at least one heterologous amino acid sequence.
The second aspect of the present disclosure provides a nucleic acid comprising a polynucleotide sequence encoding the peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence. In some embodiments, the polynucleotides encodes a peptide of at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the polynucleotides encodes a peptide of between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the polynucleotides encodes a peptide of about 20 amino acids. In some embodiments, the polynucleotides encodes a peptide of about 100 amino acids. In some embodiments, the polynucleotide sequence comprises deoxyribonucleic acid (DNA) . In some embodiments, the polyneucleotide sequence comprises ribonucleic acid (RNA) such as messenger  RNA (mRNA) . In some embodiments, the present disclosure provides an expression cassette comprises any of the peptide as described herein and the peptide is operably linked to a promoter. In some embodiments, the present disclosure provides a vector comprising the expression cassette comprising any of the peptides as described herein and the peptide is operably linked to a promoter. In some embodiments, the present disclosure provides a host cell comprising the vector comprising the expression cassette.
In a third aspect, the present disclosure provides a composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any of the peptides, the vector comprising the expression cassette, or the host cell; and (2) a pharmaceutically acceptable excipient. In some embodiments, the composition further comprises an adjuvant. In some embodiments, the composition comprises a plurality of peptides each comprising a T cell epitope set forth in Table 3 or Tables 15-47. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide  comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
In a fourth aspect, the present disclosure provides a method of eliciting an immune response in a subject in need thereof, the method comprising administering to the subject an effective amount of a composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any of the peptides, or the vector comprising the expression cassette. In some embodiments, the method comprises administering the composition to the subject by a route selected from the group consisting of, subcutaneous, intramuscular, and oral. In some embodiments, the subject is at risk of exposure to SARS-CoV, SARS-CoV-2 or other coronavirus infections. In some embodiments, the subject is at risk of exposure to SARS-CoV-2 infections. In some embodiments, the subject is at risk of developing severe illnesses when contracted with SARS-CoV-2 (e.g., COVID-19) . In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments,  the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
In a fifth aspect, the present disclosure provides a kit for eliciting an immune response in a subject in need thereof, the kit comprising a first container containing the composition comprising (1) a peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence, or a peptide comprising or consisting of at least one of the T cell epitopes, or a peptide comprising or consisting of at least one of the B cell epitopes, or a peptide comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence, the nucleic acid encoding any of the peptides, the expression cassette comprising any of the peptides, or the vector comprising the expression cassette. In some embodiments, the kit optionally contains an additional container containing a therapeutic agent against SARS-CoV-2. In some embodiments, the kit further comprises at least a second container each containing at least one different composition. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises of no more than 500, 450, 400, 350, 300, 250, 200, 150, 120, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises between about 5-500, about 20-100, about 40-200, about 50-450, about 80-350, about 100-400, or about 150-350 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 20 amino acids. In some embodiments, the peptide comprising the one or more T cell epitope set forth in Table 3 and/or Tables 15-47, and/or one or more B cell epitopes set forth in Table 4 comprises about 100 amino acids.
In a sixth aspect, the present disclosure provides a method for detecting T cell immunity against SARS-CoV-2 in a subject, comprising: (1) contacting T cells obtained from the subject  with a T cell epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having an HLA allele associated with the epitope; and (2) detecting activation of the T cells, thereby detecting presence of T cell immunity against SARS-CoV-2 in the subject. In some embodiments, step (2) of the method comprises detection of T cell proliferation or T cell secretion of one or more cytokines. In some embodiments, step (2) of the method comprises T cell proliferation assay, flow cytometry, ELISPOT, or ELISA. In some embodiments, step (1) of the method comprises contacting T cells obtained from the subject with a composition comprising a plurality of peptides each comprising a T cell epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having HLA alleles associated with each of the plurality of epitopes.
BRIEF DESCRIPTION OF THE DRAWINGS
FIGS. 1A-B show comparison of the similarity of structural proteins of SARS-CoV-2 with the corresponding proteins of SARS-CoV and MERS (Middle East Respiratory Syndrome) -CoV. FIG. 1A shows the percentage genetic similarity of the individual structural proteins of SARS-CoV-2 with those of SARS-CoV and MERS-CoV. The reference sequence of each coronavirus (Materials and Methods) was used to calculate the percentage genetic similarity. FIG. 1B is a circular phylogram of the phylogenetic trees of the four structural proteins. All trees were constructed based on the available unique sequences using PASTA and rooted with the outgroup Zaria Bat CoV strain (accession ID: HQ166910.1) .
FIGS. 2A-C depict location of SARS-CoV S protein subunits and SARS-CoV-derived B cell epitopes on the protein structure (PDB ID: 5XLR) . FIG. 2A shows subunits S1 and S2 are indicated in medium and light grey color, respectively. The receptor binding motif lies within the S1 subunit and is indicated in dark grey color. FIG. 2B shows residues of the linear B cell epitopes, that were identical in SARS-CoV-2 (Table 4) , are shown as shaded. The black and grey colors reflect the surface and buried residues, respectively. FIG. 2C shows locations of discontinuous B cell epitopes that share at least one identical residue with corresponding SARS-CoV-2 sites (Table 5) . Identical epitope residues are shown in dark grey color, while the remaining epitope residues are shown in medium grey color. Both the side view (left panel) and the top view (right panel) of the structure are shown.
FIG. 3 shows fraction of mutations in the observed sequences of the structural proteins of the three coronaviruses. Mutation is defined here as an amino acid difference from the reference  sequence of the respective coronavirus; accession IDs: NC_045512.2 (2019-nCoV) , NC_004718.3 (SARS-CoV) , and NC_019843.3 (MERS-CoV) .
FIG. 4 illustrates location of identified T cell epitopes on the SARS-CoV S protein structure (PDB ID: 5XLR) . Residues of the SARS-CoV-derived T cell epitopes (determined using positive MHC binding assays and that were identical in SARS-CoV-2) are shown with filled color. The dark and light shade reflect the surface and buried residues, respectively.
FIG. 5 shows pairwise sequence alignment of the reference sequences of the S proteins of SARS-CoV and SARS-CoV-2 (accession ID: NP_828851.1 and YP_009724390.1, respectively) . Identical residues are indicated by *.
FIGS. 6A-C show global diversity of HLA class I alleles and the corresponding T cell responses, and summary of the experimentally-determined SARS-CoV-2 CD8+ T cell epitope data. FIG. 6A shows different HLA class I alleles are prevalent in different geographical regions (left panel) ; heatmaps depicting the diverse HLA allele distribution in North America, North East Asia, and Australia are shown as examples (middle panel) . Each square in the heatmap represents a distinct HLA class I allele, and its color shade represents the frequency of the allele in the geographical region with a dark shade representing high frequency and vice versa; and different HLA alleles present different SARS-CoV-2 peptides, and consequently the T cell responses elicited against SARS-CoV-2 are expected to differ among geographical regions (right panel) . Thus, peptide pools to test for SARS-CoV-2 CD8+ T cell responses need to be designed specific to the HLA alleles prevalent in the population being tested. This figure was created with BioRender. com. FIG. 6B shows the number of HLA alleles in each geographical region for which at least one SARS-CoV-2 experimentally-determined CD8+ T cell epitope is known. FIG. 6C shows the number of experimentally-determined epitopes within spike and other SARS-CoV-2 proteins determined so far for class I HLA alleles (left panel) and the frequencies of these HLA alleles in different geographical regions (right panel) (for details of experimentally-determined epitope data, see Methods) .
FIGS . 7A-B show an optimized in silico strategy to predict SARS-CoV-2 CD8 + T cell epitopes. FIG. 7A shows a performance comparison of in silico HLA class I epitope prediction methods based on predicting experimentally-determined SARS-CoV-2 CD8 + T cell epitopes. Intersection of the top predictions of the analysed 12 in silico methods showed that MHCflurry2.0P  predicts the highest number of epitopes not predicted by any other method. For each method, the set of top-ranked predicted peptides consisted of n = 200 predictions, with 20 predictions for each of the 10 HLA alleles having the most experimental data available (see Methods) . FIG. 7B shows the set of top-ranked predictions (n = 341) obtained from the proposed approach, union of MHCflurry2.0P and NetMHCpan4.1BA (see Methods) , comprised of more experimentally-determined epitopes than the sets of top-ranked predictions of the individual methods
FIGS. 8A-B show a schematic of SARS2TPools framework and snapshot of the platform’s interface for a use case. FIG. 8A is a schematic showing how the SARS2TPools, based on user-selected options (protein, geographical region, specific HLA alleles) , combines experimentally-determined epitope data, in silico predictions, and information of prevalent HLA alleles across regions to obtain optimized peptide pools for assessing vaccine-induced T cell responses. This figure was created with BioRender. com. FIG. 8B is a snapshot of the SARS2TPools interface for the “Region-specific” tab with selected options (Region: Oceania; Protein: NSP12; Pool-size of peptides: small; Threshold for in silico predictions: Default; and Length of peptides: 9) .
FIGS. 9A-C are summary of region-specific CD8+ T cell peptide pools provided by SARS2TPools. FIG. 9A shows the number of peptides in region specific pools from experimental studies and in silico predictions. FIG. 9B shows overlap between pairs of region-specific pools. FIG. 9C shows a source of peptides for each region-specific pool. Each circle represents one of the top 30 prevalent HLAs for a specific region and the color of the circle indicates whether the peptides were obtained from only experimental studies, or only from in silico predictions, or from both.
FIG. 10 shows histograms of ranks assigned by different in silico epitope prediction methods to the experimentally-determined SARS-CoV-2 CD8 + T cell epitopes associated with the 10 HLA alleles with the most experimental data (see Methods) . Each bin here represents a set of peptides within the specified rank range of all 10 HLA alleles.
FIGS. 11A-D show the observation that predictions of MHCflurry2.0P are distinct from those of other in silico methods (FIG. 12) was robust to the size of the set of top-ranked predictions compared. Distinct sets of predicted epitopes ranked by the considered 12 in silico methods in their  (FIG. 11A) top 10, (FIG. 11B) top 15, (FIG. 11C) top 20, and (FIG. 11D) top 25 predictions corresponding to each of the 10 HLA alleles having the most experimental data (see Methods) .
FIG. 12 shows the strategy of combining top-ranked predictions of MHCflurry2.0P with any of the other 11 in silico methods considered in this study had a higher hit-rate on experimentally-determined SARS-CoV-2 CD8 + epitope data than either of the individual methods. Here, hit-rate represents the fraction of experimentally known epitopes present in the set of top-ranked predicted peptides.
FIG. 13 shows hit-rates of union methods comprising of MHCflurry2.0P and one of the other 11 in silico epitope prediction methods. While none of these clearly outperformed the rest, the 4 union methods comprising of MHCflurry2.0P and either NetMHCpan4.1BA, NetMHCpan4.0BA, NetMHC4.0, or NetMHCpan4.1EL were among the top 4 performing methods (also see Table 13) . Here, hit-rate represents the fraction of experimentally known epitopes present in the set of top-ranked predicted peptides.
FIGS. 14A-B illustrate COVIDep providing an up-to-date set of B-cell and T-cell epitopes that can serve as potential vaccine targets for SARS-CoV-2. FIG. 14A shows the identified epitopes are experimentally-derived from SARS-CoV and have a close genetic match with the available SARS-CoV-2 sequences. FIG. 14B is an example of the T-cell epitopes reported by COVIDep (as of 20 May 2020) for the spike protein of SARS-CoV-2. Here, the Search box (in the top right) was used to select only the HLA-A*02: 01-restricted epitopes. (An explanation of all interactive COVIDep visualizations is incorporated in the “How to use COVIDep page” of the platform. ) Of the 14 epitopes listed in the display, 9 of them ( IEDB IDs  36724, 54507, 54725, 69657, 71663, 2801, 54680, 16156, and 37289) overlap with epitopes against which cytotoxic CD8+ T cell responses have been observed in peripheral blood mononuclear cells isolated from COVID-19 patients. T cell responses were also recorded against protein regions overlapping with the epitope with IEDB ID 71663 in a pre-clinical trial of a DNA vaccine candidate.
FIG. 15 is an epitope screening protocol used by COVIDep for providing vaccine target recommendations for SARS-CoV-2. COVIDep periodically pools SARS-CoV-2 sequence data from the GISAID database (www. gisaid. org) and compares with experimentally-determined T cell and B cell epitopes of SARS-CoV, obtained from the ViPR database (www. viprbrc. org) . The T cell epitopes were determined based on either positive T cell assays or positive MHC binding assays  for SARS-CoV. For the B cell epitopes, both linear and discontinuous epitopes were considered. The system outputs those epitopes that are genetically similar in SARS-CoV-2, based on an epitope screening parameter. This user-defined parameter allows the user to select epitopes based on their conservation in the SARS-CoV-2 sequence data, where conservation is defined as the fraction of SARS-CoV-2 sequences with the exact epitope sequence. The value of this parameter is set to 0.95 as default; however, the user may change this value to adjust the stringency of the screening criterion. For example, reducing the value of the parameter will allow for the consideration of epitopes with greater genetic variation, potentially increasing the set of recommended SARS-CoV-2 vaccine targets. For the identified T cell epitopes, the population coverage analysis tool available at IEDB (www. iedb. org) is used to estimate the percentage of a specified population that can elicit a response against them.
DEFINITIONS
The term "isolated, " when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It is preferably in a homogeneous state although it can be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified. In particular, an isolated gene is separated from open reading frames that flank the gene and encode a protein other than the gene of interest. The term "purified" denotes that a nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. Particularly, it means that the nucleic acid or protein is at least 60%, 70%, 80%, 85%, 90%, 95%, or 99%pure, more preferably at least 95%pure, and most preferably at least 99%pure.
In this application, the term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Unnatural (non-naturally occurring) amino acids include, without limitation, amino acid analogs, amino acid mimetics, synthetic amino acids, N-substituted glycines, and N-methyl amino acids in either the L-or D-configuration that function in  a manner similar to the naturally-occurring amino acids. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an α carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure different from the general chemical structure of an amino acid, but capable of functioning in a manner similar to a naturally occurring amino acid.
The term “nucleic acid” or “polynucleotide” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single-or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19: 5081 (1991) ; Ohtsuka et al., J. Biol. Chem. 260: 2605-2608 (1985) ; and Rossolini et al., Mol. Cell. Probes 8: 91-98 (1994) ) . The term nucleic acid is used interchangeably with gene, cDNA, or mRNA encoded by a gene.
When the relative locations of elements in a polynucleotide sequence are concerned, a "downstream" location is one at the 3' side of a reference point, and an "upstream" location is one at the 5' side of a reference point.
The terms “polypeptide, ” “peptide, ” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers. As used herein, the terms encompass amino acid chains of any length, including full-length proteins (i.e., antigens) , wherein the amino acid residues are linked by covalent peptide bonds. In this application, the amino acid sequence of a polypeptide is presented  from the N-terminus to the C-terminus. In other words, when describing an amino acid sequence of a peptide, the first amino acid from the N-terminus is referred to as the “first amino acid. ” 
When used in the context of describing partners of a fusion peptide, the term "heterologous" refers to the relationship of one peptide fusion partner to the another peptide fusion partner: the manner in which the fusion partners are present in the fusion peptide is not one that can be found a naturally occurring protein. For instance, a "heterologous polypeptide" fused with a T cell or B cell epitope to form a fusion peptide may be one that is originated from a protein other than the antigen from which the T cell or B cell epitope is derived, such as a granulocyte-macrophage colony-stimulating factor (GM-CSF) . On the other hand, a "heterologous polypeptide" may be one derived from another portion of the T cell or B cell protein that is not immediately contiguous to the T cell or B cell epitope. A "heterologous polypeptide" may contain modifications of a naturally occurring protein sequence or a portion thereof, such as deletions, additions, or substitutions of one or more amino acid residues. Regardless of the origin of the "heterologous polypeptide" (i.e., whether it is derived from the antigen from which the T cell or B cell epitope is derived or another protein) , the fusion peptide should not contain a subsequence of the human T cell or B cell that encompasses the amino acid sequences in Table 3 and Tables 15-47 and have more than 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 180, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more amino acids in length. In some exemplary embodiments, the fusion peptide should not contain a subsequence of the human T cell or B cell that encompasses the amino acid sequences in Table 3 and Tables 15-47 and have more than 500 amino acids in length. In some exemplary embodiments, a "heterologous polypeptide" for use in the present disclosure has no more than 15-20 amino acids in length; in other embodiments, a "heterologous polypeptide" has at least 100 amino acids in length. In some exemplary embodiments, a "heterologous polypeptide" is a recombinant polynucleotide (or a copy or complement of a recombinant polynucleotide) that has been manipulated using well known methods. For example, a "heterologous polypeptide" can comprise a recombinant expression cassette comprising a promoter operably linked to a second polynucleotide (e.g., a coding sequence of the antigen from which the T cell or B cell epitope is derived) . In this context, the promoter is heterologous to the second polynucleotide as the result of human manipulation (e.g., by methods described in Sambrook et al., Molecular Cloning -A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, New York, (1989) or Current Protocols in Molecular Biology  Volumes 1-3, John Wiley &Sons, Inc. (1994-1998) ) . As another example, the "heterologous polypeptide" may comprise a promoter that is heterologous to a second polynucleotide encoding the polypeptide of interest (e.g., a T cell and/or a B cell epitope) or a fragment thereof, and a third polynucleotide encoding a detecting (e.g., a tag or identifier) molecule. The detecting molecule can be a fluorescent protein, e.g., a green fluorescent protein (GFP) or a variant or the like, such as DsRed and other red fluorescent protein.
The word "fuse" or "fused, " as used in the context of describing a peptide of this disclosure that comprises a T cell or B cell epitope joined with a heterologous polypeptide, refers to a connection between the epitope and the heterologous polypeptide by any covalent bond, including a peptide bond.
The phrase "a nucleic acid sequence encoding" refers to a nucleic acid which contains sequence information for a structural RNA such as rRNA, a tRNA, or the primary amino acid sequence of a specific protein or peptide, or a binding site for a trans-acting regulatory agent. This phrase specifically encompasses degenerate codons (i.e., different codons which encode a single amino acid) of the native sequence or sequences that may be introduced to conform to codon preference in a specific host cell.
An “expression cassette” is a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular polynucleotide sequence in a host cell. An expression cassette may be part of a plasmid, viral genome, or nucleic acid fragment. Typically, an expression cassette includes a polynucleotide to be transcribed, operably linked to a promoter.
The term “recombinant, ” when used with reference, e.g., to a cell, or nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a nucleic acid or protein from an outside source or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified. Thus, for example, recombinant cells express genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under-expressed or not expressed at all.
The term "administration" or "administering" refers to various methods of contacting a substance with a mammal, especially a human. Modes of administration may include, but are not limited to, methods that involve contacting the substance intravenously, intraperitoneally, intranasally, transdermally, topically, subcutaneously, parentally, intramuscularly, orally, or systemically, and via injection, ingestion, inhalation, implantation, or adsorption by any other means. One exemplary means of administration of a T cell and/or B cell peptide of this disclosure or a fusion peptide comprising a T cell and/or B cell epitope (e.g., a T cell or B cell epitope derived from the S protein, N protein, or full length of SARS-CoV and/or SARS-CoV-2) and a heterologous polypeptide is via intramuscular delivery, where the peptide or fusion peptide can be formulated as a pharmaceutical composition in the form suitable for intramuscular injection, such as an aqueous solution, a suspension, or an emulsion, etc. Other means for delivering a T cell and/or B cell epitope or a fusion peptide of this disclosure includes intradermal injection, subcutaneous injection, intravenous injection, or transdermal application as with a patch.
An "effective amount" of a certain substance refers to an amount of the substance that is sufficient to effectuate a desired result. For instance, an effective amount of a composition comprising a peptide of this disclosure that is intended to induce an anti-SARS-CoV or SARS-CoV-2 immunity is an amount sufficient to achieve the goal of inducing the immunity when administered to a subject. The effect to be achieved may include the prevention, correction, or inhibition of progression of the symptoms of a disease/condition and related complications to any detectable extent. The exact quantity of an "effective amount" will depend on the purpose of the administration, and can be ascertainable by one skilled in the art using known techniques (see, e.g., Lieberman, Pharmaceutical Dosage Forms (vols. 1-3, 1992) ; Lloyd, The Art, Science and Technology of Pharmaceutical Compounding (1999) ; and Pickar, Dosage Calculations (1999) ) .
A "therapeutically effective amount" of a substance/molecule of the disclosure may vary depending on factors such as: the disease state, age, sex and weight of the individual, and the ability of the substance/molecule to induce a desired response in the individual. A therapeutically effective amount is also an amount that has a therapeutically beneficial effect over any toxic or detrimental effect of the substance/molecule.
A "prophylactically effective amount" refers to an amount of dosage and time necessary to effectively achieve the desired prophylactic result. Since prophylactic doses are used  in individuals prior to or early in the disease, the prophylactically effective amount is typically (but not necessarily) less than the therapeutically effective amount.
A "physiologically or pharmaceutically acceptable excipient" is an inert ingredient used in the formulation of a composition of this disclosure, which contains the active ingredient (s) of a T cell and/or B cell peptide or a fusion peptide comprising a T cell and/or B cell peptide and a heterologous polypeptide and is suitable for use, e.g., by injection into a patient in need thereof. This inert ingredient may be a substance that, when included in a composition of this disclosure, provides a desired pH, consistency, color, smell, or flavor of the composition.
As used herein, the term "T cell immune response" refers to activation of antigen specific T cells as measured by proliferation or expression of molecules on the cell surface or secretion of proteins such as cytokines.
As used herein, the term “B cell immune response” refers to either a T cell-independent immune response or a T cell-dependent immune response. In a T cell-independent response, B cells respond directly to the antigen. In a T cell-dependent immune response, B cells rely on the assistance from T cells to respond. Activated B cells may express IgA, IgE, IgG or retain IgM expression. Cytokines produced by T cells and others may determine what isotype the B cells express. Several general techniques are commonly used to identify antigen-specific B cells. Non-limiting examples are B cell enzyme linked immunospot (ELISPOT) , limiting dilution, flow cytometry, adoptive transfer, microscopy, and B cell receptor (BCR) transgenic mice.
As used herein, the term “immune response” generally refers to the immune system recognizing the antigens, e.g., proteins, on the surface of substances or microorganisms, such as bacteria or virus, and attacks and destroys, or tries to destroy, them. As described herein, the term “immune response” refers to a cell-mediated (T-cell) immune response and/or an antibody (B-cell) immune response.
As described herein, the terms “immunogenic composition, ” or “vaccine” are used interchangeably and refer to a composition that elicits an immune response in a subject, especially a human. An immunogenic composition or vaccine can be used prophylactically to prevent COVID-19 (SARS-CoV-2) or other coronavirus diseases (e.g., SARS-CoV) .
As described herein, the term “immunity” refers to protection from an infectious disease. For instance, if a subject is immune to a disease, the object can be exposed to it without becoming infected.
As described herein, the term “vaccine” generally refers to a preparation that is used to stimulate the body’s immune system to generate immunity for a disease. For instance, it refers to an antigen-containing formulation consisting of whole pathogenic organisms (killed or attenuated) or components of these organisms (such as proteins, peptides or polysaccharides) for conferring immunity against disease caused by these organisms. Vaccine formulations may be natural, synthetic or obtained by recombinant DNA techniques. A vaccine may be administered in any route known in the field.
As described herein, the terms “vaccination” or “immunization” or “inoculation” are used interchangeably. They refer to a process by which a person becomes protected against a disease by introducing a vaccine into the body to produce protection from the specific disease. For example, vaccines are usually administered intramuscularly through needle injections. As another example, vaccines can be administered by mouth (oral) or sprayed into the nose (intranasal) .
As described herein, the term "immunogenic" refers to the ability of an immunogen, antigen or vaccine to stimulate an immune response.
As used herein, the term “antigen" is defined as any substance capable of eliciting an immune response. For example, an “antigen” may be a small molecule or a macromolecule such as a protein, a peptide, a polysaccharide, a nucleic acid, a lipid, or a biomolecule.
As used herein, the term " antigen-specific" refers to a property of a population of cells such that the provision of a particular antigen or antigen fragment causes the proliferation of a specific cell.
The phrase “specifically (or selectively) binds” to an antibody or “specifically (or selectively) immunoreactive with, ” when referring to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein in a heterogeneous population of proteins and other biologics. Thus, under designated immunoassay conditions, the specified antibodies bind to a particular protein at least two times the background and do not substantially bind in a significant amount to other proteins present in the sample. Specific binding to an antibody  under such conditions may require an antibody that is selected for its specificity for a particular protein. For example, polyclonal antibodies raised to fusion proteins can be selected to obtain only those polyclonal antibodies that are specifically immunoreactive with fusion protein and not with individual components of the fusion proteins. This selection may be achieved by subtracting out antibodies that cross-react with the individual antigens. A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow &Lane, Antibodies, A Laboratory Manual (1988) , for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity) . Typically a specific or selective reaction will be at least twice background signal or noise and more typically more than 10 to 100 times background.
The terms "antibody" and "immunoglobulin" are used interchangeably in a broad sense and include monoclonal antibodies (e.g., full-length or intact monoclonal antibodies) , polyclonal antibodies, multivalent antibodies, multispecific antibodies (e.g., bispecific antibodies, so long as they exhibit the desired biological activity) , and may also include certain antibody fragments (as described in more detail herein) . The antibody may be a chimeric antibody, a human antibody, a humanized antibody, and/or an affinity matured antibody.
The terms “a, ” “an, ” and “the” as used herein not only include aspects with one member, but also include aspects with more than one member. For instance, the singular forms “a, ” “an, ” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells and reference to “the agent” includes reference to one or more agents known to those skilled in the art, and so forth.
As used herein, the term “about” denotes a range of ±10%of a specified value. For instance, “about 10” denotes a range of 9-11.
The term “subject” or "subject in need of treatment" refers to an individual who seeks medical attention due to risk of, or actual sufferance from, a condition involving undesirable inflammation (e.g., pneumonia or an infection that inflames air sacs) or a condition involving infection of the respiratory tract (e.g., the upper respiratory tract, the lungs) . Subjects or individuals in need of treatment include those that demonstrate symptoms of infection of the respiratory tract or those are at risk of later developing the disease or disorder and/or its symptoms.  For example, the subject may experience or is at risk of sufferance from a SARS-CoV infection. The subject may have one or more symptoms of SARS-CoV infection as defined by the Centers for Disease Control and Prevention (www. cdc. gov/coronavirus/2019-ncov/symptoms-testing/symptoms. html) . The subject may experience or have a wide range of symptoms ranging from mild to severe illness. Exemplary symptoms are fever or chills, cough, shortness of breath or difficulty breathing, fatigue, muscle or body aches, headache, new loss of taste or smell, sore throat, congestion or running nose, nausea or vomiting, or diarrhea, trouble breathing, persistent pain or pressure in the chest, new confusion, inability to wake or stay awake, and/or pale, gray, or blue-colored skin, lips, or nail beds, depending on skin tone, or combinations thereof. Symptoms may appear 0-60 days, 1-30 days, or 2-14 days after exposure to the source of infection (e.g., SARS-CoV or SARS-CoV-2 viruses) . The term subject can include both animals, especially mammals, and humans.
DETAILED DESCRIPTION
I. INTRODUCTION
Efforts have been made to develop vaccines against the novel SARS-CoV-2. Described herein are peptides, their related compositions, and methods of uses thereof. The disclosed peptides comprise at least one T cell epitope and/or at least one B cell epitope for eliciting an immune response in a subject against SARS-CoV-2. Epitopes that provide broad coverage broadly and region-specific are also explored.
We sought to gain insights for vaccine design against SARS-CoV-2 by considering the high genetic similarity between SARS-CoV-2 and SARS-CoV, which caused the outbreak in 2003, and leveraging existing immunological studies of SARS-CoV. By screening the experimentally-determined SARS-CoV-derived B cell and T cell epitopes in the immunogenic structural proteins of SARS-CoV, we identified a set of B cell and T cell epitopes derived from the spike (S) and nucleocapsid (N) proteins that map identically to SARS-CoV-2 proteins. As no mutation has been observed in these identified epitopes among the 120 available SARS-CoV-2 sequences (as of 21 February 2020) , immune targeting of these epitopes may potentially offer protection against this novel virus. For the T cell epitopes, we performed a population coverage analysis of the associated MHC alleles and proposed a set of epitopes that is estimated to provide broad coverage globally,  as well as in China. The findings provide a screened set of epitopes that can help guide experimental efforts towards the development of vaccines against SARS-CoV-2.
Unless specifically stated, SARS-CoV-2 is referred to the original SARS-CoV-2 (GenBank accession no. NC_045512) or any variants thereof. A variant has one or more mutations that differentiate it from other variants of the SARS-CoV-2 viruses. According to The Center for Disease Control and Prevention (CDC) , current known variants include, but are not limited to, Alpha (B. 1.1.7 and Q lineages) , Beta (B. 1.351 and descendent lineages) , Gamma (P. 1 and descendent lineages) , Epsilon (B. 1.427 and B. 1.429) , Eta (B. 1.525) , Iota (B. 1.526) , Kappa (B. 1.617.1) , 1.617.3, Mu (B. 1.621, B. 1.621.1) , Zeta (P. 2) , Delta (B. 1.617.2 and AY lineages) , and Omicron (B. 1.1.529 and BA lineages) . In the context of this disclosure, a polynucleotide of the full length SARS-CoV-2 (GenBank accession no. NC_045512) , a variant, or a fragment thereof may activates a T cell and/or B cell response. The polynucleotides may have at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more sequence identity with the full length SARS-CoV-2 (GenBank accession no. NC_045512) , a variant, or a fragment thereof.
II. CHEMICAL SYNTHESIS OF PEPTIDES
The peptides of the present disclosure, particularly those of relatively short length (e.g., no more than 50-100 amino acids, no more than 10-500 amino acids) , may be synthesized chemically using conventional peptide synthesis or other protocols well known in the art. In some embodiments, the peptides are of no more than about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids, or in the range of from about 10 to about 50, 100, 200, 300, 400 or 500 amino acids, from about 20 to about 50, 100, 200, 300, 400 or 500 amino acids, from about 50 to about 100, 150, 200, 300, 400 or 500 amino acids, from about 100 to about 150, 200, 250, 300, 400 or 500 amino acids, or from about 200 to about 250, 300, 350, 400 or 500 amino acids.
Peptides may be synthesized by solid-phase peptide synthesis methods using procedures similar to those described by Merrifield et al., J. Am. Chem. Soc., 85: 2149-2156 (1963) ; Barany and Merrifield, Solid-Phase Peptide Synthesis, in The Peptides: Analysis, Synthesis, Biology Gross and Meienhofer (eds. ) , Academic Press, N.Y., vol. 2, pp. 3-284 (1980) ; and Stewart et al., Solid Phase Peptide Synthesis 2nd ed., Pierce Chem. Co., Rockford, Ill. (1984) . During synthesis, N-α- protected amino acids having protected side chains are added stepwise to a growing polypeptide chain linked by its C-terminal and to a solid support, i.e., polystyrene beads. The peptides are synthesized by linking an amino group of an N-α-deprotected amino acid to an α-carboxy group of an N-α-protected amino acid that has been activated by reacting it with a reagent such as dicyclohexylcarbodiimide. The attachment of a free amino group to the activated carboxyl leads to peptide bond formation. The most commonly used N-α-protecting groups include Boc, which is acid labile, and Fmoc, which is base labile.
Materials suitable for use as the solid support are well known to those of skill in the art and include, but are not limited to, the following: halomethyl resins, such as chloromethyl resin or bromomethyl resin; hydroxymethyl resins; phenol resins, such as 4- (α- [2, 4-dimethoxyphenyl] -Fmoc-aminomethyl) phenoxy resin; tert-alkyloxycarbonyl-hydrazidated resins, and the like. Such resins are commercially available and their methods of preparation are known by those of ordinary skill in the art.
Briefly, the C-terminal N-α-protected amino acid is first attached to the solid support. The N-α-protecting group is then removed. The deprotected α-amino group is coupled to the activated α-carboxylate group of the next N-α-protected amino acid. The process is repeated until the desired peptide is synthesized. The resulting peptides are then cleaved from the insoluble polymer support and the amino acid side chains deprotected. Longer peptides can be derived by condensation of protected peptide fragments. Details of appropriate chemistries, resins, protecting groups, protected amino acids and reagents are well known in the art and so are not discussed in detail herein (See, e.g., Atherton et al., Solid Phase Peptide Synthesis: A Practical Approach, IRL Press (1989) , and Bodanszky, Peptide Chemistry, A Practical Textbook, 2nd Ed., Springer-Verlag (1993) ) .
III. RECOMBINANT PRODUCTION OF PEPTIDES
A.  General Recombinant Technology
Basic texts disclosing general methods and techniques in the field of recombinant genetics include Sambrook and Russell, Molecular Cloning, A Laboratory Manual (3rd ed. 2001) ; Kriegler, Gene Transfer and Expression: A Laboratory Manual (1990) ; and Ausubel et al., eds., Current Protocols in Molecular Biology (1994) .
For nucleic acids, sizes are given in either kilobases (kb) or base pairs (bp) . These are estimates derived from agarose or acrylamide gel electrophoresis, from sequenced nucleic acids, or from published DNA sequences. For proteins, sizes are given in kilodaltons (kDa) or amino acid residue numbers. Proteins sizes are estimated from gel electrophoresis, from sequenced proteins, from derived amino acid sequences, or from published protein sequences.
Oligonucleotides that are not commercially available can be chemically synthesized, e.g., according to the solid phase phosphoramidite triester method first described by Beaucage &Caruthers, Tetrahedron Lett. 22: 1859-1862 (1981) , using an automated synthesizer, as described in Van Devanter et. al., Nucleic Acids Res. 12: 6159-6168 (1984) . Purification of oligonucleotides is performed using any art-recognized strategy, e.g., native acrylamide gel electrophoresis or anion-exchange HPLC as described in Pearson &Reanier, J. Chrom. 255: 137-149 (1983) .
Recombinant production is an effective means to obtain peptides of this disclosure, particularly those of relatively large molecular weight, for example, a fusion peptide of a HER-2/Neu epitope and a GM-CSF. The sequence of a polynucleotide encoding a peptide of this disclosure, and synthetic oligonucleotides can be verified after cloning or subcloning using, e.g., the chain termination method for sequencing double-stranded templates of Wallace et al., Gene 16: 21-26 (1981) .
B.  Construction of an Expression Cassette
i.  Obtaining a Polynucleotide Sequence Encoding a Peptide
A polynucleotide sequence encoding a peptide of this disclosure can be obtained by chemical synthesis, or can be purchased from a commercial supplier, which may then be further manipulated using standard techniques of molecular cloning.
ii.  Modification of Nucleic Acids for Preferred Codon Usage in Host  Organism
The polynucleotide sequence encoding a peptide of this disclosure can be optionally altered to coincide with the preferred codon usage of a particular host. For example, the preferred codon usage of one strain of bacterial cells can be used to derive a polynucleotide that encodes a peptide of the disclosure and includes the codons favored by this strain. The frequency of preferred codon usage exhibited by a host cell can be calculated by averaging frequency of preferred codon usage  in a large number of genes expressed by the host cell (e.g., calculation service is available from web site of the Kazusa DNA Research Institute, Japan) . This analysis is preferably limited to genes that are highly expressed by the host cell.
At the completion of modification, the coding sequences are verified by sequencing and are then subcloned into an appropriate expression vector for recombinant production of the peptides of this disclosure.
Following verification of the coding sequence, the peptide of the present disclosure can be produced using routine techniques in the field of recombinant genetics.
C.  Expression Systems
To obtain high level expression of a nucleic acid encoding a peptide of the present disclosure, one typically subclones a polynucleotide encoding the peptide into an expression vector that contains a strong promoter to direct transcription, a transcription/translation terminator and a ribosome binding site for translational initiation. Suitable bacterial promoters are well known in the art and described, e.g., in Sambrook and Russell, supra, and Ausubel et al., supra. Bacterial expression systems for expressing a peptide of this disclosure are available in, e.g., E. coli, Bacillus sp., Salmonella, and Caulobacter. Kits for such expression systems are commercially available. Eukaryotic expression systems for mammalian cells, yeast, and insect cells are well known in the art and are also commercially available.
In one embodiment, the eukaryotic expression vector is an adenoviral vector, an adeno-associated vector, a retroviral vector, or a yeast artificial chromosome (YAC) . In some embodiments, the eukaryotic (e.g., mammalian) expression vector is a pEAK10, pEAK12, pEAK13, pCDNA3.0, pCDNA4.0, pCDM7, pCDM8, pCDM10, pCDM12. In some embodiments, the promoter is CMV, EF1α, or SV40.
For instance, the expression system comprises a promoter operably linked to the peptide nucleic acid (e.g., cDNA) sequence such that it is under control of a promoter. The term "promoter" describes the combination of the promoter (RNA polymerase binding site) and operators. The promoter may function in vivo or in a cell-free system. Any number of promoters may be used depending on the needs and preferences of the practitioner. Promoters for controlling recombinant protein, (e.g., peptide) expression can be any promoter for any DNA dependent RNA polymerase,  for example. Generally a promoter (e.g., T7, T3, and SP6 RNA promoters and compatible RNA polymerases) is selected for in vitro expression for producing recombinant protein in a bacterial system, e.g., E. coli, which usually requires the molecular inducer isopropyl-β-D-thiogalactoside (IPTG) for regulating the promoter’s transcriptional activity. In some cases, a modified Self-Inducible Expression system that utilizes lactose as an inducer may be used, as described in Briand et al, A self-inducible heterologous protein expression system in Escherichia coli. Sci Rep 6, 33037 (2016) .
The promoter used to direct expression of a heterologous nucleic acid depends on the particular application. The promoter is optionally positioned about the same distance from the heterologous transcription start site as it is from the transcription start site in its natural setting. As is known in the art, however, some variation in this distance can be accommodated without loss of promoter function.
In addition to the promoter, the expression vector typically includes a transcription unit or expression cassette that contains all the additional elements required for the expression of a peptide of this disclosure in host cells. A typical expression cassette thus contains a promoter operably linked to the polynucleotide sequence encoding the peptide and signals required for efficient polyadenylation of the transcript, ribosome binding sites, and translation termination. The nucleic acid sequence encoding the peptide is typically linked to a cleavable signal peptide sequence to promote secretion of the peptide by the transformed cell. Such signal peptides include, among others, the signal peptides from tissue plasminogen activator, insulin, and neuron growth factor, and juvenile hormone esterase of Heliothis virescens. Additional elements of the cassette may include enhancers and, if genomic DNA is used as the structural gene (e.g., encoding the heterologous polypeptide) , introns with functional splice donor and acceptor sites.
In addition to a promoter sequence, the expression cassette should also contain a transcription termination region downstream of the structural gene to provide for efficient termination. The termination region may be obtained from the same gene as the promoter sequence or may be obtained from different genes.
The particular expression vector used to transport the genetic information into the cell is not particularly critical. Any of the conventional vectors used for expression in eukaryotic or prokaryotic cells may be used. Standard bacterial expression vectors include plasmids such as  pBR322 based plasmids, pSKF, pET23D, and fusion expression systems such as GST and LacZ. Epitope tags can also be added to recombinant proteins to provide convenient methods of isolation, e.g., c-myc.
Expression vectors containing regulatory elements from eukaryotic viruses are typically used in eukaryotic expression vectors, e.g., SV40 vectors, papilloma virus vectors, and vectors derived from Epstein-Barr virus. Other exemplary eukaryotic vectors include pMSG, pAV009/A +, pMTO10/A +, pMAMneo-5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV40 early promoter, SV40 later promoter, metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.
Some expression systems have markers that provide gene amplification such as thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate reductase. Alternatively, high yield expression systems not involving gene amplification are also suitable, such as a baculovirus vector in insect cells, with a polynucleotide sequence encoding the peptide of this disclosure under the direction of the polyhedrin promoter or other strong baculovirus promoters.
The elements that are typically included in expression vectors also include a replicon that functions in E. coli, a gene encoding antibiotic resistance to permit selection of bacteria that harbor recombinant plasmids, and unique restriction sites in nonessential regions of the plasmid to allow insertion of eukaryotic sequences. The particular antibiotic resistance gene chosen is not critical, any of the many resistance genes known in the art are suitable. The prokaryotic sequences are optionally chosen such that they do not interfere with the replication of the DNA in eukaryotic cells, if necessary. Similar to antibiotic resistance selection markers, metabolic selection markers based on known metabolic pathways may also be used as a means for selecting transformed host cells.
When periplasmic expression of a recombinant protein (e.g., a peptide of the present disclosure) is desired, the expression vector further comprises a sequence encoding a secretion signal, such as the E. coli OppA (Periplasmic Oligopeptide Binding Protein) secretion signal or a modified version thereof, which is directly connected to 5'of the coding sequence of the protein to be expressed. This signal sequence directs the recombinant protein produced in cytoplasm through the cell membrane into the periplasmic space. The expression vector may further comprise  a coding sequence for signal peptidase 1, which is capable of enzymatically cleaving the signal sequence when the recombinant protein is entering the periplasmic space. More detailed description for periplasmic production of a recombinant protein can be found in, e.g., Gray et al., Gene 39: 247-254 (1985) , U.S. Patent Nos. 6,160,089 and 6,436,674.
D.  Transfection Methods
Standard transfection methods are used to produce bacterial, mammalian, yeast, insect, or plant cell lines that express large quantities of a peptide of this disclosure, which are then purified using standard techniques (see, e.g., Colley et al., J. Biol. Chem. 264: 17619-17622 (1989) ; Guide to Protein Purification, in Methods in Enzymology, vol. 182 (Deutscher, ed., 1990) ) . Transformation of eukaryotic and prokaryotic cells are performed according to standard techniques (see, e.g., Morrison, J. Bact. 132: 349-351 (1977) ; Clark-Curtiss &Curtiss, Methods in Enzymology 101: 347-362 (Wu et al., eds, 1983) .
Any of the well-known procedures for introducing foreign nucleotide sequences into host cells may be used. These include the use of calcium phosphate transfection, polybrene, protoplast fusion, electroporation, liposomes, microinjection, plasma vectors, viral vectors and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into a host cell (see, e.g., Sambrook and Russell, supra) . It is only necessary that the particular genetic engineering procedure used be capable of successfully introducing at least one gene into the host cell capable of expressing the peptide of this disclosure.
E.  Detection of Recombinant Expression of a Peptide in Host Cells
After the expression vector is introduced into appropriate host cells, the transfected cells are cultured under conditions favoring expression of the peptide of this disclosure. The cells are then screened for the expression of the recombinant peptide, which is subsequently recovered from the culture using standard techniques (see, e.g., Scopes, Protein Purification: Principles and Practice (1982) ; U.S. Patent No. 4,673,641; Ausubel et al., supra; and Sambrook and Russell, supra) .
Several general methods for screening gene expression are well known among those skilled in the art. First, gene expression can be detected at the nucleic acid level. A variety of methods of specific DNA and RNA measurement using nucleic acid hybridization techniques are  commonly used (e.g., Sambrook and Russell, supra) . Some methods involve an electrophoretic separation (e.g., Southern blot for detecting DNA and Northern blot for detecting RNA) , but detection of DNA or RNA can be carried out without electrophoresis as well (such as by dot blot) . The presence of nucleic acid encoding a peptide of this disclosure in transfected cells can also be detected by PCR or RT-PCR using sequence-specific primers.
Second, gene expression can be detected at the polypeptide level. Various immunological assays are routinely used by those skilled in the art to measure the level of a gene product, particularly using polyclonal or monoclonal antibodies that react specifically with a peptide of the present disclosure, particularly one containing a sufficiently large heterolougs polypeptide (e.g., Harlow and Lane, Antibodies, A Laboratory Manual, Chapter 14, Cold Spring Harbor, 1988; Kohler and Milstein, Nature, 256: 495-497 (1975) ) . Such techniques require antibody preparation by selecting antibodies with high specificity against the peptide or an antigenic portion thereof. The methods of raising polyclonal and monoclonal antibodies are well established and their descriptions can be found in the literature, see, e.g., Harlow and Lane, supra; Kohler and Milstein, Eur. J. Immunol., 6: 511-519 (1976) .
F.  Purification of Peptides
i.  Purification of Chemically Synthesized Peptides
Purification of synthetic peptides is accomplished using various methods of chromatography, such as reverse phase HPLC, gel permeation, ion exchange, size exclusion, affinity, partition, or countercurrent distribution. The choices of appropriate matrices and buffers are well known in the art.
ii.  Purification of Chemically Synthesized Peptides
1. Purification of Peptides from Bacterial Inclusion Bodies
a. Solubility Fractionation
Often as an initial step, and if the protein mixture is complex, an initial salt fractionation can separate many of the unwanted host cell proteins (or proteins derived from the cell culture media) from the recombinant protein of interest, e.g., a peptide of the present disclosure. The preferred salt is ammonium sulfate. Ammonium sulfate precipitates proteins by effectively reducing the amount of water in the protein mixture. Proteins then precipitate on the basis of their  solubility. The more hydrophobic a protein is, the more likely it is to precipitate at lower ammonium sulfate concentrations. A typical protocol is to add saturated ammonium sulfate to a protein solution so that the resultant ammonium sulfate concentration is between 20-30%. This will precipitate the most hydrophobic proteins. The precipitate is discarded (unless the protein of interest is hydrophobic) and ammonium sulfate is added to the supernatant to a concentration known to precipitate the protein of interest. The precipitate is then solubilized in buffer and the excess salt removed if necessary, through either dialysis or diafiltration. Other methods that rely on solubility of proteins, such as cold ethanol precipitation, are well known to those of skill in the art and can be used to fractionate complex protein mixtures.
b. Size Differential Filtration
Based on a calculated molecular weight, a protein of greater and lesser size can be isolated using ultrafiltration through membranes of different pore sizes (for example, Amicon or Millipore membranes) . As a first step, the protein mixture is ultrafiltered through a membrane with a pore size that has a lower molecular weight cut-off than the molecular weight of a protein of interest, e.g., a peptide of the present disclosure. The retentate of the ultrafiltration is then ultrafiltered against a membrane with a molecular cut off greater than the molecular weight of the peptide of interest. The recombinant protein will pass through the membrane into the filtrate. The filtrate can then be chromatographed as described below.
c. Column Chromatography
A protein of interest (such as a peptide of the present disclosure) can also be separated from other proteins on the basis of its size, net surface charge, hydrophobicity, or affinity for ligands. In addition, antibodies raised against a peptide of this disclosure can be conjugated to column matrices and the peptide immunopurified. All of these methods are well known in the art.
It will be apparent to one of skill that chromatographic techniques can be performed at any scale and using equipment from many different manufacturers (e.g., Pharmacia Biotech) .
When a peptide of the present disclosure is produced recombinantly by transformed bacteria in large amounts, typically after promoter induction, although expression can be constitutive, the peptides may form insoluble aggregates. There are several protocols that are suitable for purification of protein inclusion bodies. For example, purification of aggregate  proteins (hereinafter referred to as inclusion bodies) typically involves the extraction, separation and/or purification of inclusion bodies by disruption of bacterial cells, e.g., by incubation in a buffer of about 100-150 μg/ml lysozyme and 0.1%Nonidet P40, a non-ionic detergent. The cell suspension can be ground using a Polytron grinder (Brinkman Instruments, Westbury, NY) . Alternatively, the cells can be sonicated on ice. Alternate methods of lysing bacteria are described in Ausubel et al. and Sambrook and Russell, both supra, and will be apparent to those of skill in the art.
The cell suspension is generally centrifuged and the pellet containing the inclusion bodies resuspended in buffer which does not dissolve but washes the inclusion bodies, e.g., 20 mM Tris-HCl (pH 7.2) , l mM EDTA, 150 mM NaCl and 2%Triton-X 100, a non-ionic detergent. It may be necessary to repeat the wash step to remove as much cellular debris as possible. The remaining pellet of inclusion bodies may be resuspended in an appropriate buffer (e.g., 20 mM sodium phosphate, pH 6.8, 150 mM NaCl) . Other appropriate buffers will be apparent to those of skill in the art.
Following the washing step, the inclusion bodies are solubilized by the addition of a solvent that is both a strong hydrogen acceptor and a strong hydrogen donor (or a combination of solvents each having one of these properties) . The proteins that formed the inclusion bodies may then be renatured by dilution or dialysis with a compatible buffer. Suitable solvents include, but are not limited to, urea (from about 4 M to about 8 M) , formamide (at least about 80%, volume/volume basis) , and guanidine hydrochloride (from about 4 M to about 8 M) . Some solvents that are capable of solubilizing aggregate-forming proteins, such as SDS (sodium dodecyl sulfate) and 70%formic acid, may be inappropriate for use in this procedure due to the possibility of irreversible denaturation of the proteins, accompanied by a lack of immunogenicity and/or activity. Although guanidine hydrochloride and similar agents are denaturants, this denaturation is not irreversible and renaturation may occur upon removal (by dialysis, for example) or dilution of the denaturant, allowing re-formation of the immunologically and/or biologically active protein of interest. After solubilization, the protein can be separated from other bacterial proteins by standard separation techniques. For further description of purifying recombinant polypeptides from bacterial inclusion body, see, e.g., Patra et al., Protein Expression and Purification 18: 182-190 (2000) .
Alternatively, it is possible to purify recombinant polypeptides, e.g., a peptide of this disclosure, from bacterial periplasm. Where the recombinant polypeptide is exported into the periplasm of the bacteria, the periplasmic fraction of the bacteria can be isolated by cold osmotic shock in addition to other methods known to those of skill in the art (see e.g., Ausubel et al., supra) . To isolate recombinant peptides from the periplasm, the bacterial cells are centrifuged to form a pellet. The pellet is resuspended in a buffer containing 20%sucrose. To lyse the cells, the bacteria are centrifuged and the pellet is resuspended in ice-cold 5 mM MgSO 4 and kept in an ice bath for approximately 10 minutes. The cell suspension is centrifuged and the supernatant decanted and saved. The recombinant peptides present in the supernatant can be separated from the host proteins by standard separation techniques well known to those of skill in the art.
2. Standard Protein Separation Techniques for Purification
When a recombinant polypeptide, e.g., a peptide of the present disclosure, is expressed in host cells in a soluble form, its purification can follow the standard protein purification procedure described below. This standard purification procedure is also suitable for purifying peptides obtained from chemical synthesis.
G.  Confirmation of Peptide Sequence
The amino acid sequence of a peptide of this disclosure can be confirmed by a number of well established methods. For example, the conventional method of Edman degradation can be used to determine the amino acid sequence of a peptide. Several variations of sequencing methods based on Edman degradation, including microsequencing, and methods based on mass spectrometry are also frequently used for this purpose.
H.  Modification of Peptides
The peptides of the present disclosure can be modified to achieve more desirable properties. The design of chemically modified peptides and peptide mimics that are resistant to degradation by proteolytic enzymes or have improved solubility or binding ability is well known.
Modified amino acids or chemical derivatives of the T cell or B cell peptides or fusion peptides of this disclosure may contain additional chemical moieties of modified amino acids not normally a part of the T cell or B cell protein. Covalent modifications of the peptides are within the scope of the present disclosure. Such modifications may be introduced into a peptide by  reacting targeted amino acid residues of the peptide with an organic derivatizing agent that is capable of reacting with selected side chains or terminal residues. The following examples of chemical derivatives are provided by way of illustration and not by way of limitation.
The design of peptide mimics which are resistant to degradation by proteolytic enzymes is known to those skilled in the art. See e.g., Sawyer, Structure-Based Drug Design, P. Verapandia, Ed., N.Y. (1997) ; U.S. Patent Nos. 5,552,534 and 5,550,251. Both peptide backbone and side chain modifications may be used in designing secondary structure mimicry. Possible modifications include substitution of D-amino acids, N α-Me-amino acids, C α-Me-amino acids, and dehydroamino acids. To this date, a variety of secondary structure mimetics have been designed and incorporated in peptides or peptidomimetics.
Other modifications include substitution of a natural amino acid with an unnatural hydroxylated amino acid, substitution of the carboxy groups in acidic amino acids with nitrile derivatives, substitution of the hydroxyl groups in basic amino acids with alkyl groups, or substitution of methionine with methionine sulfoxide. In addition, an amino acid of a HER-2/Neu peptide or a fusion peptide of this disclosure can be replaced by the same amino acid but of the opposite chirality, i.e., a naturally-occurring L-amino acid may be replaced by its D-configuration.
I.  Recombinant Proteins and Synthetic Peptides
Full length recombinant whole or variants thereof of SARS-CoV or SARS-CoV-2 virus and/or other coronavirous, or recombinant functional proteins, e.g., S protein, N protein, M protein, from SARS-CoV or SARS-CoV-2, and/or other coronavirous may be generated and purified using technologies known in the field. For example, full length SARS-CoV-2 can be generated based on “transformation-associated recombination” (TAR) in yeast (Thao et al. (2020) , “Rapid reconstruction of SARS-CoV-2 using a synthetic genomics platform, ” . bioRxiv, 2020.02.21.959817) . As an example, synthetic proteins (e.g., defined as >35 amino acids) and peptides from the SARS-CoV or SARS-CoV-2 recombinant binding domain (RBD) may be synthesized in the Peptide Core at Los Alamos National Laboratory. Positive control rabbit serum (polyclonal, against SARS/SARS-CoV-2 Coronavirus spike protein subunit 1) is Invitrogen PA5-81795. See e.g., Schein et al., 2021. “Synthetic proteins for COVID-19 diagnostics. ” Peptides. 143: 170583. doi: 10.1016/j. peptides. 2021.170583.
In some embodiments, provided herein is any protein fragment (meaning a polypeptide sequence at least one amino acid residue shorter than a reference polypeptide sequence but otherwise identical) of a reference protein having a length of 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500 or longer than 500 amino acids. In another example, any protein that includes a stretch of 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 (contiguous) amino acids that are 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%identical to any of the sequences described herein can be utilized in accordance with the disclosure. In some embodiments, a polypeptide includes 2, 3, 4, 5, 6, 7, 8, 9, 10, or more mutations as shown in any of the sequences provided herein or referenced herein. In another example, any protein that includes a stretch of 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, or 500 amino acids that are greater than 40%, 50, 60%, 70%, 80%, 90%, 95%, or 100%identical to any of the sequences described herein, wherein the protein has a stretch of 5, 10, 15, 20, 25, or 30 amino acids that are less than 95%, 90%, 85%, 80%, 75%, 70%, 65%to 60%identical to any of the sequences described herein can be utilized in accordance with the disclosure.
IV. FUSING T CELL AND/OR B CELL EPITOPE WITH A HETEROLOGOUS POLYPEPTIDE
In one aspect of this disclosure, a peptide corresponding to a SARS-CoV or SARS-CoV-2 promiscuous T cell epitope or B cell epitope is attached to a heterologous polypeptide via a covalent bond to form a fusion peptide, such that the ability of the SARS-CoV or SARS-CoV-2 epitope to induce a T cell or B cell response is enhanced. Usually, this covalent bond is a peptide bond and the SARS-CoV and/or SARS-CoV-2 epitope and the heterologous polypeptide form a new polypeptide. This peptide bond may be a direct peptide bond between the SARS-CoV and/or SARS-CoV-2 epitope and the heterologous polypeptide, or it may be an indirect peptide bond provided by way of a peptide linker between the SARS-CoV or SARS-CoV-2 epitope and the heterologous polypeptide. As an example, the SARS-CoV and/or SARS-CoV-2 epitope may be linked to a detecting molecule such as a GFP, RFP, dsRed, or the like. As another example, the SARS-CoV and/or SARS-CoV-2 epitope may be linked to an affinity tag such as a six histidine (6His) on the N-terminal or the C-terminal, or the like.
Other covalent bonds are also suitable for the purpose of fusing the SARS-CoV or SARS-CoV-2 peptide with the heterologous polypeptide. For instance, a functional group (such as a non- terminal amine group, a non-terminal carboxylic acid group, a hydroxyl group, and a sulfhydryl group) of one peptide may easily react with a functional group of the other peptide and establish a covalent bond, other than a peptide bond, that conjugates the two peptides. A covalent connection between a peptide of a SARS-CoV or SARS-CoV-2 epitope and a heterologous polypeptide can also be provided by way of a linker molecule with suitable functional group (s) . Such a linker molecule can be a peptide linker or a non-peptide linker. A linker may be derivatized to expose or to attach additional reactive functional groups prior to conjugation. The derivatization may involve attachment of any of a number of molecules such as those available from Pierce Chemical Company, Rockford, Illinois.
V. FUNCTIONAL ASSAYS
Adaptive immune response is mediated by the T lymphocytes and B lymphocytes. T lymphocytes express T-cell receptors (TCR) and detect short peptides produced by a proteolytic machinery and presented by major histocompatibility complex (MHC) molecules at the cell surface. (Bercovici et al., 2000) . Through antigen processing, T lymphocytes are able to detect foreign peptides synthesized by infected cells. T cells which express CD8 molecules have the capacity to lyse directly the target cells. A subset of CD4 + T lymphocytes is specialized in regulating the immune response via cytokine secretion and activation of the antigen-presenting cells (APC) . Chromium release assays and limiting-dilution analyses have been commonly used to measure specific T-cell responses. Additionally, various assays are available for immune monitoring of specific T-cell responses. These various assays are schematically divided into functional assays, which measure the secretion of a particular cytokine (ELISPOT and intracellular cytokines) ; assays which assess the specificity of the T cells irrespective of their functionality and which are based on structural features of the TCR (tetramers and immunoscope) ; and assays aimed at detecting T-cell precursors by amplifying cells that proliferate in response to antigenic stimulation.
A SARS-CoV or SARS-CoV-2 epitope of this disclosure (or a fusion peptide comprising a SARS-CoV or SARS-CoV-2 epitope and a heterologous polypeptide) is useful for its capability to induce a T cell immune response specific to a SARS-CoV or SARS-CoV-2 protein, when the epitope is presented by an APC that may have one of a HLA-A, HLA-B, and/or HLA-DR allele. Various functional assays can be used to confirm the ability of a SARS-CoV or SARS-CoV-2  epitope to induce such a SARS-CoV or SARS-CoV-2 specific T cell immune response in a promiscuous manner with regard to antigen presenting cells of different HLA alleles, including proliferation assay and flow cytometry assays detecting the binding between a T cell receptor and a peptide epitope or the production of cytokines by T cells.
As an illustration, the function assay ELISPOT may be used for this purpose. The ELISPOT (enzyme-linked immunospot) technique detects T cells that secrete a given cytokine (e.g., gamma interferon [IFN-γ] ) in response to an antigenic stimulation (e.g., an antigen presented by SARS-CoV or SARS-CoV-2) . Briefly, T cells are cultured with antigen-presenting cells in wells which have been coated with anti-IFN-γ antibodies. The secreted IFN-γ is captured by the coated antibody and then revealed with a second antibody coupled to a chromogenic substrate. Thus, locally secreted cytokine molecules form spots, with each spot corresponding to one IFN-γ-secreting cell. The number of spots allows one to determine the frequency of IFN-γ-secreting cells specific for a given antigen in the analyzed sample. The ELISPOT assay has also been described for the detection of tumor necrosis factor alpha, interleukin-4 (IL-4) , IL-5, IL-6, IL-10, IL-12, granulocyte-macrophage colony-stimulating factor, and granzyme B-secreting lymphocytes.
As another illustration, the structural assay tetramers may be used for this purpose. T cells recognize short peptides presented by MHC molecules through their clonotypic TCR. Tetramers of MHC class I-peptide complexes have been used in cytometry to enumerate, characterize, and purify peptide-specific CD8 cells. The heavy and light chains of the MHC are produced in Escherichia coli, solubilized in urea, and refolded in vitro in the presence of high concentrations of the antigenic peptide. The refolded complexes are purified by gel filtration, and a single biotin is added at the C-terminal end of the heavy chain using the bacterial BirA enzyme. Incubation with fluorescent streptavidin yields tetramers which can be used like any clonotypic antibody. Tetramers of MHC class II molecules may also be produced and used to analyze CD4 + T-cell responses.
B lymphocytes recognize intact proteins and produce immunoglobulins (Ig) . (Chaplin, 2010) . B cell response is activated in two pathways: T cell-dependent and T cell-independent. In a T-cell dependent manner, antigens that activate T cells as well as B cells establish Ig responses in which T cells provide ‘help’ for the B cells to mature. T cell-independent B cell activation occurs without the assistance of T cell co-stimulatory proteins. In the absence of co-stimulators,  monomeric antigens are unable to activate B cells. Polymeric antigens with a repeating structure, in contrast, are able to activate B cells, probably because they can crosslink and cluster Ig molecules on the B cell surface. T cell-independent antigens include bacterial lipopolysaccharide (LPS) , certain other polymeric polysaccharides, and certain polymeric proteins.
Several techniques have been used for probing the B cell response in vitro and in vivo by taking advantage of the specificity of B cell receptor (BCR) -associated and secreted antibodies. These include ELISPOT, flow cytometry, mass cytometry, and fluorescence microscopy to identify and/or isolate primary antigen-specific B cells. (Jim Boonyaratanakornkit and Justin Taylor, 2019) . The ELISPOT assay is particularly suitable for the purpose of detecting B cell epitope response to SARS-CoV and/or SARS-CoV-2 infection in relation to the disclosure described herein.
Another example for confirming B cell immune response specific to a SARS-CoV or SARS-CoV-2 protein is flow cytometry-based analysis of antigen-specific B cells. This technique is dependent on labeling antigen with a fluorescent tag to allow detection. Fluorochromes can either be attached covalently via chemical conjugation to the antigen, expressed as a recombinant fusion protein, or attached non-covalently by biotinylating the antigen. After biotinylation, fluorochrome-conjugated streptavidin is added to generate a labeled tetramer of the antigen. Biotinylation of the antigen may be set at a ratio ≤1 biotin to 1 antigen. Alternatively, site directed biotinylation can be accomplished by adding either an AviTag or BioEase tag to the recombinant antigen prior to expression.
VI. METHOD FOR ELICITING AN IMMUNE RESPONSE IN A SUBJECT
The present disclosure further provides a method for eliciting an immune response in a subject in need thereof such as at risk of exposure to SARS-CoV or SARS-CoV-2 infection. According to the CDC, the term exposure refers to contact with infectious agents (bacteria or virus, e.g., SARS-CoV or SARS-CoV-2) in a manner that promotes transmission and increases the likihood of disease. For instance, a subject who has close contact with someone who has COVID-19 is considered at risk of exposure to the disease. The term close contact refers to within 6 feet of someone for a cumulative total of 15 minites or more over a 24-hour period. Subjects who have underlying medical conditions (e.g., immune-comprimised due to treatment of immunosuppressants, under chemotherapy, prior heart and cardio diseases, prior pulmonary  diseases, cancer, diabetes, obesity, etc. ) , age (e.g., 65 or over) , geneteic predisposition, and/or pregnant or recently pregnant. The subject may be at risk for severe illness when contracted with COVID-19 such that the subject may need hospitalization, intensive care, a ventilator to help them breathe or they may even die. Elicitation an immune response in a subject may provide prophylactic and therapeutic applications. The method includes the following steps: first, lymphocytes including at least a T cell epitope, and/or at least a B cell epitope, and/or optionally an antigen-presenting cell are obtained from a patient. Suitable samples that yield such lymphocytes include blood, and lymph nodes or lymphatic fluids. Second, at least a T cell epitope, and/or at least a B cell epitope, and/or optionally an antigen-presenting cell are exposed to a SARS-CoV or SARS-CoV-2 peptide (or a fusion peptide comprising the SARS-CoV or SARS-CoV-2 peptide and a heterologous peptide) of this disclosure under conditions that would allow, e.g., proper presentation of a T cell epitope by the antigen-presenting cell to the T cell. Third, signs of a T cell response and/or B cell response is measured in vitro by means well known in the art such as ELISPOT, ELISA, proliferation assay, or flow cytometry. When a T cell response and/or B cell response is detected by any of these methods, it can be concluded that there exists a T cell and/or a B cell immune response specific to a SARS-CoV or SARS-CoV-2 protein in the patient.
A.  Vaccines
In certain preferred embodiments of the present disclosure, vaccines are provided. The vaccines will generally comprise a peptide comprising one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2, such as those discussed above, in combination with an immunostimulant. An immunostimulant may be any substance that enhances or potentiates an immune response (antibody and/or cell-mediated) to an exogenous antigen. Examples of immunostimulants include adjuvants, biodegradable microspheres (e.g., polylactic galactide) and liposomes (into which the compound is incorporated; see, e.g., U.S. Patent No. 4,235,877) . Vaccine preparation is generally described in, for example, Powell &Newman, eds., Vaccine Design (the subunit and adjuvant approach) (1995) . Pharmaceutical compositions and vaccines within the scope of the present disclosure may also contain other compounds, which may be biologically active or inactive. For example, one or more T cell and/or B cell epitopes derived from immunogenic portions of other coronavirus antigens (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV- NL63) may be present, either incorporated into a fusion polypeptide or as a separate compound, within the composition or vaccine. Accordingly, a pharmaceutical vaccine formulation may comprise the petide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2, one or more adjuvnts, pharmaceutically-acceptable carriers or other ingredients including immunological adjuvants routinely provided in vaccine formulation. Suitable adjuvants are described in detail below.
In some embodiments, the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is biologically produced such as using an expression cassette in a host cell. Short to medium length peptides, for example peptides that do no require specific folding, can be chemically synthesized. Large scale production of chemically synthesized peptides may be used for manufacturing large quantities of peptide vaccines. See e.g., Bray, B.L., 2003. Large-scale manufacture of peptide therapeutics by chemical synthesis. Nature Reviews Drug Discovery, 2 (7) , pp. 587-593.
In some embodiments, the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is a cationic peptide. As described herein, a “cationic peptide” refers to a peptide that is positively charged at a pH in the range of 5.0 to 8.0. The net charge on the peptide or peptide cocktails is calculated by assigning a +1 charge for each lysine (K) , arginine (R) or histidine (H) , a -1 charge for each aspartic acid (D) or glutamic acid (E) and a charge of 0 for the other amino acid within the sequence. The charge contributions from the N-terminal amine (+1) and C-terminal carboxylate (-1) end groups of each peptide effectively cancel each other when unsubstituted. The charges are summed for each peptide and expressed as the net average charge. A suitable peptide has a net average positive charge of +1. Preferably, the peptide has a net positive charge in the range that is larger than +2.
In some embodiments, the peptide comprising the one or more T cell and/or B cell epitopes derived from the S protein, N protein, or full length or a fragment thereof, of a SARS-CoV and/or SARS-CoV-2 is an anionic peptide. As described herein, an “anionic molecule” refers to a molecule that is negatively charged at a pH in the range of 5.0-8.0. The net negative  charge on the oligomer or polymer is calculated by assigning a -1 charge for each phosphodiester or phosphorothioate group in the oligomer.
Illustrative vaccines may contain polynucleotides encoding one or more of the polypeptides (e.g., T cell and/or B cell epitopes derived from one or more immunogenic portions of a SARS-CoV, SARS-CoV-2, and/or other coronaviruses) as described above, such that the polypeptide is generated in situ. As noted above, the polynucleotides may be present within any of a variety of delivery systems known to those of ordinary skill in the art, including nucleic acid expression systems, bacteria and viral expression systems. Numerous gene delivery techniques are well known in the art, such as those described by Rolland, Crit. Rev. Therap. Drug Carrier Systems 15: 143-198 (1998) , and references cited therein. Appropriate nucleic acid expression systems contain the necessary polynucleotides sequences for expression in the patient (such as a suitable promoter and terminating signal) . Bacterial delivery systems involve the administration of a bacterium (such as Bacillus-Calmette-Guerrin) that expresses an immunogenic portion of the polypeptide on its cell surface or secretes such an epitope. In one embodiment, the polynucleotides may be introduced using a viral expression system (e.g., vaccinia or other pox virus, retrovirus, or adenovirus) , which may involve the use of a non-pathogenic (defective) , replication competent virus. Suitable systems are disclosed, for example, in Fisher-Hoch et al., Proc. Natl. Acad. Sci. USA 86: 317-321 (1989) ; Flexner et al., Ann. N.Y. Acad. Sci. 569: 86-103 (1989) ; Flexner et al., Vaccine 8: 17-21 (1990) ; U.S. Patent Nos. 4,603,112, 4,769,330, and 5,017,487; WO 89/01973; U.S. Patent No. 4,777,127; GB 2,200,651; EP 0,345,242; WO 91/02805; Berkner, Biotechniques 6: 616-627 (1988) ; Rosenfeld et al., Science 252: 431-434 (1991) ; Kolls et al., Proc. Natl. Acad. Sci. USA 91: 215-219 (1994) ; Kass-Eisler et al., Proc. Natl. Acad. Sci. USA 90: 11498-11502 (1993) ; Guzman et al., Circulation 88: 2838-2848 (1993) ; and Guzman et al., Cir. Res. 73: 1202-1207 (1993) . Techniques for incorporating polynucleotides into such expression systems are well known to those of ordinary skill in the art. The polynucleotides may also be “naked, ” as described, for example, in Ulmer et al., Science 259: 1745-1749 (1993) and reviewed by Cohen, Science 259: 1691-1692 (1993) . The uptake of naked polynucleotides may be increased by coating the polynucleotides onto biodegradable beads, which are efficiently transported into the cells. It will be apparent that a vaccine may comprise both a polynucleotide and a polypeptide component. Such vaccines may provide for an enhanced immune response.
In some embodiments, the polynucleotides is formulated in lipid nanoparticles.. In some embodiments, the lipid nanoparticle is a mucus penetrating lipid nanoparticle. In some embodiments, the lipid nanoparticle is a solid lipid nanoparticle. Formulations for polynucleotides in lipid nanoparticles are discussed in U.S. Patent No: 10272150, which is incorporated herein in its entirety.
It will be apparent that a vaccine may contain pharmaceutically acceptable salts of the polynucleotides and polypeptides provided herein. Such salts may be prepared from pharmaceutically acceptable non-toxic bases, including organic bases (e.g., salts of primary, secondary and tertiary amines and basic amino acids) and inorganic bases (e.g., sodium, potassium, lithium, ammonium, calcium and magnesium salts) .
In some embodiments the peptide vaccines and/or the polynucleotide vaccines of the disclosure are superior to conventional vaccines by a factor of at least 2 fold, 5 fold, 10 fold, 20 fold, 40 fold, 50 fold, 100 fold, 500 fold or 1,000 fold.
While any suitable carrier known to those of ordinary skill in the art may be employed in the vaccine compositions of this disclosure, the type of carrier will vary depending on the mode of administration. Compositions of the present disclosure may be formulated for any appropriate manner of administration, including for example, topical, oral, nasal, intravenous, intracranial, intraperitoneal, subcutaneous or intramuscular administration. For parenteral administration, such as subcutaneous injection, the carrier preferably comprises water, saline, alcohol, a fat, a wax or a buffer. For oral administration, any of the above carriers or a solid carrier, such as mannitol, lactose, starch, magnesium stearate, sodium saccharine, talcum, cellulose, glucose, sucrose, and magnesium carbonate, may be employed. Biodegradable microspheres (e.g., polylactate polyglycolate) may also be employed as carriers for the pharmaceutical compositions of this disclosure. Suitable biodegradable microspheres are disclosed, for example, in U.S. Patent Nos. 4,897,268; 5,075,109; 5,928,647; 5,811,128; 5,820,883; 5,853,763; 5,814,344 and 5,942,252. One may also employ a carrier comprising the particulate-protein complexes described in U.S. Patent No. 5,928,647, which are capable of inducing a class I-restricted cytotoxic T lymphocyte responses in a host.
Such compositions may also comprise buffers (e.g., neutral buffered saline or phosphate buffered saline) , carbohydrates (e.g., glucose, mannose, sucrose or dextrans) , mannitol, proteins,  polypeptides or amino acids such as glycine, antioxidants, bacteriostats, chelating agents such as EDTA or glutathione, adjuvants (e.g., aluminum hydroxide) , solutes that render the formulation isotonic, hypotonic or weakly hypertonic with the blood of a recipient, suspending agents, thickening agents and/or preservatives. Alternatively, compositions of the present disclosure may be formulated as a lyophilizate. Compounds may also be encapsulated within liposomes using well known technology.
B.  Adjuvants
Any of a variety of immunostimulants may be employed in the vaccines of this disclosure. For example, an adjuvant or an immune potentiator may be included. An adjuvant may act as a co-signal to prime T-cells and/or B-cells and/or NK cells as to the existence of an infection. Adjuvants useful in the present disclosure may include, but are not limited to, natural or synthetic. They may be organic or inorganic. In some embodiments, adjuvants useful in the present disclosure include adjuvants for vaccines (e.g., influenza vaccines) as shown in Table 48. Adjuvants for DNA nucleic acid vaccines (DNA) have been disclosed in, for example, Kobiyama, et al Vaccines, 2013, 1 (3) , 278-292, the contents of which are incorporated herein by reference in their entirety.
Table 48. Adjuvants.
Figure PCTCN2022070948-appb-000001
Figure PCTCN2022070948-appb-000002
Most adjuvants contain a substance designed to protect the antigen from rapid catabolism, such as aluminum hydroxide or mineral oil, and a stimulator of immune responses, such as lipid A, Bortadella pertussis or Mycobacterium species or Mycobacterium derived proteins. For example, delipidated, deglycolipidated M. vaccae ( “pVac” ) can be used. Suitable adjuvants are commercially available as, for example, Freund’s Incomplete Adjuvant and Complete Adjuvant (Difco Laboratories, Detroit, MI) ; Merck Adjuvant 65 (Merck and Company, Inc., Rahway, NJ) ; AS-2 and derivatives thereof (SmithKline Beecham, Philadelphia, PA) ; CWS, TDM, Leif, aluminum salts such as aluminum hydroxide gel (alum) or aluminum phosphate; salts of calcium, iron or zinc; an insoluble suspension of acylated tyrosine; acylated sugars; cationically or anionically derivatized polysaccharides; polyphosphazenes; biodegradable microspheres; monophosphoryl lipid A and quil A. Cytokines, such as GM-CSF or interleukin-2, -7, or -12, may also be used as adjuvants.
Selection of appropriate adjuvants will be evident to one of ordinary skill in the art. Specific adjuvants may include, without limitation, cationic liposome-DNA complex JVRS-100, aluminum hydroxide vaccine adjuvant, aluminum phosphate vaccine adjuvant, aluminum potassium sulfate adjuvant, alhydrogel, ISCOM (s)  TM, Freund's Complete Adjuvant, Freund's  Incomplete Adjuvant, CpG DNA Vaccine Adjuvant, Cholera toxin, Cholera toxin B subunit, Liposomes, Saponin Vaccine Adjuvant, DDA Adjuvant, Squalene-based Adjuvants, Etx B subunit Adjuvant, IL-12 Vaccine Adjuvant, LTK63 Vaccine Mutant Adjuvant, TiterMax Gold Adjuvant, Ribi Vaccine Adjuvant, Corynebacterium-derived P40 Vaccine Adjuvant, MPL TM Adjuvant, AS04, AS02, Lipopolysaccharide Vaccine Adjuvant, Muramyl Dipeptide Adjuvant, CRL1005, Killed Corynebacterium parvum Vaccine Adjuvant, Montanide ISA 51, Bordetella pertussis component Vaccine Adjuvant, Cationic Liposomal Vaccine Adjuvant, Adamantylamide Dipeptide Vaccine Adjuvant, Arlacel A, VSA-3 Adjuvant, Aluminum vaccine adjuvant, Polygen Vaccine Adjuvant, Adjumer TM, Algal Glucan, Bay R1005, 
Figure PCTCN2022070948-appb-000003
Stearyl Tyrosine, Specol, Algammulin, 
Figure PCTCN2022070948-appb-000004
Calcium Phosphate Gel, CTA1-DD gene fusion protein, DOC/Alum Complex, Gamma Inulin, Gerbu Adjuvant, GM-CSF, GMDP, Recombinant hIFN-gammallnterferon-g, Interleukin-1β, Interleukin-2, Interleukin-7, Sclavo peptide, Rehydragel LV, Rehydragel HPA, Loxoribine, MF59, MTP-PE Liposomes, Murametide, Murapalmitine, D-Murapalmitine, NAGO, Non-Ionic Surfactant Vesicles, PMMA, Protein Cochleates, QS-21, SPT (Antigen Formulation) , nanoemulsion vaccine adjuvant, AS03, Quil-A vaccine adjuvant, LTR192G Vaccine Adjuvant, E. coli heat-labile toxin, LT, amorphous aluminum hydroxyphosphate sulfate adjuvant, Calcium phosphate vaccine adjuvant, Montanide Incomplete Seppic Adjuvant, Imiquimod, Resiquimod, AF03, Flagellin, Poly (I: C) , 
Figure PCTCN2022070948-appb-000005
Abisco-100 vaccine adjuvant, Albumin-heparin microparticles vaccine adjuvant, AS-2 vaccine adjuvant, B7-2 vaccine adjuvant, DHEA vaccine adjuvant, Immunoliposomes Containing Antibodies to Costimulatory Molecules, Sendai Proteoliposomes, Sendai-containing Lipid Matrices, Threonyl muramyl dipeptide (TMDP) , Ty Particles vaccine adjuvant, Bupivacaine vaccine adjuvant, DL-PGL (Polyester poly (DL-lactide-co-glycolide) ) vaccine adjuvant, IL-15 vaccine adjuvant, LTK72 vaccine adjuvant, MPL-SE vaccine adjuvant, non-toxic mutant E112K of Cholera Toxin mCT-E112K, and/or Matrix-S.
Other adjuvants which may be co-administered with the polypeptides of the disclosure include, but are not limited to interferons, TNF-alpha, TNF-beta, chemokines such as CCL21, eotaxin, HMGB1, SA100-8alpha, GCSF, GMCSF, granulysin, lactoferrin, ovalbumin, CD-40L, CD28 agonists, PD-1, soluble PD1, L1 or L2, or interleukins such as IL-1, IL-2, IL-4, IL-6, IL-7, IL-10, IL-12, IL-13, IL-21, IL-23, IL-15, IL-17, and IL-18.
Within the vaccines provided herein, the adjuvant composition is preferably designed to induce an immune response predominantly of the Th1 type. High levels of Th1-type cytokines (e.g., IFN-γ, TNFα, IL-2 and IL-12) tend to favor the induction of cell mediated immune responses to an administered antigen. In contrast, high levels of Th2-type cytokines (e.g., IL-4, IL-5, IL-6 and IL-10) tend to favor the induction of humoral immune responses. Following application of a vaccine as provided herein, a patient will support an immune response that includes Th1-and Th2-type responses. Within a preferred embodiment, in which a response is predominantly Th1-type, the level of Th1-type cytokines will increase to a greater extent than the level of Th2-type cytokines. The levels of these cytokines may be readily assessed using standard assays. For a review of the families of cytokines, see Mosmann &Coffman, Ann. Rev. Immunol. 7: 145-173 (1989) .
Preferred adjuvants for use in eliciting a predominantly Th1-type response include, for example, a combination of monophosphoryl lipid A, preferably 3-de-O-acylated monophosphoryl lipid A (3D-MPL) , together with an aluminum salt. MPL adjuvants are available from Corixa Corporation (Seattle, WA; see US Patent Nos. 4,436,727; 4,877,611; 4,866,034 and 4,912,094) . CpG-containing oligonucleotides (in which the CpG dinucleotide is unmethylated) also induce a predominantly Th1 response. Such oligonucleotides are well known and are described, for example, in WO 96/02555, WO 99/33488 and U.S. Patent Nos. 6,008,200 and 5,856,462. Immunostimulatory DNA sequences are also described, for example, by Sato et al., Science 273: 352 (1996) . Another preferred adjuvant comprises a saponin, such as Quil A, or derivatives thereof, including QS21 and QS7 (Aquila Biopharmaceuticals Inc., Framingham, MA) ; Escin; Digitonin; or Gypsophila or Chenopodium quinoa saponins . Other preferred formulations include more than one saponin in the adjuvant combinations of the present disclosure, for example combinations of at least two of the following group comprising QS21, QS7, Quil A, β-escin, or digitonin.
Alternatively the saponin formulations may be combined with vaccine vehicles composed of chitosan or other polycationic polymers, polylactide and polylactide-co-glycolide particles, poly-N-acetyl glucosamine-based polymer matrix, particles composed of polysaccharides or chemically modified polysaccharides, liposomes and lipid-based particles, particles composed of glycerol monoesters, etc. The saponins may also be formulated in the presence of cholesterol to form particulate structures such as liposomes or ISCOMs. Furthermore,  the saponins may be formulated together with a polyoxyethylene ether or ester, in either a non-particulate solution or suspension, or in a particulate structure such as a paucilamelar liposome or ISCOM. The saponins may also be formulated with excipients such as Carbopol R to increase viscosity, or may be formulated in a dry powder form with a powder excipient such as lactose.
In one embodiment, the adjuvant system includes the combination of a monophosphoryl lipid A and a saponin derivative, such as the combination of QS21 and
Figure PCTCN2022070948-appb-000006
adjuvant, as described in WO 94/00153, or a less reactogenic composition where the QS21 is quenched with cholesterol, as described in WO 96/33739. Other preferred formulations comprise an oil-in-water emulsion and tocopherol. Another particularly preferred adjuvant formulation employing QS21, 
Figure PCTCN2022070948-appb-000007
adjuvant and tocopherol in an oil-in-water emulsion is described in WO 95/17210.
Another enhanced adjuvant system involves the combination of a CpG-containing oligonucleotide and a saponin derivative particularly the combination of CpG and QS21 as disclosed in WO 00/09159. Preferably the formulation additionally comprises an oil in water emulsion and tocopherol.
Other adjuvants include Montanide ISA 720 (Seppic, France) , SAF-1 (Chiron, California, United States) , ISCOMS (CSL) , MF-59 (Chiron) , the SBAS series of adjuvants (e.g., SBAS-2, AS2’, AS2, ” SBAS-4, or SBAS6, available from SmithKline Beecham, Rixensart, Belgium) , Detox (Corixa, Hamilton, MT) , RC-529 (Corixa, Hamilton, MT) and other aminoalkyl glucosaminide 4-phosphates (AGPs) , such as those described in pending U.S. Patent Application Serial Nos. 08/853,826 and 09/074, 720, the disclosures of which are incorporated herein by reference in their entireties, and polyoxyethylene ether adjuvants such as those described in WO 99/52549A1.
Other adjuvants include adjuvant molecules of the general formula (I) : HO (CH 2CH 2O)  n-A-R, wherein, n is 1-50, A is a bond or –C (O) -, R is C 1-50 alkyl or Phenyl C 1-50 alkyl.
One embodiment of the present disclosure consists of a vaccine formulation comprising a polyoxyethylene ether of general formula (I) , wherein n is between 1 and 50, preferably 4-24, most preferably 9; the R component is C 1-50, preferably C 4-C 20 alkyl and most preferably C 12 alkyl, and A is a bond. The concentration of the polyoxyethylene ethers should be in the range 0.1-20%, preferably from 0.1-10%, and most preferably in the range 0.1-1%. Preferred polyoxyethylene  ethers are selected from the following group: polyoxyethylene-9-lauryl ether, polyoxyethylene-9-steoryl ether, polyoxyethylene-8-steoryl ether, polyoxyethylene-4-lauryl ether, polyoxyethylene-35-lauryl ether, and polyoxyethylene-23-lauryl ether. Polyoxyethylene ethers such as polyoxyethylene lauryl ether are described in the Merck index (12 th edition: entry 7717) . These adjuvant molecules are described in WO 99/52549.
The polyoxyethylene ether according to the general formula (I) above may, if desired, be combined with another adjuvant. For example, a preferred adjuvant combination is preferably with CpG as described in the pending UK patent application GB 9820956.2.
Any vaccine provided herein may be prepared using well known methods that result in a combination of antigen, immune response enhancer and a suitable carrier or excipient. The compositions described herein may be administered as part of a sustained release formulation (i.e., a formulation such as a capsule, sponge or gel (composed of polysaccharides, for example) that effects a slow release of compound following administration) . Such formulations may generally be prepared using well known technology (see, e.g., Coombes et al., Vaccine 14: 1429-1438 (1996) ) and administered by, for example, oral, rectal or subcutaneous implantation, or by implantation at the desired target site. Sustained-release formulations may contain a polypeptide, polynucleotide or antibody dispersed in a carrier matrix and/or contained within a reservoir surrounded by a rate controlling membrane. In some embodiments, the vaccine is formulated for induction of systemic or localized mucosal immunity through immunogen entrapment and co-administration with microparticles.
Carriers for use within such formulations are biocompatible, and may also be biodegradable; preferably the formulation provides a relatively constant level of active component release. Such carriers include microparticles of poly (lactide-co-glycolide) , polyacrylate, latex, starch, cellulose, dextran and the like. Other delayed-release carriers include supramolecular biovectors, which comprise a non-liquid hydrophilic core (e.g., a cross-linked polysaccharide or oligosaccharide) and, optionally, an external layer comprising an amphiphilic compound, such as a phospholipid (see, e.g., U.S. Patent No. 5,151,254 and PCT applications WO 94/20078, WO/94/23701 and WO 96/06638) . The amount of active compound contained within a sustained release formulation depends upon the site of implantation, the rate and expected duration of release and the nature of the condition to be treated or prevented.
Vaccines and pharmaceutical compositions may be presented in unit-dose or multi-dose containers, such as sealed ampoules or vials. Such containers are preferably hermetically sealed to preserve sterility of the formulation until use. In general, formulations may be stored as suspensions, solutions or emulsions in oily or aqueous vehicles. Alternatively, a vaccine or pharmaceutical composition may be stored in a freeze-dried condition requiring only the addition of a sterile liquid carrier immediately prior to use.
VII. KITS
The disclosure provides a variety of kits for conveniently and/or effectively carrying out methods of the present disclosure. Typically kits will comprise sufficient amounts and/or numbers of components to allow a user to perform multiple treatments of a subject (s) . The disclosure provides kits for eliciting an immune response in a subject according to the method of the present disclosure. The kits typically include a container that contains a pharmaceutical composition having an effective amount of SARS-CoV-derived or SARS-CoV-2-derived T cell epitope polypeptides or a function fragment thereof, and/or B cell epitope polypeptides or a function fragment thereof) , and/or the polypeptides optionally further comprising at least one heterologous polypeptide sequence. The kit may comprise nucleic acid sequences encoding the SARS-CoV-derived or SARS-CoV-2-derived T cell and/or B cell polypeptides, and/or the expression cassettes for expressing the T cell and/or B cell polypeptides, and/or the vectors for expressing the expression cassettes. The kit may optionally comprise an additional container containing a therapeutic agent against SARS-CoV or variants thereof, SARS-CoV-2 or variants thereof, and/or other coronavirus or variants (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) .
In some embodiments, the present disclosure provides kits comprising the SARS-CoV-and/or SARS-CoV-2-derived T cell and/or B cell epitopes (including any polypeptides, proteins or function fragments thereof) of the disclosure. The T cell and/or B cell epitopes may be derived from one or more functional proteins (e.g., S protein, N protein, or M protein) or full length SARS-CoV or SARS-CoV-2. In some embodiments, the kit further comprises polypeptides comprising one or more T cell and/or B cell epitopes derived from one or more of a SARS-CoV and/or SARS-CoV-2 variants. In some embodiments, the kit further comprises polypeptides comprising one or more T cell and/or B cell epitopes derived from one or more coronaviruses or variants thereof (e.g.,  MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) . The kit may further comprise packaging and instructions and/or a delivery agent to form a formulation composition. The delivery agent may comprise a saline, a buffered solution, a lipidoid or any delivery agent disclosed herein.
The kits can be for protein or polypeptide production, comprising a polynucleotide comprising a translatable region of one or more T cell and/or B cell epitopes derived from a functional protein (e.g., S protein, N protein, or M protein) or full length SARS-CoV or SARS-CoV-2. In some embodiments, the kit further comprises a polynucleotide comprising a translatable region of one or more T cell and/or B cell epitopes derived from one or more coronaviruses (e.g., MERS-CoV, HCoV-HKU1, HCoV-229E, HCoV-OC43, and HCoV-NL63) . The kit may further comprise packaging and instructions and/or a delivery agent to form a formulation composition. The delivery agent may comprise a saline, a buffered solution, a lipidoid or any delivery agent disclosed herein.
In some embodiments, the kit further comprises devices for administering the polypeptides or a protein or polypeptide product descried herein. For instance, the kit may comprise syringes, needles, and instructions on how to dispense the pharmaceutical composition, including description of the type of patients who may be treated (e.g., subjects who are at risk of having COVID-19 due to individual health condition, pre-existing condition of any kind, compromised immune system, living condition, age, or occupation; subjects who has close contact with someone who has COVID-19) .
All publications and patent applications cited in this specification are herein incorporated by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.
Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to one of ordinary skill in the art in light of the teachings of this disclosure that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
EXAMPLES
The following examples are provided by way of illustration only and not by way of limitation. Those of skill in the art will readily recognize a variety of non-critical parameters that could be changed or modified to yield essentially similar results.
EXAMPLE 1 –PRELIMINARY IDENTIFICATION OF POTENTIAL VACCINE TARGETS FOR THE COVID-19 CORONAVIRUS (SARS-COV-2) BASED ON SARS-COV IMMUNOLOGICAL STUDIES
Introduction
In this study, we sought to gain insights for vaccine design against SARS-CoV-2 by considering the high genetic similarity between SARS-CoV-2 and SARS-CoV, which caused the outbreak in 2003, and leveraging existing immunological studies of SARS-CoV. By screening the experimentally-determined SARS-CoV-derived B cell and T cell epitopes in the immunogenic structural proteins of SARS-CoV, we identified a set of B cell and T cell epitopes derived from the spike (S) and nucleocapsid (N) proteins that map identically to SARS-CoV-2 proteins. As no mutation has been observed in these identified epitopes among the 120 available SARS-CoV-2 sequences (as of 21 February 2020) , immune targeting of these epitopes may potentially offer protection against this novel virus. For the T cell epitopes, we performed a population coverage analysis of the associated MHC alleles and proposed a set of epitopes that is estimated to provide broad coverage globally, as well as in China. Our findings provide a screened set of epitopes that can help guide experimental efforts towards the development of vaccines against SARS-CoV-2.
Worldwide collaborative efforts from scientists are working on this disease and SARS-CoV-2 to develop effective interventions for controlling and preventing it [6–9] .
Coronaviruses are positive-sense single-stranded RNA viruses belonging to the family Coronaviridae. These viruses mostly infect animals, including birds and mammals. In humans, they generally cause mild respiratory infections, such as those observed in the common cold. However, some recent human coronavirus infections have resulted in lethal endemics, which include the SARS (Severe Acute Respiratory Syndrome) and MERS (Middle East Respiratory Syndrome) endemics. Both of these are caused by zoonotic coronaviruses that belong to the genus Betacoronavirus within Coronaviridae. SARS-CoV originated from Southern China and caused an  endemic in 2003. A total of 8098 cases of SARS were reported globally, including 774 associated deaths, and an estimated case-fatality rate of 14%–15% [10] . The first case of MERS occurred in Saudi Arabia in 2012. Since then, a total of 2, 494 cases of infection have been reported, including 858 associated deaths, and an estimated high case-fatality rate of 34.4% [11] . While no case of SARS-CoV infection has been reported since 2004, MERS-CoV has been around since 2012 and has caused multiple sporadic outbreaks in different countries.
Like SARS-CoV and MERS-CoV, the recent SARS-CoV-2 belongs to the Betacoronavirus genus [12] . It has a genome size of ~30 kilobases which, like other coronaviruses, encodes for multiple structural and non-structural proteins. The structural proteins include the spike (S) protein, the envelope (E) protein, the membrane (M) protein, and the nucleocapsid (N) protein. With SARS-CoV-2 being discovered very recently, there is currently a lack of immunological information available about the virus (e.g., information about immunogenic epitopes eliciting antibody or T cell responses) . Preliminary studies suggest that SARS-CoV-2 is quite similar to SARS-CoV based on the full-length genome phylogenetic analysis [9, 12] , and the putatively similar cell entry mechanism and human cell receptor usage [9, 13, 14] . Due to this apparent similarity between the two viruses, previous research that has provided an understanding of protective immune responses against SARS-CoV may potentially be leveraged to aid vaccine development for SARS-CoV-2.
Various reports related to SARS-CoV suggest a protective role of both humoral and cell-mediated immune responses. For the former case, antibody responses generated against the S protein, the most exposed protein of SARS-CoV, have been shown to protect from infection in mouse models [15–17] . In addition, multiple studies have shown that antibodies generated against the N protein of SARS-CoV, a highly immunogenic and abundantly expressed protein during infection [18] , were particularly prevalent in SARS-CoV-infected patients [19, 20] . While being effective, the antibody response was found to be short-lived in convalescent SARS-CoV patients [21] . In contrast, T cell responses have been shown to provide long-term protection [21–23] , even up to 11 years post-infection [24] , and thus have also attracted interest for a prospective vaccine against SARS-CoV [reviewed in [25] ] . Among all SARS-CoV proteins, T cell responses against the structural proteins have been found to be the most immunogenic in peripheral blood mononuclear cells of convalescent SARS-CoV patients as compared to the non-structural proteins  [26] . Further, of the structural proteins, T cell responses against the S and N proteins have been reported to be the most dominant and long-lasting [27] .
Here, by analyzing available experimentally-determined SARS-CoV-derived B cell epitopes (both linear and discontinuous) and T cell epitopes, we identify and report those that are completely identical and comprise no mutation in the available SARS-CoV-2 sequences (as of 21 February 2020) . These epitopes have the potential, therefore, to elicit a cross-reactive/effective response against SARS-CoV-2. We focused particularly on the epitopes in the S and N structural proteins due to their dominant and long-lasting immune response previously reported against SARS-CoV. For the identified T cell epitopes, we additionally incorporated the information about the associated MHC alleles to provide a list of epitopes that seek to maximize population coverage globally, as well as in China. Our presented results can potentially narrow down the search for potent targets for an effective vaccine against SARS-CoV-2, and help guide experimental studies focused on vaccine development.
Materials and Methods
Acquisition and Processing of Sequence Data. A total of 120 whole genome sequences of SARS-CoV-2 were downloaded on 21 February 2020 from the GISAID database (gisaid. org/CoV2020/) (Table 6) . We excluded sequences that likely had spurious mutations resulting from sequencing errors, as indicated in the comment field of the GISAID data. These nucleotide sequences were aligned to the GenBank reference sequence (accession ID: NC_045512.2) and then translated into amino acid residues according to the coding sequence positions provided along the reference sequence for SARS-CoV-2 proteins (orf1a, orf1b, S, ORF3a, E, M, ORF6, ORF7a, ORF7b, ORF8, N, and ORF10) . These sequences were aligned separately for each protein using the MAFFT multiple sequence alignment program [28] . Reference protein sequences for SARS-CoV and MERS-CoV were obtained following the same procedure from GenBank using the accession IDs NC_004718.3 and NC_019843.3, respectively.
Table 6. Whole genome sequences of SARS-CoV-2 downloaded from theGISAID database.
Figure PCTCN2022070948-appb-000008
Figure PCTCN2022070948-appb-000009
Acquisition and Filtering of Epitope Data. SARS-CoV-derived B cell and T cell epitopes were searched on the NIAID Virus Pathogen Database and Analysis Resource (ViPR) (www. viprbrc. org/; accessed 21 February 2020) [29] by querying for the virus species name: “Severe acute respiratory syndrome-related coronavirus” from “human” hosts. We limited our search to include only the experimentally-determined epitopes that were associated with at least one positive assay: (i) Positive B cell assays (e.g., enzyme-linked immunosorbent assay (ELISA) -based qualitative binding) for B cell epitopes; and (ii) either positive T cell assays (such as enzyme-linked immune absorbent spot (ELISPOT) or intracellular cytokine staining (ICS) IFN-γ release) , or positive major histocompatibility complex (MHC) binding assays for T cell epitopes. Strictly speaking, the latter set of epitopes, determined using positive MHC binding assays, are antigens which are candidate epitopes, since a T cell response has not been confirmed experimentally. However, for brevity and to be consistent with the terminology used in the ViPR database, we will not make this qualification, and will simply refer to them as epitopes in this study. The number of B cell and T cell epitopes obtained from the database following the above procedure is listed in Table 1.
Table 1. Filtering criteria and corresponding number of Severe Acute Respiratory Syndrome Coronavirus (SARS-CoV) -derived epitopes obtained from the Virus Pathogen Database and Analysis Resource (ViPR) database.
Figure PCTCN2022070948-appb-000010
Population-Coverage-Based T Cell Epitope Selection. Population coverages for sets of T cell epitopes were computed using the tool provided by the Immune Epitope Database (IEDB) (tools. iedb. org/population/; accessed 21 February 2020) [30] . This tool uses the distribution of MHC alleles (with at least 4-digit resolution, e.g., A*02: 01) within a defined population (obtained from allelefrequencies. net/) to estimate the population coverage for a set of T cell epitopes. The estimated population coverage represents the percentage of individuals within the population that are likely to elicit an immune response to at least one T cell epitope from the set. To identify the set of epitopes associated with MHC alleles that would maximize the population coverage, we adopted a greedy approach: (i) We first identified the MHC allele with the highest individual population coverage and initialized the set with their associated epitopes, then (ii) we progressively added epitopes associated with other MHC alleles that resulted in the largest increase of the accumulated population coverage. We stopped when no increase in the accumulated population coverage was observed by adding epitopes associated with any of the remaining MHC alleles.
Constructing the Phylogenetic Tree. We used the publicly available software PASTA v1.6.4 [31] to construct a maximum-likelihood phylogenetic tree of each structural protein using the unique set of sequences in the available data of SARS-CoV, MERS-CoV, and SARS-CoV-2. We additionally included the Zaria Bat coronavirus strain (accession ID: HQ166910.1) to serve as an outgroup. The appropriate parameters for tree estimation are automatically selected in the software based on the provided sequence data. For visualizing the constructed phylogenetic trees, we used the publicly available software Dendroscope v3.6.3 [32] . Each constructed tree was rooted with the outgroup Zaria Bat coronavirus strain, and circular phylogram layout was used
Data and Code Availability. All sequence and immunological data, and all scripts (written in R) for reproducing the results are available online [33] .
Results
Structural Proteins of SARS-CoV-2 Are Genetically Similar to SARS-CoV, but Not to MERS-CoV. SARS-CoV-2 has been observed to be close to SARS-CoV-much more so than MERS-CoV-based on full-length genome phylogenetic analysis [9, 12] . We checked whether this is also true at the level of the individual structural proteins (S, E, M, and N) . A straightforward reference-sequence-based comparison indeed confirmed this, showing that the M, N, and E proteins of SARS-CoV-2 and SARS-CoV have over 90%genetic similarity, while that of the S protein was notably reduced (but still high) (FIG. 1A) . The similarity between SARS-CoV-2 and MERS-CoV, on the other hand, was substantially lower for all proteins (FIG. 1A) ; a feature that was also evident from the corresponding phylogenetic trees (FIG. 1B) . We note that while the former analysis (FIG. 1A) was based on the reference sequence of each coronavirus, it is indeed a good representative of the virus population, since few amino acid mutations have been observed in the corresponding sequence data (FIG. 3) . It is also noteworthy that while MERS-CoV is the more recent coronavirus to have infected humans, and is comparatively more recurrent (causing outbreaks in 2012, 2015, and 2018) (who. int/emergencies/mers-cov/en/) , SARS-CoV-2 is closer to SARS-CoV, which has not been observed since 2004.
Given the close genetic similarity between the structural proteins of SARS-CoV and SARS-CoV-2, we attempted to leverage immunological studies of the structural proteins of SARS-CoV to potentially aid vaccine development for SARS-CoV-2. We focused specifically on the S and N proteins as these are known to induce potent and long-lived immune responses in SARS-CoV [15–17, 19, 20, 25, 27] . We used the available SARS-CoV-derived experimentally-determined epitope data (see Materials and Methods) and searched to identify T cell and B cell epitopes that were identical-and hence potentially cross-reactive-across SARS-CoV and SARS-CoV-2. We first report the analysis for T cell epitopes, which have been shown to provide a long-lasting immune response against SARS-CoV [27] , followed by a discussion of B cell epitopes.
Mapping the SARS-CoV-Derived T Cell Epitopes that ARE identical in SARS-CoV-2, and Determining those with Greatest Estimated Population Coverage. The SARS-CoV-derived T cell epitopes used in this study were experimentally-determined from two different  types of assays [29] : (i) Positive T cell assays, which tested for a T cell response against epitopes, and (ii) positive MHC binding assays, which tested for epitope-MHC binding. We aligned these T cell epitopes across the SARS-CoV-2 protein sequences. Among the 115 T cell epitopes that were determined by positive T cell assays (Table 1) , we found that 27 epitope-sequences were identical within SARS-CoV-2 proteins and comprised no mutation in the available SARS-CoV-2 sequences (as of 21 February 2020) (Table 2) . Interestingly, all of these were present in either the N (16) or S (11) protein. MHC binding assays were performed for 19 of these 27 epitopes, and these were reported to be associated with only five distinct MHC alleles (at 4-digit resolution) : HLA-A*02: 01, HLA-B*40: 01, HLA-DRA*01: 01, HLA-DRB1*07: 01, and HLA-DRB1*04: 01. Consequently, the accumulated population coverage of these epitopes (see Materials and Methods for details) is estimated to not be high for the global population (59.76%) , and was quite low for China (32.36%) . For the remaining 8 epitopes, since the associated MHC alleles are unknown, they could not be used in the population coverage computation. Additional MHC binding tests to identify the MHC alleles that bind to these 8 epitopes may reveal additional distinct alleles, beyond the five determined so far, that may help to improve population coverage.
To further expand the search and identify potentially effective T cell targets covering a higher percentage of the population, we next additionally considered the set of T cell epitopes that have been experimentally-determined from positive MHC binding assays (Table 1) , but, unlike the previous epitope set, their ability to induce a T cell response against SARS-CoV was not experimentally determined. Nonetheless, they also present promising candidates for inducing a response against SARS-CoV-2. For the expanded set of epitopes, all of which have at least one positive MHC binding assay, we found that 229 epitope-sequences have an identical match in SARS-CoV-2 proteins and have associated MHC allele information available (listed in Table 7) . Of these 229 epitopes, ~82%were MHC Class I restricted epitopes (Table 8) . Importantly, 102 of the 229 epitopes were derived from either the S (66) or N (36) protein. Mapping all 66 S-derived epitopes onto the resolved crystal structure of the SARS-CoV S protein (FIG. 4) revealed that 3 of these (GYQPYRVVVL, QPYRVVVLSF, and PYRVVVLSF) were located entirely in the SARS-CoV receptor-binding motif (uniprot. org/uniprot/P59594) , known to be important for virus cell entry [34] .
Table 7. List of all SARS-CoV-derived T cell epitopes determined using positive MHC binding assays (with associated MHC allele information available at 4-digit resolution) and found to be identical in SARS-CoV-2.
Figure PCTCN2022070948-appb-000011
Figure PCTCN2022070948-appb-000012
Figure PCTCN2022070948-appb-000013
Figure PCTCN2022070948-appb-000014
Figure PCTCN2022070948-appb-000015
Figure PCTCN2022070948-appb-000016
Figure PCTCN2022070948-appb-000017
Figure PCTCN2022070948-appb-000018
Table 8. Distribution of all SARS-CoV-derived T cell epitopes obtained using positive MHC binding assays (with associated MHC allele information available at 4-digit resolution) and that are identical in SARS-CoV-2.
Protein MHC Allele Class I MHC Allele Class II Total
S
40 26 66
orf1b 45 0 45
N 32 4 36
orf1a 29 0 29
E 1 0 1
M 19 3 22
ORF7a 5 1 6
ORF7b 15 8 23
ORF6 1 0 1
Total 187 42 229
Similar to previous studies on HIV and HCV [35–38] , we estimated population coverages for various combinations of MHC alleles associated with these 102 epitopes. Our aim was to determine sets of epitopes associated with MHC alleles with maximum population coverage, potentially aiding the development of vaccines against SARS-CoV-2. For selection, we adopted a greedy computational approach (see Materials and Methods) , which identified a set of T cell epitopes estimated to maximize global population coverage. This set comprised of multiple T cell epitopes associated with 20 distinct MHC alleles and was estimated to provide an accumulated population coverage of 96.29% (Table 3) . Interestingly, the majority of the T cell epitopes for which a positive immune response has been determined using T cell assays (Table 2) were presented by the globally most-prevalent MHC allele (shown in bold text in Table 3) . Moreover, the functionally important epitopes located in the SARS-CoV receptor binding motif were associated with the second and third most-prevalent MHC alleles (underlined in Table 3) . Thus, while the ordering of T cell epitopes in Table 3 is based on the estimated global population coverage of the associated MHC alleles, it is also a natural order in which these epitopes should be tested experimentally for determining their potential to induce a positive immune response against SARS-CoV-2. We also computed the population coverage of this specific set of epitopes in China, the country most affected by the COVID-19 outbreak, which was estimated to be slightly lower (88.11%) , as certain MHC alleles (e.g., HLA-A*02: 01) associated with some of these epitopes are less frequent in the Chinese population (Table 3) . Repeating the same greedy approach but focusing on the Chinese population, instead of a global population, the maximum population coverage was estimated to be 92.76% (Table 9) .
Table 2. SARS-CoV-derived T cell epitopes obtained using positive T cell assays that are identical in SARS-CoV-2 (27 epitopes in total) .
Figure PCTCN2022070948-appb-000019
1NA: Not available.
Table 9. Set of SARS-CoV-derived S and N protein T cell epitopes (obtained using positive MHC binding assays) that are identical in SARS-CoV-2 and that maximize estimated population coverage in China (86 distinct epitopes) .
Figure PCTCN2022070948-appb-000020
Figure PCTCN2022070948-appb-000021
Figure PCTCN2022070948-appb-000022
Due to the promiscuous nature of binding between peptides and MHC alleles, multiple S and N peptides were reported to bind to individual MHC alleles. Thus, while we list all the S and N epitopes that bind to each MHC allele (Table 3) , the estimated maximum population coverage may be achieved by selecting at least one epitope for each listed MHC allele. Likewise, many individual S and N epitopes were found to be presented by multiple alleles and thereby estimated to have varying global population coverage (listed in Table 10) .
Table 3. Set of the SARS-CoV-derived spike (S) and nucleocapsid (N) protein T cell epitopes (obtained from positive MHC binding assays) that are identical in SARS-CoV-2 and that maximize estimated population coverage globally (87 distinct epitopes) .
Figure PCTCN2022070948-appb-000023
Figure PCTCN2022070948-appb-000024
Figure PCTCN2022070948-appb-000025
1 Multiple SARS-CoV-derived epitopes that were determined using MHC binding assays are shown for each allele. Epitopes that were also tested for positive T cell response (listed also in Table 2) are shown in bold text. Epitopes that lie within the SARS-CoV receptor-binding motif are underlined.  2 Epitopes are ordered according to the estimated global accumulated population coverage.
Table 10. Estimated global and Chinese population coverages for the individual SARS-CoV-derived S or N protein T cell epitopes (obtained using positive MHC binding assays) that are identical in SARS-CoV-2.
Figure PCTCN2022070948-appb-000026
Figure PCTCN2022070948-appb-000027
Figure PCTCN2022070948-appb-000028
Figure PCTCN2022070948-appb-000029
Mapping the SARS-CoV-Derived B cell Epitopes that Are Identical in SARS-CoV-2 . Similar to T cell epitopes, we used in our study the SARS-CoV-derived B cell epitopes that have been experimentally-determined from positive B cell assays [29] . These epitopes were classified as: (i) Linear B cell epitopes (antigenic peptides) , and (ii) discontinuous B cell epitopes (conformational epitopes with resolved structural determinants) .
We aligned the 298 linear B cell epitopes (Table 1) across the SARS-CoV-2 proteins and found that 49 epitope-sequences, all derived from structural proteins, have an identical match and comprised no mutation in the available SARS-CoV-2 protein sequences (as of 21 February 2020) .  Interestingly, a large number (45) of these were derived from either the S (23) or N (22) protein (Table 4) , while the remaining (4) were from the M protein (Table 11) .
Table 4. SARS-CoV-derived linear B cell epitopes from S (23; 20 of which are located in subunit S2) and N (22) proteins that are identical in SARS-CoV-2 (45 epitopes in total) .
Figure PCTCN2022070948-appb-000030
Table 11. SARS-CoV-derived linear B cell epitopes, excluding those in S and N proteins, that are identical in SARS-CoV-2.
Protein IEDB ID Epitope
M 21996 GRCDIKDLPKEITVATSR
M 29127 ITVATSRT
M 48052 PKEITVATSRTLSYYKL
M 66409 TSRTLSYYKLGASQRV
On the other hand, all 6 SARS-CoV-derived discontinuous B cell epitopes obtained from the ViPR database (Table 5) were derived from the S protein. Based on the pairwise alignment  between the SARS-CoV and SARS-CoV-2 reference sequences (FIG. 5) , we found that none of these mapped identically to the SARS-CoV-2 S protein, in contrast to the linear epitopes. For 3 of these discontinuous B cell epitopes (corresponding to antibodies S230, m396, and 80R [39–41] ) , there was a partial mapping, with at least one site having an identical residue at the corresponding site in the SARS-CoV-2 S protein (Table 5) .
Table 5. SARS-CoV-derived discontinuous B cell epitopes (and associated known antibodies [39–41] ) that have at least one site with an identical amino acid to the corresponding site in SARS-CoV-2.
Figure PCTCN2022070948-appb-000031
Mapping the residues of the linear and discontinuous B cell epitopes onto the available structure of the SARS-CoV S protein revealed their distinct association with the two functional subunits of the S protein [42] : S1, important for interaction with the host cell receptor, and S2, involved in fusion of the cellular and virus membranes (FIG. 2A) . Specifically, 20 of the 23 linear epitopes (Table 4) mapped to S2 (FIG. 2B) . Thus, the antibodies targeting the identified linear epitopes in the S2 subunit might cross-react and neutralize both SARS-CoV and SARS-CoV-2, as suggested in a very recent study [43] . While S2 is comparatively less exposed than S1, it may be accessible to antibodies during the complex conformational changes involved in viral entry of coronaviruses [44–46] ; though this remains to be more clearly understood. In contrast, the 3 discontinuous B cell epitopes (Table 5) mapped onto the more exposed S1 subunit (FIG. 2C, left panel) , which contains the receptor-binding motif of the SARS-CoV S protein [34] . We observed that very few residues of the 3 discontinuous epitopes were identical within SARS-CoV and SARS-CoV-2 (Figure 2c, right panel) . These differences suggest that the SARS-CoV-specific antibodies S230, m396, and 80R known to bind to these epitopes in SARS-CoV might not be able to bind to the same regions in SARS-CoV-2 S protein. Interestingly, while this paper was under review, this has been confirmed experimentally [47] . Further studies are currently under way to  identify other SARS-CoV antibodies that may bind to discontinuous epitopes of the SARS-CoV-2 S protein [48] .
Discussion
The quest for a vaccine against the novel SARS-CoV-2 is recognized as an urgent problem. Effective vaccination could indeed play a significant role in curbing the spread of the virus, and help to eliminate it from the human population. However, scientific efforts to address this challenge are only just beginning. Much remains to be learnt about the virus, its biological properties, epidemiology, etc. At this early stage, there is also a lack of information about specific immune responses against SARS-CoV-2, which presents a challenge for vaccine development.
This study has sought to assist with the initial phase of vaccine development by providing recommendations of epitopes that may potentially be considered for incorporation in vaccine designs. Despite having limited understanding of how the human immune system responds naturally to SARS-CoV-2, these epitopes are motivated by responses they have recorded in SARS-CoV (or, for the case of T cell epitopes, to at least confer MHC binding) , and the fact that they map identically to SARS-CoV-2, based on the available sequence data (as of 21 February 2020) . This important observation should not be taken for granted. Despite the apparent similarity between SARS-CoV and SARS-CoV-2, there is still considerable genetic variation between the two, and it is not obvious a-prior if epitopes that elicit an immune response against SARS-CoV are likely to be effective against SARS-CoV-2. We found that only 23%and 16%of known SARS-CoV T cell and B cell epitopes map identically to SARS-CoV-2, respectively, and with no mutation having been observed in these epitopes among the available SARS-CoV-2 sequences (as of 21 February 2020) . This provides a strong indication of their potential for eliciting a robust T cell or antibody response in SARS-CoV-2.
EXAMPLE 2 –OPTIMIZED PEPTIDE POOLS FOR ASSESSING REGION-SPECIFIC SARS-COV2 CD8+ T CELL RESPONSES AND USES THEREOF
Introduction
Global efforts to combat COVID-19 have led to the rapid development of multiple vaccines. These vaccines have been shown to induce a robust neutralizing antibody response and provide protection against severe disease and hospitalization (Hall et al. 2021; Mor et al. 2021) .  As the virus continues to circulate worldwide, virus variants have emerged in several regions, concerns about their potential to escape vaccine-induced antibody responses. Preliminary results suggest that most current vaccines remain effective against emerging virus variants (Abdool Karim and de Oliveira 2021; Abu-Raddad et al. 2021; Collier et al. 2021; Emary et al. 2021; Liu et al. 2021; Planas et al. 2021) . However, specific variants such as Beta (B. 1.351) , Gamma (P. 1) and Delta (B1.617.2) , that first emerged in South Africa, Brazil, and India respectively, are currently under investigation due to the observed reduction in neutralizing antibody titres against them in sera of vaccinated individuals (Abdool Karim and de Oliveira 2021; Liu et al. 2021; Madhi et al. 2021; Planas et al. 2021; Wall et al. 2021; Zhou et al. 2021) .
In addition to eliciting neutralizing antibodies, SARS-CoV-2 vaccines in use or development also stimulate T cell responses. There is increasing evidence of the role of T cells in protection from severe disease in SARS-CoV-2 infected patients (Altmann and Boyton 2020; Chen and John Wherry 2020; Liao et al. 2020; Mazzoni et al. 2020; Reynolds et al. 2020; Rydyznski Moderbacher et al. 2020; Wyllie et al. 2020; Bergamaschi et al. 2021; Bertoletti et al. 2021; Cohen et al. 2021) , and of their robustness to mutations associated with SARS-CoV-2 variants (Quadeer et al. 2021; Tarke et al. 2021; Woldemeskel et al. 2021) . However, in contrast to neutralizing antibody responses (Piccoli et al. 2020; Pinto et al. 2020; Fedry et al. 2021) , specific T cell responses have not been characterized in detail for any COVID-19 vaccine thus far.
For most currently administered SARS-CoV-2 vaccines, T cell responses have been coarsely measured using immune assays that stimulate blood samples of vaccinated individuals using overlapping peptide pools (Anderson et al. 2020; Folegatti et al. 2020; Keech et al. 2020; Logunov et al. 2020; Ramasamy et al. 2020; Sahin et al. 2020; Zhu et al. 2020; Ella et al. 2021; Klasse et al. 2021; Sadoff et al. 2021) . Use of these pools, however, may underestimate T cell responses due to peptide competition, where immunogenic peptides compete with a large number of irrelevant peptides in the pool that are not recognized by T cells (Pala et al. 1988; Sahin et al. 2021) . In contrast, assays based on optimized peptide pools would be more efficient at estimating T cell responses as these comprise of a selected set of most relevant peptides against which a T cell response is expected to be stimulated. Such pools also enable identifying precise T cell epitopes in the context of the associated human leukocyte antigen (HLA) alleles presenting them.
Designing optimized peptides pools for assessing T cell responses is challenging due to the diversity of these responses. This is because T cells recognize peptides restricted by an individual’s HLA alleles which are highly diverse across the global population (albeit with some commonalities in a given region) . Consequently, the peptides restricted by these HLA alleles are also different, and hence T cell responses are expected to differ between geographical regions, even for the same vaccine. Moreover, current vaccines employ different SARS-CoV-2 antigens (e.g., based on the spike (S) protein only or employing the whole inactivated virion) , and these are expected to elicit distinct T cell responses. To our knowledge, no tool or platform is currently available that provides peptide pools for measuring SARS-CoV-2-specific T cell responses in a particular region where population is immunized by a specific COVID-19 vaccine. Here, we fill this important gap by developing a software platform, SARS2TPools, that provides optimized peptide pools for assessing region-specific vaccine-induced SARS-CoV-2 T cell responses. These pools are designed by exploiting information of prevalent HLA alleles in a population, the experimentally-determined and in-silico-predicted SARS-CoV-2 T cell epitopes associated with these alleles, and the antigen employed in the vaccine. The optimized pools provided by SARS2TPools, in addition to characterizing the vaccine-induced T cell responses in detail, can be useful for designing T cell based diagnostics, monitoring durability of T cell response, and any change in T cell responses due to emerging SARS-CoV-2 variants.
Methods and Materials
Data Collection. We downloaded experimentally-determined HLA class I and class II restricted SARS-CoV-2 T cell epitope data (CD8 + and CD4 +, respectively) from the immune epitope database (IEDB) (Vita et al. 2019) on March 10, 2021. We included all epitopes that were reported in positive T cell assays with associated HLA information available. The data consisted of 768 and 445 unique class I and class II epitope-HLA pairs, respectively. Majority of the HLA class I restricted epitopes (474/768) were nine residues long, which is the canonical length of epitopes restricted by HLA class I alleles. The epitope data was found to be biased towards a handful of HLA alleles, with only 10 HLA class I alleles (HLA-A*02: 01, HLA-A*03: 01, HLA-A*11: 01, HLA-A*24: 02, HLA-A*29: 02, HLA-A*68: 01, HLA-B*07: 02, HLA-B*35: 01, HLA-B*51: 01, HLA-B*57: 01) having 20 or more nine-residue-long epitopes. Collectively, the epitopes restricted by these 10 HLA alleles corresponded to ~62% (295/474) of nine-residue-long epitopes  in the data. In the case of HLA class II restricted epitopes, all the available epitopes were 15 resides long, and only 3 HLA alleles had more than 20 epitopes in the data.
In Silico Prediction Methods. Performance of several in silico epitope prediction methods were benchmarked against the set of experimentally-determined SARS-CoV-2 epitopes associated with the 10 HLA class I alleles having the most data. The considered methods included the current state-of-the-art methods such as MHCflurry (O’ Donnell et al. 2020) , NetMHCpan4.1 (Reynisson et al. 2020) , HLAthena (Sarkizova et al. 2020) , NetMHCpan4.0 (Jurtz et al. 2017) , NetMHC4.0 (Andreatta and Nielsen 2016) , along with other common prediction methods that have been employed for predicting SARS-CoV-2 epitopes (Sohail et al. 2021) such as NetMHCpan3.0 (Nielsen and Andreatta 2016) , SMM (Peters and Sette 2005) , SMMPMBEC (Kim et al. 2009) , and IEDB consensus (Moutaftsi et al. 2006) . We considered the eluted ligand and binding affinity predictions (denoted by suffix BA and EL respectively) of NetMHCpan4.1 and NetMHCpan4.0 as separate methods, as was done for the latter method in (Sarkizova et al. 2020) and (Paul et al. 2020) . Similarly, we considered the binding affinity and the presentation score predictions of MHCflurry as two separate methods, referred to as MHCflurry2.0BA and MHCflurry2.0P. In cases where a method required an input other than the protein sequence, HLA allele, and length of the predicted peptides, we used the default parameter settings for that method.
Union Approach. In this work, we have proposed a union approach based on combining the top-ranked predictions of MHCflurry2.0P and NetMHCpan4.1BA to obtain a set of peptides restricted by a given HLA. This approach was motivated by performance comparison analysis of the 12 in silico epitope prediction methods listed above. Briefly, we ranked peptides in ascending order of their predicted score using each method and compared the histograms of ranks of experimentally-determined SARS-CoV-2 CD8 + T cell epitopes associated with the 10 HLA class I alleles with the most data. We found that these histograms were bi-modal for all 12 methods. That is, while the top predictions of each method contained a large number of experimentally-determined epitopes, a good number of epitopes were also ranked quite low by each method (FIG. 10) . Exploring the relationships among the set of top 20 ranked peptides per HLA allele predicted by these methods revealed that the predictions of MHCflurry2.0P were most distinct from those of other methods (FIGS. 11A-D) . Consistent results were obtained when this set was constructed by pooling the top 10 to top 25 ranked peptides restricted by each HLA allele (FIGS. 11A-D) .  Predictions of MHCflurry2.0P also contained a large number of experimentally-determined SARS-CoV-2 epitopes that were not present in the set of top-ranked peptides predicted by any other method (FIG. 7A) . Given the uniqueness of the predictions of MHCflurry2.0P, we asked if a strategy that combines the predictions of MHCflurry2.0P with any of the other 11 methods would work better than any individual method. The union strategy combines the top x predictions of any two methods and provides a set of peptides whose size can vary between x and 2x depending on the number of common peptides predicted by each method. We fixed one of the methods as MHCflurry2.0P and predicted 11 peptide pools by combining predictions of MHCflurry2.0P with those of the other 11 methods. Our analysis showed that the pool predicted by the union approach always had a higher hit-rate (the fraction of experimentally-determined SARS-CoV-2 epitopes present in the set of top-ranked peptides) than those predicted by the individual methods (FIG. 12) . While comparison among the various union approaches did not readily reveal a clear winner, the unions of MHCflurry2.0P with the in silico methods NetMHC4.0, NetMHCpan4.0BA, NetMHCpan4.1BA, and NetMHCpan4.1EL ranked among the top (FIG. 13, Table 13) .
The union method implemented in SARS2TPools combines the predictions of MHCflurry2.0P with NetMHCpan4.1BA. SARS2TPools provides optimized peptide pools by supplementing experimentally-determined epitopes with small or large sized group of in silico predicted epitopes corresponding respectively to top 10 and top 20 predictions of each method being combined. We used the default thresholds of MHCflurry2.0P and NetMHCpan4.1BA to assess whether or not a peptide is predicted to be an epitope. However, the platform also provides a relaxed threshold which can be particularly useful for specific proteins with very limited number of predicted epitopes.
Statistical Analysis. Statistical analyses were performed using the R language (version 3.6) on the RStudio server (version 1.3) . The software platform was developed using the open source R Shiny (version 1.5) development framework.
Results
SARS2TPools provides optimized peptide pools for assessing T cell responses leveraging experimentally-determined SARS-CoV-2 epitope data. The designed CD8+ T cell peptide pools also include in silico predictions obtained using a computational approach optimized for predicting SARS-CoV-2 CD8 + T cell epitopes. This approach is based on benchmarking predictions of state- of-the-art in silico methods against the ample experimentally-determined SARS-CoV-2 CD8+epitope data that is now available (Methods) .
Global HLA class I diversity and summary of experimentally-determined SARS-CoV-2 CD8+ T cell epitope data. Each individual possesses three major HLA class I alleles, HLA-A, HLA-B, and HLA-C, which are among the most polymorphic loci of the human genome (Jin et al. 2018) . In fact, more than 6, 500 alleles for each of these three loci have been identified so far (Robinson et al. 2015) . In order to design a peptide pool for assessing T cell response in a specific geographical region, information of the set of most common HLA alleles in that population is required. Based on extensive population studies, the allele frequency net database (AFND) has curated a list of ten most common HLA alleles per locus for 11 distinct geographical regions encompassing the global population (Gonzalez-Galarza et al. 2019) . These sets of HLA alleles have a population coverage of 96%or more in the respective regions. By combining alleles prevalent in all regions, we obtained a total of 136 distinct HLA class I alleles (FIG. 6A, Table 12) . Ideally, a peptide pool for assessing T cell responses in a specific region should include a set of experimentally-determined epitopes against all HLA alleles prevalent in that region. However, the available experimental data of SARS-CoV-2 CD8 + T cell epitopes with associated HLA class I information is limited (see Methods) . Comparing HLA class I alleles for which experimental data is available with the compiled list of 30 most prevalent HLA alleles per region suggested that a peptide pool comprising of only experimentally-determined epitopes does not cover more than half of the alleles prevalent in majority of the regions (FIG. 6B) . This is because epitopes associated with several of the prevalent HLA alleles in different regions have not been experimentally determined yet. Further analysis showed that even at the level of individual HLA alleles, there is large disparity in the number of experimentally reported epitopes, particularly for proteins other than S (FIG. 6C) . These data limitations are addressed in the developed platform SARS2TPools by supplementing experimentally-determined SARS-CoV-2 epitope data with in silico predictions to design peptide pools optimized for assessing vaccine-induced T cell responses in a specific region.
Table 12. List of 136 HLA alleles that are ranked in the top 10 most-prevalent HLA alleles in at least one of the 13 geographical regions defined by AFND.
Figure PCTCN2022070948-appb-000032
Figure PCTCN2022070948-appb-000033
In silico prediction of CD8+ T cell epitopes leveraging SARS-CoV-2 experimental epitope data. We have developed an in silico strategy optimized to predict CD8 + T cell epitopes by leveraging experimentally-determined SARS-CoV-2 immunological data. Briefly, we used the available information of experimentally-determined SARS-CoV-2 immunological data. Briefly, we used the available information of experimentally-determined SARS-CoV-2 CD8 + T cell epitopes as ground-truth data for comparing the sets of top-ranked peptides predicted by 12 in silico HLA class I epitope prediction methods (Methods) . Using the number of experimentally-determined epitopes present in the set of top ranked peptides as a metric to quantify the ability of a method to predict SARS-CoV-2 epitopes, we found that the NetMHCpan family of methods outperformed others (FIG. 7A) . Importantly, this comparison revealed the uniqueness of  MHCflurry2.0P (O’Donnell et al. 2020) predictions; i.e., it predicted the most number of SARS-CoV-2 epitopes that were not predicted by any other method (FIG. 7A) . Our analysis showed that the optimum SARS-CoV-2 epitope prediction strategy was one that combined the predictions of MHCflurry2.0P and NetMHCpan4.1BA (Reynisson et al. 2020) (see Methods for details) . This union strategy performed better than either of the individual methods (FIG. 7B) and was thus used to obtain a set of in silico predicted epitopes to supplement the experimentally-determined epitopes in the optimized peptide pools provided by SARS2TPools. We note that a similar approach to predict CD4 + T cell epitopes could not be pursued at present due to the scarcity of experimentally-determined SARS-CoV-2 CD4 + epitopes having the information of cognate HLA allele (Methods) .
Optimized peptide pools from SARSTPools –Software platform. The platform SARS2TPools integrates experimentally-determined SARS-CoV-2 epitope data, in silico predictions, and information of prevalent HLA alleles across regions to provide optimized peptide pools for assessing vaccine-induced SARS-CoV-2 T cell responses (FIG. 8A) . It enables users to obtain region-specific, host-specific, and protein-specific optimized peptide pools through a simple and user-friendly interface (FIG. 8B) . Peptide pools optimized for each of the 11 regions, by taking into account the information of HLA alleles prevalent in these regions (FIG. 6A) , can be obtained by selecting the ‘Region-specific’ tab on SARS2TPools interface (FIG. 8B) . This optimization is important due to the heterogeneity of prevalent HLA alleles among regions that may result in presentation of different epitopes and consequently different T cell responses (FIG. 6A) . Region-specific pools can be useful to contrast T cell responses induced by the same vaccine in different geographical regions, and to understand the role of population heterogeneity in mediating different disease outcomes.
Host-specific (or cohort-specific) peptide pools, for a range of peptide lengths (8–11 residues) , optimized for HLA haplotypes of the host (or cohort) can be obtained by selecting the ‘Host-specific’ tab on SARS2TPools. This optimization can be useful for cases when host (or cohort) HLA typing has been performed.
SARS2TPools also provides the flexibility to select protein-specific pools derived from the entire SARS-CoV-2 proteome or any number of specific proteins for both region-specific and host-specific options. Moreover, the user can also select peptides belonging to a specific domain of a protein (e.g., receptor binding domain of the spike protein) by specifying a range of amino- acid positions. Protein-specific peptide pools are important for assessing T cell responses induced by vaccine comprising of different antigens, e.g., S only, S with other proteins, or whole-virion.
The CD8 + T cell peptide pools provided by SARS2TPools comprise of both experimentally-determined and in silico predicted epitopes, while those for measuring CD4 + T cell responses comprise only of experimentally-determined epitopes. In the former case, the platform indicates whether a peptide is an experimentally-determined epitope, predicted epitope, or both. It also provides users the flexibility to select peptide pools consisting exclusively of experimentally-determined CD8 + T cell epitopes. SARS2TPools ranks peptides within an optimized pool based on HLA promiscuity, which can be used as an additional prioritization criterion among peptides. In summary, SARS2TPools allows users to obtain pools for assessing T cell responses by flexibly selecting various optimization criteria.
Region-specific CD8+ T cell pools from the whole SARS-CoV-2 proteome. As an illustrative example, we used SARS2TPools to obtain region-specific optimized pools, comprising of peptides derived from the whole SARS-CoV-2 proteome, for measuring CD8 + T cell responses for each of the 11 regions defined by AFND. These pools were designed to include around 20 peptides corresponding to each HLA allele prevalent in a region as follows: (i) If an HLA allele prevalent in a region had 20 or more associated experimentally-determined SARS-CoV-2 epitopes, we selected 20 of these based on response frequency (proportion of responding donors (Quadeer et al. 2020) ) ; (ii) if an HLA allele had less than 20 associated experimentally-determined epitopes, we complemented them with in silico predictions based on the union approach; and (iii) if an HLA allele had no associated experimentally-determined epitopes, all peptides associated with it were predicted based on the union approach.
We found that the number of peptides in each of the region-specific optimized pools was roughly similar, with the fraction of experimentally-determined epitopes in these pools varying between ~35%for Oceania to ~78%for Europe (FIG. 9A) . The HLA alleles covered by each optimized pool can be grouped into 3 classes based on whether their associated peptides in the pool are all (i) experimentally-determined, (ii) in silico predicted, or (iii) a mix of both (FIG. 9B) . Importantly, in silico predictions help to fill the gap for alleles for which limited or no experimentally-determined epitopes are available at present. Comparing the 11 region-specific optimized pools, we found that these pools are largely distinct from each other (FIG. 9C) . These  differences in region-specific pools underscore further the importance of designing optimized peptide pools for assessing T cell responses in different regions.
Listing of all region-specific optimized pools. In addition to providing pools for the whole SARS-CoV-2 proteome, SARS2TPools provides optimized peptide pools for assessing CD8 + T cell responses against any of the 30 individual proteins of SARS-CoV-2 for each of the 11 regions. Thus, a total of 31 x 11 = 341 region-specific optimized peptide pools are provided by the platform. We list all the 3, 860 peptides belonging to any of the 341 pools in Table 15, while the peptides belonging to each specific pools are indicated Table 16 (whole proteome) and Tables 17-47 (individual proteins) .
Table 15. List of all peptide sequences (with their IDs) for CD8 T cell pools.
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1 STECSNLLLQY 51 AIVSTIQRK 101 VYSSANNCTF
2 FADDLNQLTGY 52 YMSALNHTK 102 VYRGTTTYKL
3 VTVKNGSIHLY 53 RIAGHHLGR 103 AYILFTRFF
4 SSANNCTFEY 54 QLRARSVSPK 104 EYHDVRVVL
5 FSAVGNICY 55 SVYAWNRKR 105 IFFITGNTL
6 VVDYGARFY 56 TLADAGFIK 106 DVFYKENSY
7 YKIEELFYSY 57 GVYYHKNNK 107 NTVKSVGKF
8 SSEAFLIGCNY 58 AIVSTIQRKYK 108 DTFCAGSTF
9 LADAGFIKQY 59 ALDPLSETK 109 FTISVTTEI
10 TDEMIAQY 60 RASANLAATK 110 YVNTFSSTF
11 SSPDDQIGYYR 61 IQITISSFK 111 STAALGVLM
12 AGDSGFAAY 62 RMYIFFASFY 112 DSAEVAVKM
13 TSEDMLNPNY 63 TSFGPLVRK 113 DTIANYAKPF
14 YTELEPPCRF 64 KTFPPTEPK 114 DVVAIDYKHY
15 AAISDYDYY 65 VTNNTFTLK 115 ETIQITISSF
16 FTCASEYTGNY 66 KLFDRYFKY 116 EVARDLSLQF
17 LGDVRETMSY 67 ATVVIGTSK 117 EVAVKMFDAY
18 TIEVNSFSGY 68 FAVSKGFFK 118 EVGHTDLMAAY
19 TITQMNLKY 69 AISDYDYYR 119 NSTNVTIATY
20 TSSGDATTAY 70 VVSTGYHFR 120 SVPWDTIANY
21 WLDMVDTSL 71 YIATNGPLK 121 TVKNGSIHLY
22 FLYENAFLP 72 KVAGFAKFLK 122 ELIRQGTDY
23 FLPGVYSV 73 QTVKPGNFNK 123 EVTPSGTWLTY
24 YLITPVHV 74 KSAAEASKK 124 ETKCTLKSF
25 YLTNDVSFLA 75 AGFSLWVYK 125 WTFGAGAAL
26 AQFAPSASA 76 VVNARLRAK 126 TLKEILVTY
27 KLDDKDPNF 77 ASMPTTIAK 127 YIFFASFYY
28 KLNDLCFTNV 78 STFNVPMEK 128 HSYFTSDYY
29 FLAFVVFL 79 GTHWFVTQR 129 LEAPFLYLY
30 YLGTGPEAGL 80 SASKIITLK 130 SFYYVWKSY
31 NTASWFTAL 81 KTIQPRVEK 131 AGLEAPFLY
32 VLQLPQGTTL 82 SAFAMMFVK 132 VGGNYNYLY
33 ALWEIQQV 83 GVYFASTEK 133 KFCLEASFNY
34 SLIYSTAAL 84 TISLAGSYK 134 WFVTQRNFY
35 TLMNVLTLV 85 QYIKWPWYIW 135 VLKGVKLHY
36 KLKDCVMYA 86 LYDKLVSSF 136 GAAAYYVGY
37 WLLWPVTLA 87 SYATHSDKF 137 KVGGNYNYLY
38 TVYSHLLLV 88 VYDPLQPELDSF 138 SFKEELDKY
39 VLSEARQHL 89 YYVGYLQPRTF 139 SWMESEFRVY
40 GLEAPFLYL 90 IYLYLTFYL 140 LVAEWFLAY
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
41 FVVPGLPGT 91 YYKKDNSYF 141 VFAQVKQIY
42 FVENPDILRV 92 LYLYALVYF 142 CVADYSVLY
43 YQDVNCTEV 93 IYQTSNFRV 143 WTAGAAAYY
44 FLAHIQWMV 94 MFTPLVPFW 144 FVFKNIDGY
45 VVFLHVTYV 95 SFLPGVYSV 145 VASQSIIAY
46 ASFDNFKFV 96 AYILFTRF 146 RIKASMPTT
47 VMYMGTLSY 97 AYVDNSSLTI 147 RLFARTRSM
48 QVVNVVTTK 98 AYVNTFSSTF 148 KASMPTTIA
49 VVTTKIALK 99 KMFDAYVNTF 149 MSALNHTKK
50 KSAGFPFNK 100 GNYNYLYRLF 150 TTIAKNTVK
151 YSRYRIGNYK 201 LTAVVIPTK 251 VLAAECTIF
152 RNRFLYIIK 202 ESKPSVEQR 252 KLVSSFLEM
153 GTRNPANNA 203 DTVIEVQGYK 253 VVYRAFDIY
154 AGLPYGANK 204 SSSDNIALL 254 ITILDGISQY
155 RVYSTGSNV 205 FVLAAVYRI 255 LVKQGDDYVY
156 VTYVPAQEK 206 VPHISRQRL 256 QLYLGGMSYY
157 RYRIGNYKL 207 RARSVSPKL 257 VLTESNKKF
158 ASRELKVTF 208 FPLKLRGTA 258 SQRVAGDSGF
159 KVFRSSVLH 209 LPKEITVAT 259 LVQMAPISAM
160 ATSRTLSYYK 210 LPFAMGIIAM 260 FVVEVVDKY
161 QLTPTWRVY 211 NPIQLSSYSL 261 YLKLTDNVY
162 VTPSGTWLTY 212 RPLLESEL 262 KQFDTYNLW
163 LAYYFMRFR 213 KPFERDISTEI 263 LLNKHIDAY
164 GAMDTTSYR 214 LPNNTASWF 264 KIEELFYSY
165 WVLNNDYYR 215 KPVETSNSFDVL 265 VVQQLPETY
166 WFFSNYLKR 216 RPQGLPNNTA 266 QRNAPRITF
167 GSVAYESLR 217 VPGLPGTIL 267 LVSDIDITF
168 KSNLKPFER 218 QPYRVVVLSF 268 VVVNAANVY
169 YNYLYRLFR 219 VPLHGTIL 269 VPFWITIAY
170 QTNSPRRAR 220 RIRGGDGKM 270 TPSKLIEY
171 KFLPFQQFGR 221 APHGVVFLHV 271 QIPFAMQMAY
172 RFASVYAWNR 222 QPGQTFSVL 272 TPSGTWLTY
173 LSYFIASFR 223 LEIPRRNVATL 273 SPDDQIGYY
174 RLFRKSNLK 224 APHGVVFL 274 DVLLPLTQY
175 STGSNVFQTR 225 SPIFLIVAA 275 VAAGLEAPF
176 KLMGHFAWW 226 SAMVRMYIF 276 LPAADLDDF
177 YVMHANYIF 227 TPKYKFVRI 277 LPLTQYNRY
178 KLINIIIWF 228 TLKKRWQLA 278 SIIQFPNTY
179 KVAGFAKFL 229 HLRIAGHHL 279 DASGKPVPY
180 AMYTPHTVL 230 SLYVNKHAF 280 FAPSASAFF
181 SLDNVLSTF 231 HLKDGTCGL 281 TNVLEGSVAY
182 GVVFLHVTY 232 TFKVSIWNL 282 FAMQMAYRF
183 STNVTIATY 233 LTIKKPNEL 283 LGAENSVAY
184 HVTFFIYNK 234 NLKTLLSL 284 NATRFASVY
185 MASLVLARK 235 MLRIMASL 285 LPPLLTDEM
186 FTIGTVTLK 236 FRLFARTRSM 286 IPFAMQMAY
187 AVILRGHLR 237 HPLADNKFAL 287 SEIIGYKAI
188 HVSGTNGTK 238 INITRFQTL 288 FGEYSHVVAF
189 AAISDYDYYR 239 KIYSKHTPI 289 YENFNQHEV
190 FVVSTGYHFR 240 ANRNRFLYI 290 LEMELTPVV
191 YAISAKNRAR 241 NITRFQTL 291 VEVQPQLEM
192 QIAPGQTGK 242 MFDAYVNTF 292 SELLTPLGI
193 NSASFSTFK 243 LLADKFPVL 293 FELDERIDKVL
194 LVIGAVILR 244 LIIMRTFKV 294 LEFGATSAAL
195 FASVYAWNR 245 VPQEHYVRI 295 MELTPVVQTI
196 SVLNDILSR 246 SSAKSASVY 296 SEDAQGMDNL
197 NASVVNIQK 247 VVAIDYKHY 297 FDEDDSEPVL
198 GTITVEELK 248 RLYYDSMSY 298 FEYVSQPFLM
199 FVIRGDEVR 249 YLFDESGEF 299 SEPVLKGVKL
200 DSGFAAYSR 250 EIKESVQTF 300 WEPEFYEAM
301 LEYHDVRVVL 351 YLPYPDPSRI 401 YAKPFLNKV
302 GETLPTEVL 352 IPYNSVTSSIVI 402 VRQALLKTV
303 TEVVGDIIL 353 LPFGWLIV 403 ERHSLSHFV
304 FERDISTEI 354 APYIVGDVV 404 SAKNRARTV
305 AEVQIDRL 355 VPMEKLKTL 405 MYKGLPWNV
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
306 NESLIDLQEL 356 EGYLNSTNV 406 YQPYRVVVL
307 FEPSTQYEY 357 EAKKVKPTV 407 YRYNLPTMC
308 IEVNSFSGY 358 MPTTIAKNTV 408 LHKPIVWHV
309 VELGTEVNEF 359 MPYFFTLL 409 VKNGSIHLY
310 QELIRQGTDY 360 YPQVNGLTSI 410 LRPDTRYV
311 AELAKNVSL 361 LPYGANKDGI 411 VGYQPYRVV
312 SELVIGAVI 362 LPYPDPSRI 412 VRIQPGQTF
313 YERHSLSHF 363 LPLVSSQCV 413 WRNTNPIQL
314 AEVAVKMF 364 FAYTKRNVI 414 VRNLQHRLY
315 AEWFLAYILF 365 EVFAQVKQI 415 SRVLGLKTL
316 EESSAKSASVY 366 NPLLYDANY 416 LRVEAFEYY
317 SEYKGPITDVFY 367 FLPFFSNVTW 417 VRETMSYLF
318 SEYTGNYQCGHY 368 IHADQLTPTW 418 FKNLREFVF
319 TETDLTKGPHEF 369 LPPAYTNSF 419 NVIPTITQM
320 VENPDILRVY 370 IAIPTNFTI 420 ARAGEAANF
321 VENPHLMGWD 371 FPQSAPHGV 421 WKYPQVNGL
322 YSDVENPHLMGW 372 AVHFISNSW 422 FWRNTNPIQL
323 AAGLEAPFLYLY 373 NSIAIPTNF 423 FRSSVLHST
324 ADAGFIKQY 374 KMKDLSPRW 424 LRGTAVMSL
325 YENQKLIANQF 375 ATIPIQASL 425 GRVDGQVDL
326 YEQYIKWPW 376 VAMPNLYKM 426 TANPKTPKY
327 VENPDILRV 377 RSVASQSII 427 KKQQTVTLL
328 EEVVENPTI 378 VARDLSLQF 428 MKDLSPRWY
329 EEVGHTDLMAAY 379 STVFPPTSF 429 YRSLPGVF
330 VENMTPRDL 380 AKSHNIALIW 430 SRYWEPEF
331 REGVFVSNGTHW 381 ITFDNLKTL 431 RNRFLYIIKL
332 EEIAIILASF 382 LTAFGLVAEW 432 MYASAVVLL
333 EEAIRHVRAW 383 NKATYKPNTW 433 TRTQLPPAY
334 QEILGTVSW 384 SAKSASVYY 434 YADVFHLYL
335 SEFSSLPSY 385 ISTKHFYW 435 FLYLYALVY
336 MEVTPSGTW 386 RTTNGDFLHF 436 KFADDLNQL
337 AEAELAKNV 387 KSAGFPFNKW 437 YFDKAGQKTY
338 AEVQIDRLI 388 LTNDNTSRYW 438 YYKKDNSY
339 AEIRASANL 389 CATVHTANKW 439 YDYLVSTQEF
340 KEIDRLNEV 390 GVAPGTAVLRQW 440 QSAPHGVVF
341 GEVFNATRF 391 GVFVSNGTHW 441 LRIMASLVL
342 QELGKYEQY 392 VRSIFSRTL 442 YFTSDYYQL
343 KQEILGTVSW 393 YRGTTTYKL 443 VRIIMRLWL
344 QEYADVFHLY 394 YASAVVLLI 444 LVKPSFYVY
345 AEHVNNSY 395 GFMGRIRSV 445 FYYVWKSY
346 IAAVITREV 396 YVYSRVKNL 446 ARLYYDSMSY
347 DAVNLLTNM 397 AHAEETRKL 447 IYKTPPIKDF
348 LPGVYSVI 398 LRKHFSMMI 448 YFIKGLNNL
349 LPRVF SAV 399 CRSKNPLLY 449 TRFASVYAW
350 DAMRNAGIV 400 NSFSGYLKL 450 FYLITPVHV
451 IYDEPTTTT 501 KMQRMLLEK 551 MTYRRLISM
452 FLLPSLATV 502 AVAKHDFFK 552 FTSDYYQLY
453 FLLNKEMYL 503 HVVGPNVNK 553 MVMCGGSLY
454 TMADLVYAL 504 TMLFTMLRK 554 EVNSFSGYL
455 YLQPRTFLL 505 MTSCCSCLK 555 EVVGDIILK
456 YLNSTNVTI 506 AQCFKMFYK 556 SFYEDFLEY
457 HLVDFQVTI 507 HLYLQYIRK 557 VVYRGTTTY
458 VLNDILSRL 508 RQFHQKLLK 558 YILFTRFFY
459 ALWEIQQVV 509 TTIKPVTYK 559 FAIGLALYY
460 SVVSKVVKV 510 GVAMPNLYK 560 SMMGFKMNY
461 ALSKGVHFV 511 SSTCMMCYK 561 GVYSVIYLY
462 NLIDSYFVV 512 GTLSYEQFK 562 ATSRTLSYY
463 KIADYNYKL 513 QTMLFTMLR 563 ASHMYCSFY
464 LLYDANYFL 514 QTFFKLVNK 564 KMNYQVNGY
465 FVNEFYAYL 515 NYMPYFFTL 565 AVKTQFNYY
466 YLYALVYFL 516 IYNDKVAGF 566 ALCEKALKY
467 ILFTRFFYV 517 VYMPASWVM 567 MMSAPPAQY
468 FLNRFTTTL 518 YFVVKRHTF 568 RISNCVADY
469 YLNTLTLAV 519 VYIGDPAQL 569 GTFTCASEY
470 YLTNDVSFL 520 VYDPLQPEL 570 VYYPDKVFR
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
471 FLPRVFSAV 521 QYIKWPWYI 571 CSLSHRFYR
472 KLNIKLLGV 522 YFIASFRLF 572 HFYSKWYIR
473 FIAGLIAIV 523 YYTSNPTTF 573 AYYFMRFRR
474 FLNGSCGSV 524 YYQLYSTQL 574 AVHECFVKR
475 KLSYGIATV 525 TYKPNTWCI 575 KAIDGGVTR
476 VLAWLYAAV 526 VYFLQSINF 576 RVVRSIFSR
477 RTIKVFTTV 527 WSMATYYLF 577 NSLLTPFAR
478 VLWAHGFEL 528 SYYSLLMPI 578 RVKNLNSSR
479 LMIERFVSL 529 TYACWHHSI 579 NFYGPFVDR
480 SMWALIISV 530 SYFIASFRL 580 VTHSKGLYR
481 YTMADLVYA 531 DYQGKPLEF 581 ALHFLLFFR
482 FVAAIFYLI 532 NYNYLYRLF 582 CMMCYKRNR
483 SLPGVFCGV 533 FFASFYYVW 583 MSKFPLKLR
484 RIMTWLDMV 534 TYASALWEI 584 RTIKGTHHW
485 LQLGFSTGV 535 LYSPIFLIV 585 KAYNVTQAF
486 AVIKTLQPV 536 RYLALYNKY 586 KLLHKPIVW
487 KVDGVVQQL 537 EYADVFHLY 587 RMYIFFASF
488 FVDGVPFVV 538 MYIFFASFY 588 KFYGGWHNM
489 YMPYFFTLL 539 QYNRYLALY 589 KTPKYKFVR
490 YLDAYNMMI 540 YYPSARIVY 590 LFALLQRYR
491 MLDMYSVML 541 ETISLAGSY 591 RAMPNMLRI
492 GLMWLSYFI 542 DTYNLWNTF 592 KLAKKFDTF
493 NLSDRVVFV 543 DVTDVTQLY 593 KSYELQTPF
494 FLARGIVFM 544 ETKAIVSTI 594 RTNVYLAVF
495 AMDEFIERY 545 EIVDTVSAL 595 NVFAFPFTI
496 TLKSFTVEK 546 EAIRHVRAW 596 KVYPIILRL
497 HLMGWDYPK 547 EIAIILASF 597 VMFTPLVPF
498 LLFFRALPK 548 VVIPDYNTY 598 DYGDAVVYR
499 KLFAAETLK 549 DAQSFLNGF 599 DFDTWFSQR
500 RLISMMGFK 550 DVRETMSYL 600 DFYDFAVSK
601 VYADSFVIR 651 KPREQIDGY 701 MKIILFLAL
602 MTQMYKQAR 652 KPRQKRTAT 702 FAVDAAKAY
603 NYAKPFLNK 653 LPSLATVAY 703 FLHFLPRVF
604 IASFRLFAR 654 YLRKHFSMM 704 QLYLGGMSY
605 NTVIWDYKR 655 YLKLRSDVL 705 LMNVLTLVY
606 YAFASEAAR 656 FVKHKHAFL 706 YLVQQESPF
607 TVIEVQGYK 657 CLLNRYFRL 707 VQMAPISAM
608 TTDPSFLGR 658 DLFMRIFTI 708 ILMTARTVY
609 FSSEIIGYK 659 YFMRFRRAF 709 MISAGFSLW
610 MSAFAMMFV 660 FLKTNCCRF 710 NMVYMPASW
611 NATNVVIKV 661 DAPAHISTI 711 LPFFSNVTW
612 STSAFVETV 662 TQMNLKYAI 712 MSMTYGQQF
613 NTFSSTFNV 663 SQLGGLHLL 713 FISNSWLMW
614 SVAALTNNV 664 GEYSHVVAF 714 NMMVTNNTF
615 QSFLNGFAV 665 MMISAGFSL 715 HADQLTPTW
616 ETFKLSYGI 666 TQWSLFFFL 716 NVLEGSVAY
617 YTACSHAAV 667 REHEHEIAW 717 HMLDMYSVM
618 FSASTSAFV 668 RLVDPQIQL 718 FMGRIRSVY
619 HTIDGSSGV 669 NQMCLSTLM 719 VPWDTIANY
620 HVISTSHKL 670 RQWLPTGTL 720 DEWSMATYY
621 FSYFAVHFI 671 MQTMLFTML 721 VEHVTFFIY
622 DAQSFLNRV 672 FQFCNDPFL 722 DEISMATNY
623 NTQEVFAQV 673 RELHLSWEV 723 FELEDFIPM
624 TTFDSEYCR 674 MQVESDDYI 724 HEFCSQHTM
625 LSTFISAAR 675 YELQTPFEI 725 LEWLAMAVM
626 FLAYILFTR 676 RQLLFVVEV 726 NETLVTMPL
627 RVYANLGER 677 RSLKVPATV 727 LEIKDTEKY
628 FPRGQGVPI 678 KQIYKTPPI 728 SEVGPEHSL
629 IPRRNVATL 679 KQLIKVTLV 729 TEETFKLSY
630 SPRRARSVA 680 MQLFFSYFA 730 EEFEPSTQY
631 KPNELSRVL 681 SQNAVASKI 731 SEFDRDAAM
632 SPRWYFYYL 682 TQYNRYLAL 732 TELEPPCRF
633 SPYNSQNAV 683 NRFLYIIKL 733 MPYFFTLLL
634 IPVAYRKVL 684 DQFKHLIPL 734 LPFNDGVYF
635 RPDTRYVLM 685 EHYVRITGL 735 NPHLMGWDY
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
636 KPCIKVATV 686 TRFQTLLAL 736 YPNASFDNF
637 MARKTLNSL 687 DKVFRSSVL 737 LPGVYSVIY
638 HPTQAPTHL 688 YRLFRKSNL 738 CPDGVKHVY
639 IPIGAGICA 689 SHVVAFNTL 739 VPFVVSTGY
640 LPQNAVVKI 690 EHFIETISL 740 KRWQLALSK
641 MPNMLRIMA 691 VRFPNITNL 741 RRLISMMGF
642 IPIQASLPF 692 FRNARNGVL 742 GRWVLNNDY
643 KPVETSNSF 693 TRFFYVLGL 743 KRVDWTIEY
644 IPTNFTISV 694 HHMELPTGV 744 ARFYFYTSK
645 TPAFDKSAF 695 NNLNRGMVL 745 ARYMRSLKV
646 FPPTSFGPL 696 SHFVNLDNL 746 NRFNVAITR
647 MIAQYTSAL 697 QHEETIYNL 747 MRIMTWLDM
648 NPAWRKAVF 698 QHMVVKAAL 748 GRIRSVYPV
649 TPKGPKVKY 699 FRYMNSQGL 749 SRYRIGNYK
650 QPRTFLLKY 700 SHFAIGLAL 750 ARTRSMWSF
751 SRLSFKELL 801 GEQKSILSP 851 DAYNMMISA
752 RRVWTLMNV 802 TERLKLFAA 852 LPFKLTCAT
753 LPVNVAFEL 803 REAACCHLA 853 SPSGVYQCA
754 FPFTIYSLL 804 LQAAVGELL 854 KSHNIALIW
755 SAPPAQYEL 805 FEYVSQPFL 855 LSDLQDLKW
756 MPASWVMRI 806 FQVTIAEIL 856 RSFIEDLLF
757 FPDLNGDVV 807 AEIVDTVSA 857 VSFLAHIQW
758 YPSLETIQI 808 LEPEYFNSV 858 MACLVGLMW
759 FGADPIHSL 809 MEKLKTLVA 859 KAYKIEELF
760 VADAVIKTL 810 APFLYLYAL 860 LAAVYRINW
761 EAVGTNLPL 811 KPTVVVNAA 861 LAGTITSGW
762 YPLECIKDL 812 VENPHLMGW 862 VMPLSAPTL
763 FSSTFNVPM 813 SEKQVEQKI 863 SAPHGVVFL
764 QPTESIVRF 814 SEMHPALRL 864 GVAPGTAVL
765 FVSLAIDAY 815 SEDMLNPNY 865 HANEYRLYL
766 YPGQGLNGY 816 CPIHFYSKW 866 FLPGVYSVI
767 YANRNRFLY 817 TAFGLVAEW 867 VSPTKLNDL
768 FAYANRNRF 818 TASDTYACW 868 NVPLHGTIL
769 LAKDTTEAF 819 TPGDSSSGW 869 QLPAPRTLL
770 HHSIGFDYV 820 GETLGVLVP 870 TAPHGHVMV
771 MHAASGNLL 821 REAVGTNLP 871 IGPERTCCL
772 YHTTDPSFL 822 QEAYEQAVA 872 SANNCTFEY
773 LHSTQDLFL 823 SEYDYVIFT 873 FVLTSHTVM
774 THHWLLLTI 824 EELFYSYAT 874 IAMSAFAMM
775 VRDPQTLEI 825 NEYRLYLDA 875 VATSRTLSY
776 KHITSKETL 826 SEAGVCVST 876 FAQDGNAAI
777 THTGTGQAI 827 YLITPVHVM 877 FSNSGSDVL
778 THLSVDTKF 828 YSSANNCTF 878 YSTAALGVL
779 VHFVCNLLL 829 FASEAARVV 879 FVSDADSTL
780 IHFYSKWYI 830 MELPTGVHA 880 ISTSHKLVL
781 YKVYYGNAL 831 FENKTTLPV 881 FCYMHHMEL
782 LRSDVLLPL 832 MPLSAPTLV 882 VAKSHNIAL
783 NRALTGIAV 833 CPAEIVDTV 883 RTAPHGHVM
784 AEWFLAYIL 834 SPFELEDFI 884 TFDNLKTLL
785 HEGKTFYVL 835 MPTIFFAGI 885 RFDNPVLPF
786 GEAANFCAL 836 FPLCANGQV 886 ILDITPCSF
787 AEAAVKPLL 837 SAFYILPSI 887 KYDFTEERL
788 LENVAFNVV 838 LAWLYAAVI 888 LFDMSKFPL
789 TEVPVAIHA 839 MAYITGGVV 889 YGDFSHSQL
790 NESGLKTIL 840 VEYCPIFFI 890 YFTEQPIDL
791 HEVLLAPLL 841 CQYLNTLTL 891 VYDDGARRV
792 AECTIFKDA 842 TQFNYYKKV 892 VTDVTQLYL
793 REFLTRNPA 843 IQLSSYSLF 893 YVDNSSLTI
794 SEFRVYS SA 844 LAAVNSVPW 894 YSDVENPHL
795 YENAFLPFA 845 MPILTLTRA 895 VVDSYYSLL
796 IELKFNPPA 846 LPFAMGIIA 896 YIDIGNYTV
797 MEIDFLELA 847 MPVCVETKA 897 MADQAMTQM
798 GECPNFVFP 848 FPFNKWGKA 898 ANDPVGFTL
799 VELKHFFFA 849 FPREGVFVS 899 VSDIDITFL
800 LEFGATSAA 850 CPFGEVFNA 900 ISDEVARDL
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
901 ISDEFSSNV 951 GQVDLFRNA 1001 HSSGVTREL
902 VAFNTLLFL 952 AMRPNFTIK 1002 VTRELMREL
903 YYRYNLPTM 953 EILPVSMTK 1003 QSASKIITL
904 YYHTTDPSF 954 TVYDDGARR 1004 VGYLQPRTF
905 FYLTNDVSF 955 TIDYTEISF 1005 LSDRVVFVL
906 LYYQNNVFM 956 YTRYVDNNF 1006 EAANFCALI
907 AADPAMHAA 957 YYRSLPGVF 1007 MGYINVFAF
908 LSDGLLLAL 958 TVNVLAWLY 1008 KVDGVDVEL
909 ITDAVDCAL 959 KLIEYTDFA 1009 FGDDTVIEV
910 VSDIDYVPL 960 VLLSVLQQL 1010 LLLDRLNQL
911 VAYFNMVYM 961 PYNMRVIHF 1011 LLLEWLAMA
912 SFSASTSAF 962 TTIQTIVEV 1012 KIFVDGVPF
913 ISAMVRMYI 963 LATNNLVVM 1013 MLVYCFLGY
914 KTLLSLREV 964 VLVPHVGEI 1014 MPLKAPKEI
915 ISAGFSLWV 965 SQSIIAYTM 1015 DGYFKIYSK
916 YTNSFTRGV 966 NTVCTVCGM 1016 DTLKNLSDR
917 HAASGNLLL 967 EAFEKMVSL 1017 VMAYITGGV
918 AVASKILGL 968 YFYTSKTTV 1018 EAVMYMGTL
919 ALNNIINNA 969 LLDDFVEII 1019 FGDSVEEVL
920 KLWAQCVQL 970 HLDGEVITF 1020 VYSTGSNVF
921 SLSHRFYRL 971 HFYWFFSNY 1021 VVNVVTTKI
922 TLIGDCATV 972 FTVLCLTPV 1022 ELTPVVQTI
923 QMAPISAMV 973 KSVNITFEL 1023 ELPDEFVVV
924 VLSDRELHL 974 TTAAKLMVV 1024 YFPLQSYGF
925 YLATALLTL 975 LSDDAVVCF 1025 AIKITEHSW
926 MLAKALRKV 976 LPNDDTLRV 1026 ALDQAISMW
927 HSIGFDYVY 977 YTVELGTEV 1027 SLIDFYLCF
928 NYSGVVTTV 978 LEYHDVRVV 1028 RVESSSKLW
929 YQKVGMQKY 979 IADKYVRNL 1029 YKKPASREL
930 TVAYFNMVY 980 THVQLSLPV 1030 LMDGSIIQF
931 ALNTLVKQL 981 LLLALHFLL 1031 LAVPYNMRV
932 YFKYWDQTY 982 KLLEQWNLV 1032 SGFAAYSRY
933 VFLGIITTV 983 YITGGVVQL 1033 QEYADVFHL
934 VTWFHAIHV 984 NVLTLVYKV 1034 FAFACPDGV
935 KVQIGEYTF 985 ALCTFLLNK 1035 QVVDMSMTY
936 FLTENLLLY 986 RVDGQVDLF 1036 FTNVYADSF
937 S SDNIALLV 987 YGIATVREV 1037 KTIGPDMFL
938 RSVSPKLFI 988 LLFNKVTLA 1038 NRGMVLGSL
939 IVDTVSALV 989 ERSEKSYEL 1039 YVFCTVNAL
940 KYTQLCQYL 990 HHWLLLTIL 1040 TLKNTVCTV
941 TSDLATNNL 991 YFNSVCRLM 1041 NLWNTFTRL
942 TVASLINTL 992 VSNGTHWFV 1042 SALWEIQQV
943 VSDADSTLI 993 FVSNGTHWF 1043 KNFKSVLYY
944 VLYENQKLI 994 ISNSWLMWL 1044 EAAVKPLLV
945 NYLKRRVVF 995 DTVIEVQGY 1045 ILFALLQRY
946 EAMYTPHTV 996 FLAFLLFLV 1046 YLCFLAFLL
947 NQKLIANQF 997 KLNVGDYFV 1047 LANECAQVL
948 FLAFVVFLL 998 AYANSVFNI 1048 LPQLEQPYV
949 LLLDDFVEI 999 YYVGYLQPR 1049 SWVMRIMTW
950 ALLADKFPV 1000 DALFAYTKR 1050 RTATKAYNV
1051 KVFTTVDNI 1101 MSDVKCTSV 1151 FEYYHTTDP
1052 KKFLPFQQF 1102 DEFIERYKL 1152 YVWKSYVHV
1053 FHQKLLKSI 1103 ISMDNSPNL 1153 LASHMYCSF
1054 NHTSPDVDL 1104 NPNYEDLLI 1154 SVNPYVCNA
1055 AHIQWMVMF 1105 SPNLAWPLI 1155 DFNLVAMKY
1056 TTTIKPVTY 1106 LIISVTSNY 1156 RLRAKHYVY
1057 KYKYFSGAM 1107 MSYEDQDAL 1157 STTTNIVTR
1058 QFAPSASAF 1108 SSLPSYAAF 1158 DIQLLKSAY
1059 YTNDKACPL 1109 ASANLAATK 1159 KENSYTTTI
1060 VPHHVVATV 1110 VQQESPFVM 1160 VLSFCAFAV
1061 KWDLTAFGL 1111 IVSTIQRKY 1161 FLALCADSI
1062 FVTVYSHLL 1112 ETMSYLFQH 1162 ETICAPLTV
1063 LRLGSPLSL 1113 FEEAALCTF 1163 FLFVAAIFY
1064 VINGDRWFL 1114 YLASGGQPI 1164 RYFRLTLGV
1065 DILSRLDKV 1115 MSNLGMPSY 1165 AIDAYPLTK
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1066 YECDIPIGA 1116 MVDTSLSGF 1166 SAGFPFNKW
1067 KSHKPPISF 1117 RAFGEYSHV 1167 IAIAMACLV
1068 TVEEAKTVL 1118 TQAPTHLSV 1168 ETTADIVVF
1069 AVKPLLVPH 1119 AQLPAPRTL 1169 HFISNSWLM
1070 ATLPKGIMM 1120 SAFFGMSRI 1170 HSWNADLYK
1071 HTDLMAAYV 1121 FLGRYMSAL 1171 AIMQLFFSY
1072 VPYCYDTNV 1122 CAMRPNFTI 1172 QHEVLLAPL
1073 NRNRFLYII 1123 AIMTRCLAV 1173 MYDPKTKNV
1074 VLYYQNNVF 1124 YYKLGASQR 1174 LSPRWYFYY
1075 KTQFNYYKK 1125 TVVIGTSKF 1175 KVSIWNLDY
1076 REETGLLMP 1126 LVAVPTGYV 1176 CSFGGVSVI
1077 QVDVVNFNL 1127 YSGVVTTVM 1177 FSTFEEAAL
1078 VMYASAVVL 1128 ILSPLYAFA 1178 EETGLLMPL
1079 AAYVDNSSL 1129 MSLSEQLRK 1179 MTNRQFHQK
1080 HFAIGLALY 1130 FEIKLAKKF 1180 SLRPDTRYV
1081 SLVKPSFYV 1131 ILHCANFNV 1181 SYLFQHANL
1082 LQLPQGTTL 1132 IGAGICASY 1182 DYKHYTPSF
1083 KLVNKFLAL 1133 LRPDTRYVL 1183 ESGLKTILR
1084 NVAFNVVNK 1134 FSKQLQQSM 1184 LSFKELLVY
1085 SRYWEPEFY 1135 AANTVIWDY 1185 NAANVYLKH
1086 TSNPTTFHL 1136 TPLIQPIGA 1186 AFPFTIYSL
1087 FYAYLRKHF 1137 NYDLSVVNA 1187 KATEETFKL
1088 QLFFSYFAV 1138 YKTPPIKDF 1188 TSMKYFVKI
1089 AMSAFAMMF 1139 AWPLIVTAL 1189 IINNTVYTK
1090 WLPTGTLLV 1140 QSINFVRII 1190 SFKWDLTAF
1091 ETAQNSVRV 1141 KHAFHTPAF 1191 MFVKHKHAF
1092 TYFTQSRNL 1142 LLMPLKAPK 1192 KLIFLWLLW
1093 FVVSTGYHF 1143 TVKPGNFNK 1193 WVYKQFDTY
1094 SQLMCQPIL 1144 YKGPITDVF 1194 LALGGSVAI
1095 GSIHLYFDK 1145 MFLARGIVF 1195 FAAYSRYRI
1096 LQTPFEIKL 1146 ALAYYNTTK 1196 FVSEETGTL
1097 TPCNGVEGF 1147 GTYEGNSPF 1197 STKHFYWFF
1098 AHGFELTSM 1148 LIDFYLCFL 1198 INFVRIIMR
1099 GTSKFYGGW 1149 TLDSKTQSL 1199 TTIVYLTIV
1100 WMESEFRVY 1150 KWDLIISDM 1200 DTYPSLETI
1201 KDLPKEITV 1251 GLPWNVVRI 1301 LLIIMRTFK
1202 TTKGGRFVL 1252 NLLKDCPAV 1302 GQQFGPTYL
1203 LVLSVNPYV 1253 DYLVSTQEF 1303 FVMMSAPPA
1204 NVLAWLYAA 1254 DSKEGFFTY 1304 QWSLFFFLY
1205 WEIQQVVDA 1255 KVNSTLEQY 1305 FVNLKQLPF
1206 ASAFFGMSR 1256 NELSPVALR 1306 SLPSYAAFA
1207 IANQFNSAI 1257 APISAMVRM 1307 FTINCQEPK
1208 WLMWLIINL 1258 FACPDGVKH 1308 DVFHLYLQY
1209 LEIPRRNVA 1259 FYWFFSNYL 1309 YFVKIGPER
1210 ILTSLLVLV 1260 IAATRGATV 1310 HVGEIPVAY
1211 TIVEVQPQL 1261 IAQYTSALL 1311 VAELEGIQY
1212 ELYHYQECV 1262 NSFTRGVYY 1312 TEQPIDLVP
1213 IIIGGAKLK 1263 FNATRFASV 1313 LPKGIMMNV
1214 VFAFPFTIY 1264 IMASLVLAR 1314 YTPSKLIEY
1215 VTANVNALL 1265 WQLALSKGV 1315 LLLTILTSL
1216 ALLTKSSEY 1266 VLSTFISAA 1316 NRQFHQKLL
1217 LTDEMIAQY 1267 KAIVSTIQR 1317 RLQSLQTYV
1218 LLALHRSYL 1268 LYQPPQTSI 1318 KHAFLCLFL
1219 LLNKEMYLK 1269 STDTCFANK 1319 AAAYYVGYL
1220 SILSPLYAF 1270 RYKLEGYAF 1320 MGIIAMSAF
1221 VVNAANVYL 1271 TTITVNVLA 1321 AIASEFS SL
1222 GFDYVYNPF 1272 MLQSCYNFL 1322 FKHLIPLMY
1223 IFTIGTVTL 1273 FLKKDAPYI 1323 EAFLIGCNY
1224 RVVFNGVSF 1274 DIAANTVIW 1324 AHSCNVNRF
1225 QQWGFTGNL 1275 RVWTLMNVL 1325 KHADFDTWF
1226 REVGFVVPG 1276 MVYMPASWV 1326 EPKLGSLVV
1227 SAQTGIAVL 1277 LTNMFTPLI 1327 LPFGWLIVG
1228 HSLSHFVNL 1278 QAWQPGVAM 1328 IAQVDVVNF
1229 RVFSAVGNI 1279 ASLPFGWLI 1329 RFRRAFGEY
1230 LYLDAYNMM 1280 MSYYCKSHK 1330 VYANGGKGF
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1231 SASVYYSQL 1281 QTFSVLACY 1331 EVQELYSPI
1232 NLHSSRLSF 1282 VSFCYMHHM 1332 VMVELVAEL
1233 SYEDQDALF 1283 AFDIYNDKV 1333 YPIILRLGS
1234 FELTSMKYF 1284 FPLNSIIKT 1334 KATNNAMQV
1235 VLGSLAATV 1285 VHFISNSWL 1335 DISGINASV
1236 NVLSTFISA 1286 ALLQRYRYK 1336 RTIAFGGCV
1237 STEKSNIIR 1287 LESELVIGA 1337 EVITFDNLK
1238 GIITTVAAF 1288 VSIINNTVY 1338 HSMQNCVLK
1239 IELSLIDFY 1289 RELKVTFFP 1339 HFDGQQGEV
1240 FTDGVCLFW 1290 DGARRVWTL 1340 VTVYSHLLL
1241 KATYKPNTW 1291 LTNDNTSRY 1341 TLGVLVPHV
1242 FRKSNLKPF 1292 LYENAFLPF 1342 DTANPKTPK
1243 VLSGHNLAK 1293 YPANSIVCR 1343 VEVEKGVLP
1244 QLIKVTLVF 1294 EVNEFACVV 1344 AARVVRSIF
1245 RVCGVSAAR 1295 EVFNATRFA 1345 SSRVPDLLV
1246 STVLSFCAF 1296 TLKATEETF 1346 REQIDGYVM
1247 MATNYDLSV 1297 VLNEKCSAY 1347 TSVDCTMYI
1248 KLVLSVNPY 1298 RDLPQGFSA 1348 RHHANEYRL
1249 TLVTMPLGY 1299 LLFLMSFTV 1349 YIVDSVTVK
1250 VQLHNDILL 1300 LVYAADPAM 1350 TILDGISQY
1351 TLAVPYNMR 1401 RNYVFTGYR 1451 IEDLLFNKV
1352 RQEEVQELY 1402 AVDAAKAYK 1452 IAKNTVKSV
1353 HEETIYNLL 1403 ITHDVSSAI 1453 CPACHNSEV
1354 GYLPQNAVV 1404 NTWCIRCLW 1454 YSLFDMSKF
1355 VGTNLPLQL 1405 FQTRAGCLI 1455 GLNDNLLEI
1356 LYLGGMSYY 1406 AIFYLITPV 1456 DKAYKIEEL
1357 TLNDLNETL 1407 YLEGSVRVV 1457 WPWYIWLGF
1358 EAARYMRSL 1408 SIIAYTMSL 1458 SLEDKAFQL
1359 STYASQGLV 1409 FQPTNGVGY 1459 KISEMHPAL
1360 FEHIVYGDF 1410 TDYKHWPQI 1460 DEFVVVTVK
1361 TAQNSVRVL 1411 HDIGNPKAI 1461 ASAVVLLIL
1362 FVFPLNSII 1412 LLLQILFAL 1462 LLTNMFTPL
1363 MKFLVFLGI 1413 YFVLTSHTV 1463 KWYIRVGAR
1364 KAPKEIIFL 1414 SVIYLYLTF 1464 TSNSFDVLK
1365 SYSGQSTQL 1415 VMHANYIFW 1465 LQIPFAMQM
1366 ATVHTANKW 1416 MESEFRVYS 1466 VYSVIYLYL
1367 MFYKGVITH 1417 YTEISFMLW 1467 SLLSVLLSM
1368 IIFWF SLEL 1418 RLSFKELLV 1468 GVRRSFYVY
1369 AQVLSEMVM 1419 SVFNICQAV 1469 TTIVNGVRR
1370 FYGGWHNML 1420 VFVSNGTHW 1470 RQVVNVVTT
1371 CVPLNIIPL 1421 DSIIIGGAK 1471 KIQEGVVDY
1372 NDLCFTNVY 1422 GEVITFDNL 1472 MYTPHTVLQ
1373 YTDFATSAC 1423 DAMMFTSDL 1473 RLYLDAYNM
1374 DEFTPFDVV 1424 SAFVNLKQL 1474 LKKRWQLAL
1375 FFGMSRIGM 1425 DQAISMWAL 1475 ATVAYFNMV
1376 ETLGVLVPH 1426 VAVKMFDAY 1476 ALVYFLQSI
1377 YWFFSNYLK 1427 RVCTNYMPY 1477 DEMIAQYTS
1378 ITVATSRTL 1428 KTSVDCTMY 1478 ATAEAELAK
1379 WPVTLACFV 1429 EKALKYLPI 1479 SEYTGNYQC
1380 LPSYAAFAT 1430 GHFAWWTAF 1480 TVYTKVDGV
1381 MVTNNTFTL 1431 SIIKTIQPR 1481 VVDKYFDCY
1382 ELLHAPATV 1432 YYFMRFRRA 1482 DVRVVLDFI
1383 KHDFFKFRI 1433 TSHKLVLSV 1483 TTLPVNVAF
1384 VYNPFMIDV 1434 FIASFRLFA 1484 LPLQLGFST
1385 KSFDLGDEL 1435 AYLRKHFSM 1485 KNFTTAPAI
1386 ILIVTTIVY 1436 LYYPSARIV 1486 NAAISDYDY
1387 RLWLCWKCR 1437 DLQDLKWAR 1487 DLPQGFSAL
1388 CVMYASAVV 1438 CTDDNALAY 1488 NSQGLLPPK
1389 KLQFTSLEI 1439 TPRDLGACI 1489 LTNIFGTVY
1390 HTDFSSEII 1440 NEEIAIILA 1490 GPFVDRQTA
1391 KQGDDYVYL 1441 TLAILTALR 1491 SALNHTKKW
1392 LTYTGAIKL 1442 RTILGSALL 1492 VGDSAEVAV
1393 ETSWQTGDF 1443 VPHVGEIPV 1493 TTFTYASAL
1394 DAVTAYNGY 1444 NYTVSCLPF 1494 QSSYIVDSV
1395 QEIQLQAAV 1445 FLCWHTNCY 1495 AVDCALDPL
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1396 NALDQAISM 1446 EVPANSTVL 1496 LVASIKNFK
1397 GHMLDMYSV 1447 QVVSDIDYV 1497 TTFHLDGEV
1398 RYMNSQGLL 1448 SVSPKLFIR 1498 DHVDILGPL
1399 LLDKRTTCF 1449 VETVKGLDY 1499 QLCQYLNTL
1400 LAHIQWMVM 1450 TESNKKFLP 1500 KLFIRQEEV
1501 GVYDYLVST 1551 TLGVYDYLV 1601 SFNPETNIL
1502 GERVRQALL 1552 TTVDNINLH 1602 FLRDGWEIV
1503 IDYVPLKSA 1553 STDVVYRAF 1603 YMPASWVMR
1504 NFCALILAY 1554 LTILTSLLV 1604 KVKYLYFIK
1505 YHFRELGVV 1555 QLMCQPILL 1605 LTFYLTNDV
1506 HKPPISFPL 1556 ARNGVLITE 1606 FALLQRYRY
1507 SFRLFARTR 1557 KHTDFSSEI 1607 WLTNIFGTV
1508 ILGTVSWNL 1558 AEYHNESGL 1608 LAAECTIFK
1509 SSNVANYQK 1559 LEQYVFCTV 1609 WHNMLKTVY
1510 TPVCINGLM 1560 KPVPEVKIL 1610 VSQPFLMDL
1511 FYYVWKSYV 1561 AKPPPGDQF 1611 GLFKDCSKV
1512 IVAGGIVAI 1562 MMFVKHKHA 1612 NELSRVLGL
1513 KLMPVCVET 1563 AMPNMLRIM 1613 GADLKSFDL
1514 AFATAQEAY 1564 IQYIDIGNY 1614 AMQTMLFTM
1515 WFSQRGGSY 1565 STCMMCYKR 1615 TPVVQTIEV
1516 TFEYVSQPF 1566 MVSLLSVLL 1616 NYYKKDNSY
1517 MLRIMASLV 1567 SFLAHIQWM 1617 LVSTQEFRY
1518 YRARAGEAA 1568 GTTTLNGLW 1618 CASEYTGNY
1519 YPSARIVYT 1569 ALLAVFQSA 1619 GHSMQNCVL
1520 LVYDNKLKA 1570 TEVPANSTV 1620 LSVCLGSLI
1521 AAVDALCEK 1571 DESGEFKLA 1621 TIAFGGCVF
1522 TLRVEAFEY 1572 GVFCGVDAV 1622 KVVKVTIDY
1523 FRVQPTESI 1573 FMRFRRAFG 1623 EQPYVFIKR
1524 YVLPNDDTL 1574 YKQFDTYNL 1624 LEILQKEKV
1525 LHDELTGHM 1575 IATNGPLKV 1625 EPVLKGVKL
1526 WVPRASANI 1576 HVVAFNTLL 1626 SIWNLDYII
1527 IVVTCLAYY 1577 RLANECAQV 1627 SRTLSYYKL
1528 TLLALHRSY 1578 YVVDDPCPI 1628 LQSCYNFLK
1529 TTPGSGVPV 1579 EEIAIILAS 1629 SVLLFLAFV
1530 VMCGGSLYV 1580 LAMAVMLLL 1630 KMVSLLSVL
1531 QMCLSTLMK 1581 RVVVLSFEL 1631 LNDFNLVAM
1532 NLKTLLSLR 1582 EFTPFDVVR 1632 NSPNLAWPL
1533 STQDLFLPF 1583 NQHEVLLAP 1633 VVNQNAQAL
1534 TLSEQLDFI 1584 CEFCGTENL 1634 VTIAEILLI
1535 FLEYHDVRV 1585 FLQSINFVR 1635 SEYCRHGTC
1536 KRFDNPVLP 1586 LQSINFVRI 1636 KYVRNLQHR
1537 REVLSDREL 1587 SVLYNSASF 1637 KMAFPSGKV
1538 QEKNFTTAP 1588 IANYAKPFL 1638 VVIGTSKFY
1539 IVDEPEEHV 1589 LPGCDGGSL 1639 QTYVTQQLI
1540 AQYTSALLA 1590 MRIFTIGTV 1640 DKFKVNSTL
1541 ETVKGLDYK 1591 VLLAPLLSA 1641 KLMVVIPDY
1542 AHVASCDAI 1592 LQGPPGTGK 1642 VYQLRARSV
1543 MGFKMNYQV 1593 LYLQYIRKL 1643 SGDGTTSPI
1544 QALLKTVQF 1594 DMFLGTCRR 1644 FELLHAPAT
1545 VPYNMRVIH 1595 MPSYCTGYR 1645 QELYSPIFL
1546 SHTVMPLSA 1596 FKMFYKGVI 1646 TFTYASALW
1547 EEAIRHVRA 1597 YYSLLMPIL 1647 TENLTKEGA
1548 VPNQPYPNA 1598 NLKQLPFFY 1648 TLVPQEHYV
1549 STDTGVEHV 1599 GFAAYSRYR 1649 HLMSFPQSA
1550 QAISMWALI 1600 DYGARFYFY 1650 AKYTQLCQY
1651 GELGDVRET 1701 LFLMSFTVL 1751 NIFGTVYEK
1652 VEAPLVGTP 1702 NYMLTYNKV 1752 ALPETTADI
1653 SVAIKITEH 1703 ADSIIIGGA 1753 FFLYENAFL
1654 RNAGIVGVL 1704 TQLGIEFLK 1754 KTLQPVSEL
1655 GYYRRATRR 1705 VGMQKYSTL 1755 SAQCFKMFY
1656 LWLLWPVTL 1706 LAPLLSAGI 1756 IMRTFKVSI
1657 GSFCTQLNR 1707 AAKKNNLPF 1757 HVVATVQEI
1658 GIATVREVL 1708 ISTIGVCSM 1758 KLTDNVYIK
1659 QECVRGTTV 1709 TFFKLVNKF 1759 DTVRTNVYL
1660 KKNNLPFKL 1710 GLLLALHFL 1760 AEETRKLMP
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1661 RLDKVEAEV 1711 AQVAKSHNI 1761 SIIIGGAKL
1662 TTAYANSVF 1712 KVTLVFLFV 1762 FFFLYENAF
1663 TGDSCNNYM 1713 VRIKIVQML 1763 MADLVYALR
1664 QESPFVMMS 1714 ARVECFDKF 1764 YINVFAFPF
1665 GYKSVNITF 1715 SYIVDSVTV 1765 SEAKCWTET
1666 EYVSQPFLM 1716 YQVNGYPNM 1766 CTSCCFSER
1667 KGLPWNVVR 1717 TVLCLTPVY 1767 YVPLKSATC
1668 KEKVNINIV 1718 KWGKARLYY 1768 QQGEVPVSI
1669 AAFHQECSL 1719 LEGNFYGPF 1769 VHTANKWDL
1670 LAYILFTRF 1720 CPAVAKHDF 1770 KDTEKYCAL
1671 KFLKTNCCR 1721 RYRYKPHSL 1771 VSTSGRWVL
1672 GVKHVYQLR 1722 TLSYEQFKK 1772 FTALTQHGK
1673 EAALCTFLL 1723 AFEKMVSLL 1773 QASLNGVTL
1674 AQKFNGLTV 1724 SLVPGFNEK 1774 YFDKAGQKT
1675 ITLATCELY 1725 LVLVQSTQW 1775 LLSAGIFGA
1676 VPLKSATCI 1726 DTTEAFEKM 1776 YSLRLIDAM
1677 IPTITQMNL 1727 HEIAWYTER 1777 NKHAFHTPA
1678 AALQIPFAM 1728 DEVARDLSL 1778 MFTPLIQPI
1679 KLNEEIAII 1729 GEFKLASHM 1779 AVPYNMRVI
1680 ILGLPTQTV 1730 TPSDFVRAT 1780 DFVNEFYAY
1681 TTLKGVEAV 1731 GSLAATVRL 1781 SANLAATKM
1682 EEVQELYSP 1732 WMVMFTPLV 1782 FAFPFTIYS
1683 KPLLVPHHV 1733 SNYLKRRVV 1783 TQSRNLQEF
1684 FYVLGLAAI 1734 AYSNNSIAI 1784 GTITSGWTF
1685 MDNSPNLAW 1735 ANGQVFGLY 1785 AAVGELLLL
1686 EVGKPRPPL 1736 VYSFLPGVY 1786 LTSHTVMPL
1687 AAVINGDRW 1737 LNRYFRLTL 1787 LAAIMQLFF
1688 IAKKPTETI 1738 LSPVALRQM 1788 LSVLQQLRV
1689 KHWPQIAQF 1739 WEIVKFIST 1789 IFWRNTNPI
1690 LAILTALRL 1740 LIVNSVLLF 1790 DLYKLMGHF
1691 AFGGCVFSY 1741 VFLFVAAIF 1791 YSQLMCQPI
1692 ASFSTFKCY 1742 IEELFYSYA 1792 TIDGSSGVV
1693 LHCANFNVL 1743 LQTYVTQQL 1793 LWPVTLACF
1694 TRVLSNLNL 1744 IVNSVLLFL 1794 WLAMAVMLL
1695 ISAARQGFV 1745 LYIDINGNL 1795 FLNKVVSTT
1696 CEESSAKSA 1746 DLSPRWYFY 1796 VFVLWAHGF
1697 CGPKKSTNL 1747 NPKTPKYKF 1797 FTPLVPFWI
1698 SSTASALGK 1748 KPYIKWDLL 1798 HFVNLDNLR
1699 DHSSSSDNI 1749 DADSKIVQL 1799 ASKIITLKK
1700 AYANRNRFL 1750 RFKESPFEL 1800 IIAMSAFAM
1801 EAVKTQFNY 1851 SLAIDAYPL 1901 IAMACLVGL
1802 STASALGKL 1852 VVENPTIQK 1902 LAFVVFLLV
1803 RITGLYPTL 1853 GYAFEHIVY 1903 TSRTLSYYK
1804 DPFLGVYYH 1854 QQLIRAAEI 1904 TLNDFNLVA
1805 QKFNGLTVL 1855 VANYQKVGM 1905 YLKSPNFSK
1806 MNVAKYTQL 1856 KYTMADLVY 1906 ELKINAACR
1807 YYVWKSYVH 1857 YAADPAMHA 1907 QTVTLLPAA
1808 NLYDKLVSS 1858 YLAVFDKNL 1908 NEKCSAYTV
1809 ALDISASIV 1859 SPFVMMSAP 1909 VAFELWAKR
1810 KWADNNCYL 1860 ETKDVVECL 1910 SLSDGLLLA
1811 EQKSILSPL 1861 SSPDAVTAY 1911 LEQPTSEAV
1812 RKSAPLIEL 1862 IGYYRRATR 1912 KTILRKGGR
1813 IAIVMVTIM 1863 KSWMESEFR 1913 VAAIVFITL
1814 AANFCALIL 1864 IFFASFYYV 1914 QFTSLEIPR
1815 MVPHISRQR 1865 KVQHMVVKA 1915 GPKVYPIIL
1816 SEAVEAPLV 1866 TTLNDFNLV 1916 HYVRITGLY
1817 HWFVTQRNF 1867 IPKEEVKPF 1917 LPIDKCSRI
1818 IMQLFFSYF 1868 NYQHEETIY 1918 QIYKTPPIK
1819 IQPGQTFSV 1869 YIKWPWYIW 1919 EAARVVRSI
1820 IIMRLWLCW 1870 RAAEIRASA 1920 VLCNSQTSL
1821 SLPINVIVF 1871 YVLGLAAIM 1921 LTSMKYFVK
1822 KIITLKKRW 1872 GMSRIGMEV 1922 MLSDTLKNL
1823 SKVGGNYNY 1873 ASIKNFKSV 1923 LTNDVSFLA
1824 IQASLPFGW 1874 AFLPFAMGI 1924 VIYLYLTFY
1825 ITLKKRWQL 1875 GVDIAANTV 1925 CEIVGGQIV
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1826 LMSFTVLCL 1876 KFKEGVEFL 1926 SLLMPILTL
1827 LAWPLIVTA 1877 NSYFTEQPI 1927 LLKSIAATR
1828 ETIQITISS 1878 FNPETNILL 1928 AYWVPRASA
1829 AVVCFNSTY 1879 TSLSGFKLK 1929 RNFYEPQII
1830 IERFVSLAI 1880 FSTFKCYGV 1930 HEHEIAWYT
1831 VTDTPKGPK 1881 INDMILSLL 1931 LVDSDLNDF
1832 LPIGINITR 1882 TEKYCALAP 1932 LLMPILTLT
1833 QAITVTPEA 1883 QINDMILSL 1933 YGQQFGPTY
1834 ILKPANNSL 1884 YTVEEAKTV 1934 HFLLFFRAL
1835 TVQFCDAMR 1885 KPFLNKVVS 1935 ALYNKYKYF
1836 QPILLLDQA 1886 KQGNFKNLR 1936 GDMVPHISR
1837 YSKHTPINL 1887 KPHSLSDGL 1937 QSTQWSLFF
1838 HMVVKAALL 1888 RLIDAMMFT 1938 LQKAAITIL
1839 KRAKVTSAM 1889 NTLTLAVPY 1939 VLDMCASLK
1840 IPYNSVTSS 1890 NENGTITDA 1940 KEGATTCGY
1841 ASCDAIMTR 1891 LTYNKVENM 1941 HSLSDGLLL
1842 TGSNVFQTR 1892 VAYRKVLLR 1942 WLSYFIASF
1843 VEQKIAEIP 1893 IEYTDFATS 1943 FLPFAMGII
1844 NIDYDCVSF 1894 CYFGLFCLL 1944 SMWSFNPET
1845 KSPNFSKLI 1895 NPPALQDAY 1945 GSVRVVTTF
1846 YYHKNNKSW 1896 SDRVVFVLW 1946 MLFTMLRKL
1847 TTCCSLSHR 1897 LTKHPNQEY 1947 YEAMYTPHT
1848 VSEETGTLI 1898 NCYDYCIPY 1948 ILLLDQALV
1849 STECSNLLL 1899 TVREVLSDR 1949 RIMASLVLA
1850 TVYEKLKPV 1900 QPITNCVKM 1950 QEGVVDYGA
1951 EKMVSLLSV 2001 NIDGYFKIY 2051 YIKWDLLKY
1952 ISTKHFYWF 2002 DLKWARFPK 2052 GTSTDVVYR
1953 SNVTWFHAI 2003 YFRLTLGVY 2053 LYIIKLIFL
1954 TSNQVAVLY 2004 AALTNNVAF 2054 EYPIIGDEL
1955 RNIKPVPEV 2005 ESSAKSASV 2055 SVCLGSLIY
1956 APLLSAGIF 2006 LLQLCTFTR 2056 NQTTTIQTI
1957 LVS SFLEMK 2007 LLTKSSEYK 2057 RIIPARARV
1958 FSNYLKRRV 2008 EEAALCTFL 2058 VVISSDVLV
1959 LTTAAKLMV 2009 CVDIPGIPK 2059 AYKIEELFY
1960 ISDYDYYRY 2010 EYFNSVCRL 2060 QIDRLITGR
1961 RLNEVAKNL 2011 FFSYFAVHF 2061 ESVQTFFKL
1962 GVYYPDKVF 2012 LPDEFVVVT 2062 MQMAYRFNG
1963 NVAFELWAK 2013 NYIAQVDVV 2063 ASVYAWNRK
1964 LSYGIATVR 2014 EANMDQESF 2064 LTALRLCAY
1965 VRATATIPI 2015 RPLLESELV 2065 MSFPQSAPH
1966 YLVSTQEFR 2016 NTYLEGSVR 2066 TIAEILLII
1967 NMLRIMASL 2017 DLKGKYVQI 2067 VVFDEISMA
1968 MMILSDDAV 2018 KLPDDFTGC 2068 VTFFIYNKI
1969 FLLVTLAIL 2019 QMAYRFNGI 2069 LATCELYHY
1970 GQGLNGYTV 2020 ASQGLVASI 2070 RLKLFDRYF
1971 GPKVKYLYF 2021 WHHSIGFDY 2071 TSTDVVYRA
1972 VLKLKVDTA 2022 SEETGTLIV 2072 QHGKEDLKF
1973 EVGFVVPGL 2023 KYAISAKNR 2073 HDELTGHML
1974 FIDTKRGVY 2024 LAVFDKNLY 2074 LPTGVHAGT
1975 AFQLTPIAV 2025 YDANYFLCW 2075 TYRRLISMM
1976 GTTTYKLNV 2026 FTTVDNINL 2076 CALAPNMMV
1977 EEMLDNRAT 2027 AQALNTLVK 2077 FLHVTYVPA
1978 LAFLLFLVL 2028 VEKGVLPQL 2078 YSRYRIGNY
1979 YLALYNKYK 2029 IAEIPKEEV 2079 IAPGQTGKI
1980 DLVYALRHF 2030 KALNLGETF 2080 CYKRNRATR
1981 TVSCLPFTI 2031 LPRVFSAVG 2081 FSSLPSYAA
1982 TPGSGVPVV 2032 STQWSLFFF 2082 KEILVTYNC
1983 DKAFQLTPI 2033 DNLKTLLSL 2083 LALSKGVHF
1984 REVRTIKVF 2034 KWDLLKYDF 2084 YEYGTEDDY
1985 LQHRLYECL 2035 LIDSYFVVK 2085 KLKPVLDWL
1986 EEHVQIHTI 2036 GDCATVHTA 2086 KIILFLALI
1987 ADQAMTQMY 2037 KSAQCFKMF 2087 AESHVDTDL
1988 AGIVGVLTL 2038 RKGGRTIAF 2088 FYEPQIITT
1989 SEISMDNSP 2039 SHEGKTFYV 2089 GYPNMFITR
1990 NLDYIINLI 2040 LLTILTSLL 2090 HSDKFTDGV
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
1991 ELFENKTTL 2041 QWMVMFTPL 2091 IRKSNHNFL
1992 VVPGLPGTI 2042 LRAKHYVYI 2092 LTLVYKVYY
1993 SLENVAFNV 2043 NFGAISSVL 2093 SVKGLQPSV
1994 KFTDGVCLF 2044 FAMMFVKHK 2094 DYIINLIIK
1995 AFLIGCNYL 2045 MRNAGIVGV 2095 VDILGPLSA
1996 LPETTADIV 2046 SALEPLVDL 2096 GPKKSTNLV
1997 KYLYFIKGL 2047 TVMFLARGI 2097 VLDWLEEKF
1998 TVMPLSAPT 2048 KPANNSLKI 2098 VFITLCFTL
1999 ITLCFTLKR 2049 ESNKKFLPF 2099 VVYCPRHVI
2000 ISQYSLRLI 2050 TALLTLQQI 2100 LVKNKCVNF
2101 YREAACCHL 2151 KVVSTTTNI 2201 NIIPLTTAA
2102 KPASRELKV 2152 MLLLQILFA 2202 EQFKKGVQI
2103 GTDLEGNFY 2153 NLKYAISAK 2203 AVITREVGF
2104 NLAKHCLHV 2154 LDYKAFKQI 2204 FWITIAYII
2105 GELLLLEWL 2155 LPAPRTLLT 2205 YAFEHIVYG
2106 DEFSSNVAN 2156 FVKRVDWTI 2206 DFQENWNTK
2107 TAAKLMVVI 2157 FPNTYLEGS 2207 ARTVAGVSI
2108 KARLYYDSM 2158 FVLALLSDL 2208 TPINLVRDL
2109 VFDEISMAT 2159 NRNYVFTGY 2209 IMRLWLCWK
2110 TCDGTTFTY 2160 YHPNCVNCL 2210 EISFMLWCK
2111 GHNLAKHCL 2161 YIRKLHDEL 2211 YAAVINGDR
2112 ITPVHVMSK 2162 LALYYPSAR 2212 GVVREFLTR
2113 PQNAVVKIY 2163 RLIIRENNR 2213 KRGDKSVYY
2114 RKAVFISPY 2164 YQIGGYTEK 2214 KHIDAYKTF
2115 SRIIPARAR 2165 NNAAIVLQL 2215 TAHSCNVNR
2116 PYPDPSRIL 2166 LAVVVCNSL 2216 FACVVADAV
2117 TIYSLLLCR 2167 VRQCSGVTF 2217 SSRSRNSSR
2118 VPATVSVSS 2168 SQLDEEQPM 2218 LATHGLAAV
2119 CSMTDIAKK 2169 GQQQQGQTV 2219 RWVLNNDYY
2120 VLHDIGNPK 2170 CTNYMPYFF 2220 LEKMADQAM
2121 RVVISSDVL 2171 GLFCLLNRY 2221 LIVTTIVYL
2122 FLYIIKLIF 2172 NLDSCKRVL 2222 SIDAFKLNI
2123 VESSSKLWA 2173 ATTAYANSV 2223 DIADTTDAV
2124 NLVAVPTGY 2174 LGDIAARDL 2224 SNFGAISSV
2125 LSAQTGIAV 2175 VEIIKSQDL 2225 YNMMISAGF
2126 MLKTVYSDV 2176 TEVNEFACV 2226 ALNLGETFV
2127 IPLTTAAKL 2177 VYRGTTTYK 2227 NLGERVRQA
2128 EQWNLVIGF 2178 NPTIQKDVL 2228 QLSLPVLQV
2129 MCDIRQLLF 2179 RHVRAWIGF 2229 ILFLALITL
2130 VQAGNVQLR 2180 RINWITGGI 2230 QPRVEKKKL
2131 GLQPSVGPK 2181 TLQCIMLVY 2231 VPVSIINNT
2132 QIGEYTFEK 2182 DIASTDTCF 2232 DFVKATCEF
2133 SSVELKHFF 2183 NVVIKVCEF 2233 LEKCDLQNY
2134 YNSASFSTF 2184 RDLICAQKF 2234 CLTPVYSFL
2135 NVVTTKIAL 2185 LIPLMYKGL 2235 SLNVAKSEF
2136 QIGGYTEKW 2186 CFANKHADF 2236 ADIVVFDEI
2137 TSQWLTNIF 2187 NSWLMWLII 2237 SNYQHEETI
2138 SLPFGWLIV 2188 WTNAGDYIL 2238 FSYVGCHNK
2139 KPSFYVYSR 2189 ALRANSAVK 2239 EPEFYEAMY
2140 TTNGDFLHF 2190 LFVTVYSHL 2240 FVSGNCDVV
2141 VVAFNTLLF 2191 ALLSDLQDL 2241 TPVYSFLPG
2142 ASLNGVTLI 2192 NYQKVGMQK 2242 WLIVGVALL
2143 WICLLQFAY 2193 DTKFKTEGL 2243 NVTQAFGRR
2144 GPEHSLAEY 2194 TEKSNIIRG 2244 KKLDGFMGR
2145 IFVDGVPFV 2195 GDELKINAA 2245 FEKMVSLLS
2146 VTSNYSGVV 2196 YNLWNTFTR 2246 GMPSYCTGY
2147 CVEEVTTTL 2197 TNNLVVMAY 2247 LEILDITPC
2148 IRQEEVQEL 2198 LLSDLQDLK 2248 VLKKCKSAF
2149 NFNKDFYDF 2199 WSFNPETNI 2249 QEGVLTAVV
2150 YPVASPNEC 2200 PYNSVTSSI 2250 LTLAVPYNM
2251 SEQLDFIDT 2301 DATPSDFVR 2351 LYKMQRMLL
2252 ELSLIDFYL 2302 NRYLALYNK 2352 ALYYPSARI
2253 LSKGRLIIR 2303 ASPNECNQM 2353 TSEAVEAPL
2254 GVKDCVVLH 2304 GRTILGSAL 2354 CIMSDRDLY
2255 AGILIVTTI 2305 VLQKAAITI 2355 AIVVTCLAY
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
2256 SEMVMCGGS 2306 WTLMNVLTL 2356 VQPQLEMEL
2257 WEVGKPRPP 2307 GTTLPKGFY 2357 RYVLMDGSI
2258 QITISSFKW 2308 CTTIVNGVR 2358 EWFLAYILF
2259 YTSALLAGT 2309 LHSYFTSDY 2359 VLITEGSVK
2260 FADDLNQLT 2310 SLSSTASAL 2360 TTCFSVAAL
2261 TFVTHSKGL 2311 ELEGIQYGR 2361 GANKDGIIW
2262 RAFDIYNDK 2312 LTSQWLTNI 2362 SSQGSEYDY
2263 EFLTRNPAW 2313 NVYLAVFDK 2363 MLWCKDGHV
2264 HLAKALNDF 2314 YSHLLLVAA 2364 EVVLKTGDL
2265 VVDADSKIV 2315 VVQEGVLTA 2365 SLRVCVDTV
2266 WKCRSKNPL 2316 PYRVVVLSF 2366 CFVLAAVYR
2267 SYGADLKSF 2317 YELKHGTFT 2367 VRGTTVLLK
2268 MADSNGTIT 2318 FVFLVLLPL 2368 VQHMVVKAA
2269 AENVTGLFK 2319 TVQEIQLQA 2369 IPARARVEC
2270 EEVGHTDLM 2320 LYRKCVKSR 2370 VLLRKNGNK
2271 MLIIFWFSL 2321 SAGFSLWVY 2371 FASFYYVWK
2272 CRMNSRNYI 2322 IIQFPNTYL 2372 GDFLHFLPR
2273 KQLSSNFGA 2323 EFSSNVANY 2373 SICSTMTNR
2274 LVGLMWLSY 2324 SKSLTENKY 2374 TGVEHVTFF
2275 NEFACVVAD 2325 AHFPREGVF 2375 YIIKLIFLW
2276 TLACFVLAA 2326 FLPFQQFGR 2376 FGAGAALQI
2277 SVQTFFKLV 2327 NSPRRARSV 2377 EAACCHLAK
2278 DEPTTTTSV 2328 KGFCKLHNW 2378 LPVLQVRDV
2279 TSCCFSERF 2329 SMQNCVLKL 2379 LSLQFKRPI
2280 VSTQEFRYM 2330 ELKKLLEQW 2380 NPTDQSSYI
2281 VIPDYNTYK 2331 VVTTFDSEY 2381 KGDYGDAVV
2282 TMCDIRQLL 2332 RKVQHMVVK 2382 CSQHTMLVK
2283 RGWIFGTTL 2333 RPNFTIKGS 2383 WADNNCYLA
2284 CQVHGNAHV 2334 NASFDNFKF 2384 QFAYANRNR
2285 KFVCDNIKF 2335 VTTIVYLTI 2385 LRANSAVKL
2286 QEFKPRSQM 2336 NSRNYIAQV 2386 ATNNLVVMA
2287 QPSVGPKQA 2337 RRIRGGDGK 2387 LFLPFFSNV
2288 IISDMYDPK 2338 IEVQGYKSV 2388 ETKFLTENL
2289 LALCADSII 2339 AVLDMCASL 2389 TYVTQQLIR
2290 YPDKVFRSS 2340 ADAQSFLNR 2390 RLNQLESKM
2291 NPDILRVYA 2341 FKLKDCVMY 2391 ASSSEAFLI
2292 FKEGSSVEL 2342 YYLGTGPEA 2392 VLSFELLHA
2293 VIGFLFLTW 2343 CAFAVDAAK 2393 SVAYSNNSI
2294 VTLAILTAL 2344 FTEQPIDLV 2394 SLREVRTIK
2295 MNVLTLVYK 2345 CRFVTDTPK 2395 ALILAYCNK
2296 TSAMQTMLF 2346 ITPCSFGGV 2396 HVMSKHTDF
2297 GIYQTSNFR 2347 VSLVKPSFY 2397 GGDAALALL
2298 LLALHFLLF 2348 ALGGSVAIK 2398 YQTQTNSPR
2299 FSHSQLGGL 2349 CPRHVICTS 2399 VIAWNSNNL
2300 GAWNIGEQK 2350 STGYHFREL 2400 YYRRATRRI
2401 VLKKLKKSL 2451 ARGIVFMCV 2501 HGFELTSMK
2402 IDHPNPKGF 2452 LNVPLHGTI 2502 MIDVQQWGF
2403 LSLREVRTI 2453 FYLCFLAFL 2503 CRKVQHMVV
2404 SVTSNYSGV 2454 NLVAMKYNY 2504 YFLQSINFV
2405 SSSSDNIAL 2455 VGFTLKNTV 2505 MMPTIFFAG
2406 FCSQHTMLV 2456 CTSVVLLSV 2506 NPYVCNAPG
2407 FLGIITTVA 2457 TSAFVETVK 2507 MHANYIFWR
2408 VVNPVMEPI 2458 VVHNQDVNL 2508 ESPFVMMSA
2409 GACIRRPFL 2459 LPVSMTKTS 2509 LMWLIINLV
2410 FLGYFCTCY 2460 TEHSWNADL 2510 GLWLDDVVY
2411 TITVNVLAW 2461 CISTKHFYW 2511 CYGVSPTKL
2412 MRSLKVPAT 2462 LLAKDTTEA 2512 YQCAMRPNF
2413 LKLFAAETL 2463 SLCLQLAVV 2513 HFVCNLLLL
2414 GVVTTVMFL 2464 IVNNWLKQL 2514 VQMLSDTLK
2415 TANKWDLII 2465 HVASCDAIM 2515 FFSNVTWFH
2416 LPDDFTGCV 2466 VTSAMQTML 2516 TFLLNKEMY
2417 KENDSKEGF 2467 TVLSFCAFA 2517 ARHINAQVA
2418 LPFFYYSDS 2468 NSVLLFLAF 2518 RIDKVLNEK
2419 KLDGFMGRI 2469 IYSLLLCRM 2519 SIVCRFDTR
2420 KACPLIAAV 2470 NVNRFNVAI 2520 SLLSKGRLI
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
2421 LRKVPTDNY 2471 ALAPNMMVT 2521 FMCVEYCPI
2422 TRPLLESEL 2472 FSLWVYKQF 2522 EYRLYLDAY
2423 ISMATNYDL 2473 NARDGCVPL 2523 SVTTEILPV
2424 TAGAAAYYV 2474 FLELAMDEF 2524 KTLNSLEDK
2425 LALYNKYKY 2475 ITSGWTFGA 2525 VFKNIDGYF
2426 LLQFAYANR 2476 ITVNVLAWL 2526 TGNYQCGHY
2427 RILGAGCFV 2477 VLLSMQGAV 2527 GTILTRPLL
2428 KQLPFFYYS 2478 LTLGVYDYL 2528 FAMGIIAMS
2429 ELTSMKYFV 2479 RTCCLCDRR 2529 VESCGNFKV
2430 SSRLSFKEL 2480 YPQVNGLTS 2530 VRTNVYLAV
2431 LDYIINLII 2481 RVDFCGKGY 2531 YEDQDALFA
2432 TGPEAGLPY 2482 TTLPKGFYA 2532 TNDKACPLI
2433 GDSEVVLKK 2483 FVENPDILR 2533 NFKDQVILL
2434 YVRNLQHRL 2484 TETAHSCNV 2534 HPLADNKFA
2435 SREETGLLM 2485 LVTMPLGYV 2535 DYVYNPFMI
2436 MLAHAEETR 2486 FSTGVNLVA 2536 FLMSFTVLC
2437 KEGFFTYIC 2487 GCDGGSLYV 2537 ELTGHMLDM
2438 QQESPFVMM 2488 RENNRVVIS 2538 KHYTPSFKK
2439 KITEHSWNA 2489 QPYRVVVLS 2539 FHPLADNKF
2440 TPFEIKLAK 2490 TVVVNAANV 2540 DYDYYRYNL
2441 DPSFLGRYM 2491 VLLFLAFVV 2541 ELYSPIFLI
2442 LLFLAFVVF 2492 VYEKLKPVL 2542 RNLQEFKPR
2443 ITREVGFVV 2493 FPNITNLCP 2543 DIGNPKAIK
2444 GLCVDIPGI 2494 DGYVMHANY 2544 VAKHDFFKF
2445 TQDLFLPFF 2495 QILFALLQR 2545 LVDPQIQLA
2446 FVTHSKGLY 2496 KSDGTGTIY 2546 CWHTNCYDY
2447 WVMRIMTWL 2497 SSRGTSPAR 2547 DRYPANSIV
2448 RFLYIIKLI 2498 NVAKSEFDR 2548 RQRLTKYTM
2449 FNVLFSTVF 2499 FTSLEIPRR 2549 LFTRFFYVL
2450 SEAARVVRS 2500 CLEASFNYL 2550 NFTIKGSFL
2551 NRVCGVSAA 2601 VAGFAKFLK 2651 AENSVAYSN
2552 VFQTRAGCL 2602 TYLEGSVRV 2652 FLPFFSNVT
2553 NVAITRAKV 2603 ASEAARVVR 2653 HPALRLVDP
2554 IDFYLCFLA 2604 ITGGIAIAM 2654 FNEKTHVQL
2555 AYIICISTK 2605 DLDDFSKQL 2655 LVKQLSSNF
2556 YVQIPTTCA 2606 PFFSNVTWF 2656 FLKRGDKSV
2557 GFIQQKLAL 2607 NVYADSFVI 2657 NLNESLIDL
2558 KSILSPLYA 2608 NQPYPNASF 2658 QLHNDILLA
2559 FCLLNRYFR 2609 HISRQRLTK 2659 GYLQPRTFL
2560 LALHFLLFF 2610 RRATCFSTA 2660 NASSSEAFL
2561 CLAVHECFV 2611 FVVFLLVTL 2661 FITESKPSV
2562 AMAVMLLLL 2612 LAMDEFIER 2662 RELGVVHNQ
2563 ALKYLPIDK 2613 KPFERDIST 2663 TFTRSTNSR
2564 DELKINAAC 2614 TPLVPFWIT 2664 MFITREEAI
2565 GVVQLTSQW 2615 ILAYCNKTV 2665 DAVRDPQTL
2566 DDTLRVEAF 2616 YQCGHYKHI 2666 ATATIPIQA
2567 RFTTTLNDF 2617 FAWWTAFVT 2667 SDRELHLSW
2568 RNPANNAAI 2618 IWLGFIAGL 2668 EWSMATYYL
2569 SWNADLYKL 2619 GGNYNYLYR 2669 VGPKVYPII
2570 TASWFTALT 2620 FISTCACEI 2670 ETGLLMPLK
2571 LSWEVGKPR 2621 FFSNYLKRR 2671 IINLIIKNL
2572 SSGDATTAY 2622 EVVLKKLKK 2672 SINFVRIIM
2573 LFWNCNVDR 2623 QLGGLHLLI 2673 RPPLNRNYV
2574 RKMAFPSGK 2624 LKVGGSCVL 2674 DGISQYSLR
2575 RLTLGVYDY 2625 RANNTKGSL 2675 TFISDEVAR
2576 LIVAAIVFI 2626 KYCALAPNM 2676 RRAFGEYSH
2577 DVEGCHATR 2627 SLVVRCSFY 2677 TTEILPVSM
2578 VPLNIIPLT 2628 VRMYIFFAS 2678 RQGFVDSDV
2579 HFLPRVFSA 2629 HVICTSEDM 2679 MLVKQGDDY
2580 WPQIAQFAP 2630 RKHFSMMIL 2680 MPNLYKMQR
2581 MGHFAWWTA 2631 KESVQTFFK 2681 HHANEYRLY
2582 VQSTQWSLF 2632 SPLSLNMAR 2682 VAPGTAVLR
2583 RLCAYCCNI 2633 PQADVEWKF 2683 FELWAKRNI
2584 DVNCTEVPV 2634 VQQLPETYF 2684 SRIKASMPT
2585 LTVFFDGRV 2635 TEILPVSMT 2685 GEIKDATPS
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
2586 KLHDELTGH 2636 NFVFPLNSI 2686 DKRAKVTSA
2587 NTLQCIMLV 2637 YSYATHSDK 2687 TNSFTRGVY
2588 VDDPCPIHF 2638 MLNPNYEDL 2688 KWPWYIWLG
2589 IIWFLLLSV 2639 AFGEYSHVV 2689 TCFSTQFAF
2590 WESGVKDCV 2640 FATSACVLA 2690 YLGKPREQI
2591 AELEGIQYG 2641 MAYCWRCTS 2691 SEDNQTTTI
2592 TPKDHIGTR 2642 GFELTSMKY 2692 LRARSVSPK
2593 LATVAYFNM 2643 QKKQQTVTL 2693 CEFQFCNDP
2594 MRFRRAFGE 2644 NKYKYFSGA 2694 IGFLFLTWI
2595 LTRNPAWRK 2645 FAQVKQIYK 2695 KSPIQYIDI
2596 KEIKESVQT 2646 HFAWWTAFV 2696 YGDSATLPK
2597 LFARTRSMW 2647 NRKRISNCV 2697 AFVETVKGL
2598 LRKQIRSAA 2648 SSQAWQPGV 2698 TPTWRVYST
2599 MPLGYVTHG 2649 SQCVNLTTR 2699 CGYLPQNAV
2600 MELTPVVQT 2650 RFQNHNPQK 2700 EAPLVGTPV
2701 KAALLADKF 2751 ALQDAYYRA 2801 GSSGVVNPV
2702 LRKLDNDAL 2752 FAAETLKAT 2802 KRSFIEDLL
2703 VSALVYDNK 2753 TLQPVSELL 2803 AYRKVLLRK
2704 ATNYDLSVV 2754 LACFVLAAV 2804 CEKALKYLP
2705 RTLLTKGTL 2755 WNVKDFMSL 2805 VAKNLNESL
2706 YPKCDRAMP 2756 LLEWLAMAV 2806 FLALITLAT
2707 NYVFTGYRV 2757 IFGTVYEKL 2807 AVFQSASKI
2708 VDTVRTNVY 2758 YDDGARRVW 2808 KVEGCMVQV
2709 NGDSEVVLK 2759 QPTSEAVEA 2809 CVDTVRTNV
2710 YLYLTFYLT 2760 WAKRNIKPV 2810 HAFLCLFLL
2711 IVNNATNVV 2761 IRASANLAA 2811 EDQDALFAY
2712 SAPLIELCV 2762 SERFQNHNP 2812 QIVESCGNF
2713 NVGPKVYPI 2763 VSKVVKVTI 2813 HSQLGGLHL
2714 TGYKKPASR 2764 TVCTVCGMW 2814 GALDISASI
2715 SVGPKQASL 2765 SSKCVCSVI 2815 LYDANYFLC
2716 SQDLSVVSK 2766 FEKGDYGDA 2816 AAMQRKLEK
2717 SLINTLNDL 2767 MVRMYIFFA 2817 VFIKRSDAR
2718 NLREMLAHA 2768 VYANLGERV 2818 GLALYYPSA
2719 QYGSFCTQL 2769 FCDLKGKYV 2819 LQPRTFLLK
2720 EMLDNRATL 2770 VSIWNLDYI 2820 NGMNGRTIL
2721 CPLIAAVIT 2771 FLIVAAIVF 2821 GRFVLALLS
2722 FMSLSEQLR 2772 NYLYRLFRK 2822 KMFDAYVNT
2723 TIFKDASGK 2773 FVDRQTAQA 2823 ATRGATVVI
2724 DALCEKALK 2774 TAFVTNVNA 2824 FGGCVFSYV
2725 APHGHVMVE 2775 KLKTLVATA 2825 VRDVLVRGF
2726 LPKGFYAEG 2776 NNCYLATAL 2826 AGTDTTITV
2727 LKQLIKVTL 2777 QQTVTLLPA 2827 KGTHHWLLL
2728 TQALPQRQK 2778 TNNVAFQTV 2828 VDWTIEYPI
2729 DFLEYHDVR 2779 IAYTMSLGA 2829 AMACLVGLM
2730 ISEMHPALR 2780 KEMYLKLRS 2830 GVTLIGEAV
2731 LRVESSSKL 2781 SWMESEFRV 2831 AQEAYEQAV
2732 DHIGTRNPA 2782 AKHDFFKFR 2832 GDCLGDIAA
2733 VTTVMFLAR 2783 CKDGHVETF 2833 GSKSPIQYI
2734 TWLTYTGAI 2784 KEGQINDMI 2834 SGLKTILRK
2735 VMFLARGIV 2785 MVLGSLAAT 2835 QLTPIAVQM
2736 TPFDVVRQC 2786 LLSVCLGSL 2836 DMSKFPLKL
2737 ILASFSAST 2787 SCDQLREPM 2837 VYDNKLKAH
2738 LGFSTGVNL 2788 VAGGIVAIV 2838 VRVLQKAAI
2739 FVCNLLLLF 2789 QWLPTGTLL 2839 PFAMQMAYR
2740 TNILLNVPL 2790 GRLQSLQTY 2840 SAYENFNQH
2741 RDAAMQRKL 2791 SAVVLLILM 2841 NLQSNHDLY
2742 NRARTVAGV 2792 SEKSYELQT 2842 FYSKWYIRV
2743 QESFGGASC 2793 AMKYNYEPL 2843 GLVEVEKGV
2744 YDKLQFTSL 2794 TLKNLSDRV 2844 NIVTRCLNR
2745 YYKKVDGVV 2795 DFSSEIIGY 2845 CEEMLDNRA
2746 LTQDHVDIL 2796 DPAQLPAPR 2846 AAITILDGI
2747 YICGFIQQK 2797 NKDFYDFAV 2847 LLSVLQQLR
2748 AQTGSSKCV 2798 YWEPEFYEA 2848 GVLTESNKK
2749 MAPISAMVR 2799 QLNRALTGI 2849 DTTITVNVL
2750 NPANNAAIV 2800 LLKDCPAVA 2850 FLFLTWICL
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
2851 MPTTIAKNT 2901 TICAPLTVF 2951 CIRCLWSTK
2852 NIVNVSLVK 2902 SATLPKGIM 2952 NTLLFLMSF
2853 ISPYNSQNA 2903 VVTCLAYYF 2953 LIRKSNHNF
2854 GYQPYRVVV 2904 VFHLYLQYI 2954 IVNGVRRSF
2855 YMRSLKVPA 2905 DEDDSEPVL 2955 NVSLVKPSF
2856 VATVQEIQL 2906 KKADETQAL 2956 LLEDEFTPF
2857 ICAPLTVFF 2907 APGQTGKIA 2957 VFMSEAKCW
2858 NQFNSAIGK 2908 HRLYECLYR 2958 IVESCGNFK
2859 NLLLLFVTV 2909 FPKSDGTGT 2959 KPTETICAP
2860 NFVRIIMRL 2910 TETDLTKGP 2960 FYEDFLEYH
2861 KPLEFGATS 2911 PPISFPLCA 2961 TTEELPDEF
2862 MQNCVLKLK 2912 LTENKYSQL 2962 KQIRSAAKK
2863 FFKLVNKFL 2913 RSGETLGVL 2963 DMILSLLSK
2864 LVQAGNVQL 2914 QPYVVDDPC 2964 MLLLLCCCL
2865 FTEERLKLF 2915 EKTHVQLSL 2965 NLYKMQRML
2866 VFTGYRVTK 2916 QKLLKSIAA 2966 GPLVRKIFV
2867 CFSTASDTY 2917 SSRSSSRSR 2967 SFGGASCCL
2868 VLPQLEQPY 2918 QAAVGELLL 2968 CHNKCAYWV
2869 MAFPSGKVE 2919 VTDFNAIAT 2969 NSVFNICQA
2870 HTANKWDLI 2920 KSREETGLL 2970 VANGDSEVV
2871 LANTCTERL 2921 RDLSLQFKR 2971 QLLFVVEVV
2872 DVNLHSSRL 2922 EHSWNADLY 2972 ALRLVDPQI
2873 KKLLEQWNL 2923 EAPFLYLYA 2973 ITEGSVKGL
2874 LELAMDEFI 2924 DAFKLNIKL 2974 GFFTYICGF
2875 TDTPKGPKV 2925 TYKLNVGDY 2975 CSHAAVDAL
2876 SGAMDTTSY 2926 VGARKSAPL 2976 DKRTTCFSV
2877 FWNCNVDRY 2927 KRRVVFNGV 2977 SGKPVPYCY
2878 AVANGDSEV 2928 TSVVLLSVL 2978 FCGPDGYPL
2879 DVVYCPRHV 2929 VFQSASKII 2979 SVEEVLSEA
2880 SDVENPHLM 2930 CANGQVFGL 2980 GIMMNVAKY
2881 SEEVVENPT 2931 KAGQKTYER 2981 EELPDEFVV
2882 AQEKNFTTA 2932 HLLIGLAKR 2982 GYVMHANYI
2883 YKLNVGDYF 2933 EIDFLELAM 2983 HLSVDTKFK
2884 YNGSPSGVY 2934 FYEAMYTPH 2984 STIGVCSMT
2885 AMMFVKHKH 2935 ITIAYIICI 2985 VLTLVYKVY
2886 NRPQIGVVR 2936 STGVNLVAV 2986 LLVPHHVVA
2887 CDQLREPML 2937 RALTAESHV 2987 DYDCVSFCY
2888 SEHDYQIGG 2938 AVVKIYCPA 2988 SALVYDNKL
2889 ITVEELKKL 2939 YYNTTKGGR 2989 HVETFYPKL
2890 FKELLVYAA 2940 TACTDDNAL 2990 NRYFRLTLG
2891 IPKDMTYRR 2941 YVLMDGSII 2991 YHNESGLKT
2892 DIVKTDGTL 2942 QVDLFRNAR 2992 YRINWITGG
2893 TLLSLREVR 2943 VECFDKFKV 2993 FMRIFTIGT
2894 NWYDFGDFI 2944 LPWNVVRIK 2994 NLVIGFLFL
2895 LVIGFLFLT 2945 ECAQVLSEM 2995 NIKPVPEVK
2896 RDAPAHIST 2946 HPNQEYADV 2996 AMMFTSDLA
2897 LGYFCTCYF 2947 DFYLCFLAF 2997 VAIVVTCLA
2898 LLLCRMNSR 2948 TIFFAGILI 2998 KVITGLHPT
2899 FNYLKSPNF 2949 VTLIGEAVK 2999 YTVSCLPFT
2900 TREAVGTNL 2950 ALTCFSTQF 3000 SKQRRPQGL
3001 EVTPSGTWL 3051 NWITGGIAI 3101 LSMQGAVDI
3002 FFYYSDSPC 3052 RVVFVLWAH 3102 HQKLLKSIA
3003 VYAWNRKRI 3053 TALRANSAV 3103 LTWICLLQF
3004 LEPLVDLPI 3054 HLLLVAAGL 3104 PHSLSDGLL
3005 GTPVCINGL 3055 YLGGMSYYC 3105 YGNALDQAI
3006 LLKSAYENF 3056 FPSGKVEGC 3106 ASKKPRQKR
3007 RQMSCAAGT 3057 IPCTCGKQA 3107 LYECLYRNR
3008 ENAFLPFAM 3058 ARLRAKHYV 3108 EGFNCYFPL
3009 SLRCGACIR 3059 TKLATTEEL 3109 ERLKLFDRY
3010 DYTEISFML 3060 CVCSVIDLL 3110 SSGVVNPVM
3011 KVTSAMQTM 3061 LRTTNGDFL 3111 LQDLKWARF
3012 RKHTTCCSL 3062 NINIVGDFK 3112 RNSTPGSSR
3013 EVAKNLNES 3063 LKLRGTAVM 3113 SHMYCSFYP
3014 VQELYSPIF 3064 MMFTSDLAT 3114 CSLCLQLAV
3015 LLFLVLIML 3065 AAVYRINWI 3115 SEDKRAKVT
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
3016 LPQGTTLPK 3066 YAISAKNRA 3116 PPTSFGPLV
3017 TEVLTEEVV 3067 FGGASCCLY 3117 KGLDYKAFK
3018 IWVATEGAL 3068 FITLCFTLK 3118 TQKGAEAAV
3019 LLQILFALL 3069 LEGYAFEHI 3119 IERYKLEGY
3020 LLVAAGLEA 3070 IIFLEGETL 3120 ILRTTNGDF
3021 YNVTQAFGR 3071 MSFTVLCLT 3121 AVTANVNAL
3022 ILANTCTER 3072 RAKVGILCI 3122 GFCDLKGKY
3023 LECIKDLLA 3073 GSVGFNIDY 3123 GDYGDAVVY
3024 LVRGFGDSV 3074 ILLIIMRTF 3124 DAIMTRCLA
3025 VLTEEVVLK 3075 FDTRVLSNL 3125 LSLPVLQVR
3026 LKFPRGQGV 3076 LKLRSDVLL 3126 GDAALALLL
3027 DGVKHVYQL 3077 IILRLGSPL 3127 TRNPANNAA
3028 LPGTILRTT 3078 EGYAFEHIV 3128 GVGGKPCIK
3029 TIQPRVEKK 3079 SDRDLYDKL 3129 LPPKNSIDA
3030 FSNVTWFHA 3080 QYGRSGETL 3130 LVQSTQWSL
3031 QASSRSSSR 3081 LEGSVRVVT 3131 YEGNSPFHP
3032 RRSFYVYAN 3082 MDTTSYREA 3132 WDYKRDAP
3033 QLRVIGHSM 3083 YRVTKNSKV 3133 IVVFDEISM
3034 EMHPALRLV 3084 GASCCLYCR 3134 SPVALRQMS
3035 LSSYSLFDM 3085 LDISASIVA 3135 QQQGQTVT
3036 DKCSRIIPA 3086 CESHGKQVV 3136 LNNIINNAR
3037 APRITFGGP 3087 LNDNLLEIL 3137 QEPKLGSLV
3038 TSTLQGCSL 3088 FGPLVRKIF 3138 VPLHGTILT
3039 NLLLQYGSF 3089 IQWMVMFTP 3139 FLTRNPAWR
3040 TRMENAVGR 3090 KEPCSSGTY 3140 KGVAPGTAV
3041 PPQTSITSA 3091 ARSVSPKLF 3141 QLRVESSSK
3042 YRFNGIGVT 3092 VIDLLLDDF 3142 QEHYVRITG
3043 LLLLFVTVY 3093 PHGHVMVEL 3143 YMHHMELPT
3044 APLVGTPVC 3094 YPDPSRILG 3144 SEGLNDNLL
3045 FGLFCLLNR 3095 TQLSTDTGV 3145 AQSFLNGFA
3046 DAAKAYKDY 3096 LIRQGTDYK 3146 VCVSTSGRW
3047 AAKAYKDYL 3097 NSYECDIPI 3147 RTVYDDGAR
3048 NPIQLSSYS 3098 VGLMWLSYF 3148 SSKLWAQCV
3049 CDHCGETSW 3099 CQEPKLGSL 3149 TFNGECPNF
3050 NGVGYQPYR 3100 NSKVQIGEY 3150 VEGCMVQVT
3151 KVTKGKAK 3201 PVNVAFELW 3251 GSLVVRCSF
3152 YKLMGHFAW 3202 KLKKSLNVA 3252 QAGNVQLRV
3153 DAALALLLL 3203 LELQDHNET 3253 GKSHFAIGL
3154 NSVTSSIVI 3204 KHKHAFLCL 3254 ACPDGVKHV
3155 LVFLFVAAI 3205 PFTIYSLLL 3255 NKDGIIWVA
3156 VPTDNYITT 3206 QTSITSAVL 3256 CAQVLSEMV
3157 DTGVEHVTF 3207 KFDTFNGEC 3257 ALLSTDGNK
3158 AAFATAQEA 3208 CWRCTSCCF 3258 VLLILMTAR
3159 TPSKLIEYT 3209 SVLYYQNNV 3259 DEAGSKSPI
3160 NVTDFNAIA 3210 LMKTIGPDM 3260 GPLSAQTGI
3161 LAVFQSASK 3211 LPTQTVDSS 3261 FLLFLVLIM
3162 CERSEAGVC 3212 DTSLSGFKL 3262 KCSRIIPAR
3163 IAEILLIIM 3213 FDKSAFVNL 3263 WYIRVGARK
3164 LSVVNARLR 3214 KGPKVKYLY 3264 FKLNEEIAI
3165 LEDKAFQLT 3215 YKFVRIQPG 3265 VFLLVTLAI
3166 LLLLEWLAM 3216 TETICAPLT 3266 GLHLLIGLA
3167 KYVQIPTTC 3217 LVPHHVVAT 3267 LYHYQECVR
3168 LITLATCEL 3218 YFNMVYMPA 3268 DAYKTFPPT
3169 KNIDGYFKI 3219 IGDPAQLPA 3269 FVRIIMRLW
3170 RKSNHNFLV 3220 YTEKWESGV 3270 MVVKAALLA
3171 FGWLIVGVA 3221 GMVLGSLAA 3271 FLWLLWPVT
3172 SKIVQLSEI 3222 RSEDKRAKV 3272 ILRGHLRIA
3173 SQGSEYDYV 3223 NAQALNTLV 3273 YEDLLIRKS
3174 KRVLNVVCK 3224 YNKYKYFSG 3274 VVRIKIVQM
3175 TEAFEKMVS 3225 TPCGTGTST 3275 ELFYSYATH
3176 LGSPLSLNM 3226 CAAGTTQTA 3276 DVVRQCSGV
3177 KQVVSDIDY 3227 VYTACSHAA 3277 DRRATCFST
3178 ITVTPEANM 3228 IFFAGILIV 3278 NSQNAVASK
3179 DKFTDGVCL 3229 VPANSTVLS 3279 DTNVLEGSV
3180 GIVFMCVEY 3230 LRLIDAMMF 3280 DVKCTSVVL
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
3181 IIKNLSKSL 3231 HTNCYDYCI 3281 IVQMLSDTL
3182 HVQLSLPVL 3232 QPGVAMPNL 3282 PFMIDVQQW
3183 TEELPDEFV 3233 VAYSNNSIA 3283 NGVRRSFYV
3184 YVFTGYRVT 3234 RSGARSKQR 3284 MQGAVDINK
3185 AIVFITLCF 3235 ILRKGGRTI 3285 GPEQTQGNF
3186 GRSGETLGV 3236 CELYHYQEC 3286 NLIIKNLSK
3187 LTGHMLDMY 3237 IQPIGALDI 3287 SFYVYSRVK
3188 TLIVNSVLL 3238 RVECFDKFK 3288 LLIGLAKRF
3189 VILRGHLRI 3239 LADNKFALT 3289 SHAAVDALC
3190 SSAINRPQI 3240 SNIIRGWIF 3290 IKDFGGFNF
3191 RGTAVMSLK 3241 APRTLLTKG 3291 SLNMARKTL
3192 VLQVRDVLV 3242 ASIVAGGIV 3292 IKGTHHWLL
3193 KSTNLVKNK 3243 VVTTVMFLA 3293 VRITGLYPT
3194 SKEGFFTYI 3244 FLKEQHCQK 3294 NKSWMESEF
3195 VGEIPVAYR 3245 GVCVSTSGR 3295 RAGKASCTL
3196 AQPCSDKAY 3246 ESELVIGAV 3296 LGIITTVAA
3197 GTGTIYTEL 3247 NPKAIKCVP 3297 YANSVFNIC
3198 ELVAELEGI 3248 ELKFNPPAL 3298 AYCCNIVNV
3199 TLCFTLKRK 3249 GHHLGRCDI 3299 VPINTNSSP
3200 PLMYKGLPW 3250 YFCTCYFGL 3300 HKDKSAQCF
3301 FYYLGTGPE 3351 QRVAGDSGF 3401 CAKEIKESV
3302 KCDRAMPNM 3352 YSPIFLIVA 3402 IIPLTTAAK
3303 LLEQWNLVI 3353 VPRASANIG 3403 LLSKGRLII
3304 SFGPLVRKI 3354 ISSVLNDIL 3404 ECPNFVFPL
3305 TGLYPTLNI 3355 GLTVLPPLL 3405 GDYILANTC
3306 TVCGMWKGY 3356 IVQLSEISM 3406 SGWTAGAAA
3307 LLQNGMNGR 3357 TLETAQNSV 3407 LVYCFLGYF
3308 TEIYQAGST 3358 LFLTWICLL 3408 QRNFYEPQI
3309 FQQFGRDIA 3359 FLIGCNYLG 3409 VRAWIGFDV
3310 SQAWQPGVA 3360 SQWLTNIFG 3410 FPVLHDIGN
3311 LITGRLQSL 3361 LCDRRATCF 3411 CTLSEQLDF
3312 LMPILTLTR 3362 LLWPVTLAC 3412 QENWNTKHS
3313 SIKNFKSVL 3363 ICISTKHFY 3413 ICYTPSKLI
3314 VELFENKTT 3364 DFATSACVL 3414 TEGLCVDIP
3315 DPNFKDQVI 3365 QAAGTDTTI 3415 TDFVNEFYA
3316 RPQGLPNNT 3366 QSRNLQEFK 3416 QMEIDFLEL
3317 VTLKQGEIK 3367 CNDPFLGVY 3417 NSTLEQYVF
3318 LTPFARCCW 3368 GSSKCVCSV 3418 RPFLCCKCC
3319 RIRSVYPVA 3369 DDYFNKKDW 3419 QECSLQSCT
3320 MDSTVKNYF 3370 FITREEAIR 3420 DQAMTQMYK
3321 QEFRYMNSQ 3371 MCVEYCPIF 3421 AISAKNRAR
3322 IRENNRVVI 3372 LPTEVLTEE 3422 PANSTVLSF
3323 VGVALLAVF 3373 QRQKKQQTV 3423 KEQHCQKAS
3324 CLAYYFMRF 3374 DLTKPYIKW 3424 QVNGYPNMF
3325 YLKRRVVFN 3375 AISMWALII 3425 NKHADFDTW
3326 YCALAPNMM 3376 KSAPLIELC 3426 CYNGSPSGV
3327 IFLWLLWPV 3377 SETKCTLKS 3427 CNLGGAVCR
3328 CLFLLPSLA 3378 TGTLIVNSV 3428 ITRAKVGIL
3329 TESIVRFPN 3379 MFHLVDFQV 3429 YDFAVSKGF
3330 CTQHQPYVV 3380 SSGVTRELM 3430 CAYWVPRAS
3331 DGWEIVKFI 3381 VEGFNCYFP 3431 EGNSPFHPL
3332 IVGVALLAV 3382 EQYIKWPWY 3432 DEVRQIAPG
3333 VEWKFYDAQ 3383 LNEEIAIIL 3433 KAFQLTPIA
3334 LIGCNYLGK 3384 NRVVISSDV 3434 NEFYAYLRK
3335 TLQGCSLCL 3385 KLALGGSVA 3435 GPLKVGGSC
3336 APAHISTIG 3386 MRTFKVSIW 3436 LQFTSLEIP
3337 TEISFMLWC 3387 KGFYAEGSR 3437 AGNGGDAAL
3338 REFVFKNID 3388 IIRENNRVV 3438 FKLSYGIAT
3339 SGVYQCAMR 3389 LGTVSWNLR 3439 NGQVFGLYK
3340 KQQTVTLLP 3390 NHNFLVQAG 3440 NTDFSRVSA
3341 LFKDCSKVI 3391 QSCTQHQPY 3441 VPVVDSYYS
3342 RQKRTATKA 3392 SDFVRATAT 3442 RMNSRNYIA
3343 CSTMTNRQF 3393 LLPLTQYNR 3443 WPLIVTALR
3344 LEGSVAYES 3394 FIKQYGDCL 3444 ERSEAGVCV
3345 TILTSLLVL 3395 CVRGTTVLL 3445 RTRSMWSFN
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
3346 HGTILTRPL 3396 NYYKKVDGV 3446 MRVIHFGAG
3347 TCLAYYFMR 3397 TTTLNGLWL 3447 ACFVLAAVY
3348 VAALTNNVA 3398 QKEMATSTL 3448 KDLSPRWYF
3349 LMWLSYFIA 3399 VGELLLLEW 3449 QLQQSMSSA
3350 ALALLLLDR 3400 SMDNSPNLA 3450 LAVHECFVK
3451 SIINNTVYT 3501 ERVRQALLK 3551 GDDYVYLPY
3452 ESIVRFPNI 3502 QVVDADSKI 3552 FMIDVQQWG
3453 ACPLIAAVI 3503 EFLRDGWEI 3553 QPYVFIKRS
3454 RRATRRIRG 3504 CVLAAECTI 3554 SRILGAGCF
3455 QLQAAVGEL 3505 TPFARCCWP 3555 KDCVVLHSY
3456 KAKKGAWNI 3506 AACRKVQHM 3556 EPEEHVQIH
3457 VDRQTAQAA 3507 KSHFAIGLA 3557 KSFTVEKGI
3458 SPLYAFASE 3508 NRDVDTDFV 3558 YPNMFITRE
3459 TKRNVIPTI 3509 FKDASGKPV 3559 ILNNLGVDI
3460 SNSGSDVLY 3510 ELWAKRNIK 3560 KELLVYAAD
3461 SQTSLRCGA 3511 ALRQMSCAA 3561 QWNLVIGFL
3462 YNLPTMCDI 3512 ILSLLSKGR 3562 KPRPPLNRN
3463 EVLTEEVVL 3513 LVDFQVTIA 3563 SYSLFDMSK
3464 HGHVMVELV 3514 NGECPNFVF 3564 LNKEMYLKL
3465 GAEAAVKPL 3515 KPGNFNKDF 3565 DILGPLSAQ
3466 LKTGDLQPL 3516 IRRPFLCCK 3566 GTLIVNSVL
3467 FSTASDTYA 3517 CGACIRRPF 3567 RSLPGVFCG
3468 RFQTLLALH 3518 PQLEQPYVF 3568 NMFITREEA
3469 VRCSFYEDF 3519 NTTKGGRFV 3569 VTLACFVLA
3470 GARRVWTLM 3520 MVMFTPLVP 3570 YRYKPHSLS
3471 FMSEAKCWT 3521 ECIKDLLAR 3571 LVEVEKGVL
3472 LIGLAKRFK 3522 LSEARQHLK 3572 AARYMRSLK
3473 GAKLKALNL 3523 FGEYSHVVA 3573 TFFIYNKIV
3474 NEKTHVQLS 3524 LDKSAGFPF 3574 LRVCVDTVR
3475 GRYMSALNH 3525 TSWQTGDFV 3575 IPLMYKGLP
3476 WDLIISDMY 3526 DELTGHMLD 3576 SYAAFATAQ
3477 SVCRLMKTI 3527 RMLLEKCDL 3577 NGLWLDDVV
3478 TRCNLGGAV 3528 YTERSEKSY 3578 TSDYYQLYS
3479 RNYIAQVDV 3529 VDIAANTVI 3579 VSFSTFEEA
3480 KRTTCFSVA 3530 LNNDYYRSL 3580 TITVEELKK
3481 LSKSLTENK 3531 PAFDKSAFV 3581 MIELSLIDF
3482 KGIMMNVAK 3532 QGSEYDYVI 3582 LLPSLATVA
3483 YDYVIFTQT 3533 CVPQADVEW 3583 KFLVFLGII
3484 MAGNGGDAA 3534 LEQWNLVIG 3584 CNVNRFNVA
3485 SYYKLGASQ 3535 FDYVYNPFM 3585 RRARSVASQ
3486 EKFKEGVEF 3536 NKGAGGHSY 3586 LFLVLIMLI
3487 VTLLPAADL 3537 MKTIGPDMF 3587 AGSKSPIQY
3488 LPETYFTQS 3538 FFTYICGFI 3588 RWFLNRFTT
3489 SAMQTMLFT 3539 TDFATSACV 3589 FNVAITRAK
3490 NTCVGSDNV 3540 CRFDTRVLS 3590 VEKGIYQTS
3491 SLRLIDAMM 3541 SEAFLIGCN 3591 PPLNRNYVF
3492 KCYGVSPTK 3542 GNICYTPSK 3592 YDYCIPYNS
3493 LHFLLFFRA 3543 HKHAFLCLF 3593 ALCADSIII
3494 QASLPFGWL 3544 GDFKLNEEI 3594 QNYGDSATL
3495 LMGHFAWW 3545 TDDNALAYY 3595 LPFTINCQE
3496 YNSVTS SIV 3546 EIPVAYRKV 3596 SHLLLVAAG
3497 NNLVVMAYI 3547 TKYTMADLV 3597 CRHHANEYR
3498 LVPQEHYVR 3548 LYAFASEAA 3598 MDQESFGGA
3499 SVRVLQKAA 3549 PWNVVRIKI 3599 TPGSSRGTS
3500 QVTCGTTTL 3550 VPTGYVDTP 3600 RYPANSIVC
3601 PYEDFQENW 3651 YSDSPCESH 3701 AVNLLTNMF
3602 GARKSAPLI 3652 LTCFSTQFA 3702 YSKWYIRVG
3603 INNTVYTKV 3653 SARIVYTAC 3703 LSVVSKVVK
3604 FKPRSQMEI 3654 NYFLCWHTN 3704 LIMLIIFWF
3605 NSSRVPDLL 3655 MSDRDLYDK 3705 KINAACRKV
3606 ALVYDNKLK 3656 ISRQRLTKY 3706 GILCIMSDR
3607 SIFSRTLET 3657 HMELPTGVH 3707 IKVTLVFLF
3608 NNFCGPDGY 3658 EWLAMAVML 3708 LRCGACIRR
3609 FNAIATCDW 3659 QKSILSPLY 3709 VVQLTSQWL
3610 QLESKMSGK 3660 WNLDYIINL 3710 SQMEIDFLE
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
3611 QPIGALDIS 3661 IELCVDEAG 3711 YALVYFLQS
3612 DPCPIHFYS 3662 LFLAFVVFL 3712 YKLGASQRV
3613 CPNFVFPLN 3663 DQIGYYRRA 3713 VQEIQLQAA
3614 GSEYDYVIF 3664 KDMTYRRLI 3714 RGGSQASSR
3615 AMQRKLEKM 3665 KDGHVETFY 3715 DEPEEHVQI
3616 DGLLLALHF 3666 KDCVMYASA 3716 YTKVDGVDV
3617 ASFSASTSA 3667 YKRDAPAHI 3717 AKVGILCIM
3618 APATVCGPK 3668 VAGDSGFAA 3718 GIAIAMACL
3619 IGNYTVSCL 3669 IAVQMTKLA 3719 YRRATRRIR
3620 GEVPVSIIN 3670 LAVTRMENA 3720 NDVSFLAHI
3621 NNELSPVAL 3671 VGCHNKCAY 3721 LNKHIDAYK
3622 VKCTSVVLL 3672 RSQMEIDFL 3722 WLCWKCRSK
3623 FCAFAVDAA 3673 YAAFATAQE 3723 DNKFALTCF
3624 NHNPQKEMA 3674 TAQEAYEQA 3724 NMRVIHFGA
3625 VLIMLIIFW 3675 PEAGLPYGA 3725 FSTQFAFAC
3626 LKQLPFFYY 3676 IMMNVAKYT 3726 FNICQAVTA
3627 AVINGDRWF 3677 GSEGLNDNL 3727 RCLAVHECF
3628 YQLRARSVS 3678 CNIVNVSLV 3728 TPIAVQMTK
3629 RTTCFSVAA 3679 ILDGISQYS 3729 MESLVPGFN
3630 FHLVDFQVT 3680 WFLNRFTTT 3730 NLAWPLIVT
3631 RSDARTAPH 3681 EQIDGYVMH 3731 MRLWLCWKC
3632 YVFIKRSDA 3682 MHHMELPTG 3732 VTCLAYYFM
3633 RELMRELNG 3683 SDYYQLYST 3733 WNLVIGFLF
3634 FVTNVNASS 3684 ECVRGTTVL 3734 KQLQQSMSS
3635 DCVVLHSYF 3685 MNLKYAISA 3735 ISDMYDPKT
3636 PSLATVAYF 3686 KAIKCVPQA 3736 YQLYSTQLS
3637 VCNSLLTPF 3687 YDFGDFIQT 3737 LPTGTLLVD
3638 PFLYLYALV 3688 RNARNGVLI 3738 KQASLNGVT
3639 VDTVSALVY 3689 IVAAIVFIT 3739 EETIYNLLK
3640 WALIISVTS 3690 SGETLGVLV 3740 WLTYTGAIK
3641 VYSHLLLVA 3691 YDYYRYNLP 3741 PEVKILNNL
3642 YCKSHKPPI 3692 YAYLRKHFS 3742 SVLLSMQGA
3643 SPNECNQMC 3693 RVAGDSGFA 3743 FALTCFSTQ
3644 EETGTLIVN 3694 GIPKDMTYR 3744 KPVPYCYDT
3645 GLPNNTASW 3695 PMDSTVKNY 3745 TENKYSQLD
3646 VQMTKLATT 3696 VCRHHANEY 3746 RDGCVPLNI
3647 KEMATSTLQ 3697 YRRLISMMG 3747 PFLCCKCCY
3648 CATVHTANK 3698 PLADNKFAL 3748 CPIFFITGN
3649 LEQPYVFIK 3699 DFVRATATI 3749 TIKGTHHWL
3650 CSKVITGLH 3700 ASANIGCNH 3750 APLTVFFDG
3751 MVVIPDYNT 3801 KVTIDYTEI 3851 LRIAGHHLG
3752 LFRNARNGV 3802 NSTVLSFCA 3852 HIVYGDFSH
3753 LLEILQKEK 3803 YDCVSFCYM 3853 QLDFIDTKR
3754 EELKKLLEQ 3804 GTTFTYASA 3854 IPDYNTYKN
3755 TTTLNDFNL 3805 YSFVSEETG 3855 SNHNFLVQA
3756 EHDYQIGGY 3806 ESPFELEDF 3856 NKTTLPVNV
3757 GTAVLRQWL 3807 KTVQFCDAM 3857 WHTNCYDYC
3758 NPQKEMATS 3808 KTEGLCVDI 3858 QIAQFAPSA
3759 FARTRSMWS 3809 VDSSQGSEY 3859 FRLFARTRS
3760 EGNFYGPFV 3810 WITGGIAIA 3860 LEDFIPMDS
3761 GYRVTKNSK 3811 REGVFVSNG 3851 LRIAGHHLG
3762 EPIYDEPTT 3812 VYPIILRLG 3852 HIVYGDFSH
3763 VDRYPANSI 3813 IVTTIVYLT 3853 QLDFIDTKR
3764 NGRTILGSA 3814 WAHGFELTS 3854 IPDYNTYKN
3765 YHYQECVRG 3815 LDQAISMWA 3855 SNHNFLVQA
3766 AIRHVRAWI 3816 FNVVNKGHF 3856 NKTTLPVNV
3767 SLLTPFARC 3817 LLLFVTVYS 3857 WHTNCYDYC
3768 LVTLAILTA 3818 AWYTERSEK 3858 QIAQFAPSA
3769 GS SVELKHF 3819 VNGYPNMFI 3859 FRLFARTRS
3770 QLEQPYVFI 3820 TLVFLFVAA 3860 LEDFIPMDS
3771 SRGTSPARM 3821 MQRKLEKMA 3851 LRIAGHHLG
3772 SDEFSSNVA 3822 LCNSQTSLR 3852 HIVYGDFSH
3773 DDNLIDSYF 3823 ATPSDFVRA 3853 QLDFIDTKR
3774 LALLLLDRL 3824 TQFAFACPD 3854 IPDYNTYKN
3775 SQASSRSSS 3825 GAVILRGHL 3855 SNHNFLVQA
Seq. ID Peptide Seq. ID Peptide Seq. ID Peptide
3776 LSVLLSMQG 3826 ELLLLEWLA 3856 NKTTLPVNV
3777 NKCAYWVPR 3827 KHLIPLMYK 3857 WHTNCYDYC
3778 IPMDSTVKN 3828 LGGSVAIKI 3858 QIAQFAPSA
3779 FKLVNKFLA 3829 FHQECSLQS 3859 FRLFARTRS
3780 DTYACWHHS 3830 YCRHGTCER 3860 LEDFIPMDS
3781 TPVHVMSKH 3831 ELQTPFEIK    
3782 FDMSKFPLK 3832 DDPCPIHFY    
3783 TDFNAIATC 3833 CDWTNAGDY    
3784 RSMWSFNPE 3834 GSALLEDEF    
3785 HCANFNVLF 3835 SKWYIRVGA    
3786 RGMVLGSLA 3836 MAVMLLLLC    
3787 FRKMAFPSG 3837 KDFYDFAVS    
3788 ISASIVAGG 3838 SGVKDCVVL    
3789 FAGILIVTT 3839 AQSFLNRVC    
3790 YRSLPGVFC 3840 SSSEAFLIG    
3791 TLATCELYH 3841 GVEHVTFFI    
3792 NQDLNGNWY 3842 MLRKLDNDA    
3793 WWTAFVTNV 3843 CVEYCPIFF    
3794 LLFVTVYSH 3844 SPIQYIDIG    
3795 MDLFMRIFT 3845 LDMYSVMLT    
3796 VDTDLTKPY 3846 LNVGDYFVL    
3797 DQVILLNKH 3847 TRAGCLIGA    
3798 APSASAFFG 3848 MCASLKELL    
3799 IKDATPSDF 3849 GQTFSVLAC    
3800 RRVVFNGVS 3850 VMSLKEGQI    
Table 16. Region-specific peptide pools derived from whole proteome for CD8 T cell assays.
Figure PCTCN2022070948-appb-000034
Figure PCTCN2022070948-appb-000035
Figure PCTCN2022070948-appb-000036
Table 17. Region-specific peptide pools derived from S protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000037
Figure PCTCN2022070948-appb-000038
Figure PCTCN2022070948-appb-000039
Figure PCTCN2022070948-appb-000040
Table 18.. Region-specific peptide pools derived from NSP3 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000041
Figure PCTCN2022070948-appb-000042
Figure PCTCN2022070948-appb-000043
Figure PCTCN2022070948-appb-000044
Figure PCTCN2022070948-appb-000045
Table 19. Region-specific peptide pools derived from NSP4 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000046
Figure PCTCN2022070948-appb-000047
Figure PCTCN2022070948-appb-000048
Figure PCTCN2022070948-appb-000049
Table 20. Region-specific peptide pools derived from NSP14 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000050
Figure PCTCN2022070948-appb-000051
Figure PCTCN2022070948-appb-000052
Table 21. Region-specific peptide pools derived from NSP16 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000053
Figure PCTCN2022070948-appb-000054
Table 22. Region-specific peptide pools derived from N protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000055
Figure PCTCN2022070948-appb-000056
Figure PCTCN2022070948-appb-000057
Table 23. Region-specific peptide pools derived from M protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000058
Figure PCTCN2022070948-appb-000059
Table 24. Region-specific peptide pools derived from NSP5 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000060
Figure PCTCN2022070948-appb-000061
Table 25. Region-specific peptide pools derived from NSP9 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000062
Table 26. Region-specific peptide pools derived from NSP12 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000063
Figure PCTCN2022070948-appb-000064
Figure PCTCN2022070948-appb-000065
Figure PCTCN2022070948-appb-000066
Table 27. Region-specific peptide pools derived from NSP6 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000067
Figure PCTCN2022070948-appb-000068
Figure PCTCN2022070948-appb-000069
Table 28. Region-specific peptide pools derived from E protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000070
Figure PCTCN2022070948-appb-000071
Table 29. Region-specific peptide pools derived from NSP8 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000072
Figure PCTCN2022070948-appb-000073
Table 30. Region-specific peptide pools derived from ORF3a protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000074
Figure PCTCN2022070948-appb-000075
Figure PCTCN2022070948-appb-000076
Table 31. Region-specific peptide pools derived from NSP1 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000077
Figure PCTCN2022070948-appb-000078
Table 32. Region-specific peptide pools derived from ORF7a protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000079
Figure PCTCN2022070948-appb-000080
Table 33. Region-specific peptide pools derived from NSP2 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000081
Figure PCTCN2022070948-appb-000082
Figure PCTCN2022070948-appb-000083
Figure PCTCN2022070948-appb-000084
Table 34. Region-specific peptide pools derived from NSP13 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000085
Figure PCTCN2022070948-appb-000086
Figure PCTCN2022070948-appb-000087
Table 35. Region-specific peptide pools derived from ORF8 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000088
Table 36. Region-specific peptide pools derived from ORF6 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000089
Figure PCTCN2022070948-appb-000090
Table 37. Region-specific peptide pools derived from NSP15 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000091
Figure PCTCN2022070948-appb-000092
Table 38. Region-specific peptide pools derived from ORF3c protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000093
Table 39. Region-specific peptide pools derived from NSP11 protein for CD8 T cell assays.
Region Seq. IDs of peptides
Australia 549, 615
Europe 549, 615
North Africa 549, 615, 3145
North America 615
North East Asia 549, 615
Oceania 549, 615
South and Central America 549, 615
South Asia 549, 615, 3145
South East Asia 549, 615
Sub-Saharan Africa 549, 615
Western Asia 549, 615, 3145
Table 40. Region-specific peptide pools derived from ORF3d protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000094
Figure PCTCN2022070948-appb-000095
Table 41. Region-specific peptide pools derived from ORF10 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000096
Table 42. Region-specific peptide pools derived from ORF9b protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000097
Figure PCTCN2022070948-appb-000098
Table 43. Region-specific peptide pools derived from NSP10 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000099
Table 44. Region-specific peptide pools derived from ORF9c protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000100
Table 45. Region-specific peptide pools derived from NSP7 protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000101
Figure PCTCN2022070948-appb-000102
Table 46. Region-specific peptide pools derived from ORF3b protein for CD8 T cell assays.
Region Seq. IDs of peptides
Australia 835, 1199, 1386, 2221, 2335, 3228
Europe 835, 1199, 1386, 2221, 2335, 2948
North Africa 835, 1199, 2221, 2255, 2335, 2505, 2948, 3228, 3813
North America 835, 1386, 2221, 3228, 3789
North East Asia 835, 1199, 1386, 2221, 2255, 2335, 2505, 2948, 3228, 3789
Oceania 835, 1199, 1386, 2221, 2335, 2505, 3228
South and Central America 835, 1199, 1386, 2221, 2255, 2335, 2505
South Asia 835, 1199, 1386, 2221, 2255, 2335, 2948, 3228
South East Asia 835, 1199, 1386, 2221, 2335, 2505, 3789
Sub-Saharan Africa 835, 1199, 1386, 2221, 2335, 2505, 2948, 3228, 3813
Western Asia 835, 1199, 1386, 2221, 2255, 2335, 2505, 2948, 3228
Table 47. Region-specific peptide pools derived from ORF7b protein for CD8 T cell assays.
Figure PCTCN2022070948-appb-000103
Discussion
SARS2TPools is a first-of-its-kind software platform that provides optimized SARS-CoV-2 peptide pools for assessing vaccine-induced T cell responses in a geographical region. These pools can be used to explore key questions related to COVID-19 vaccines and the role of T cell immunity. For example, characterizing effects of mutations within T cell epitopes in emerging SARS-CoV-2 variants on vaccine-induced T cell responses, investigating association between targeting specific T cell epitopes and protection from severe disease, characterizing the breadth, diversity, and durability of vaccine-induced T cell responses, and contrasting the T cell responses elicited by different vaccines. These questions are important and relevant for both academic and industry research focused on development and assessment of COVID-19 vaccines. Answering them can help provide pre-emptive indicators of the potential for T-cell escape (e.g., due to variants) for specific vaccines.
At present, HLA alleles commonly found in various regions (e.g., South-east Asia, South Asia, and Oceania) are underrepresented in the available experimental data (FIG. 6B) . T cell responses against SARS-CoV-2 proteins other than S are also understudied (FIG. 6C) . Characterizing these responses is particularly important for inactivated whole-virion based vaccines (already in use in more than 48 countries (Shrotri et al. 2021) ) that have been shown to elicit T cell responses against not only S but also against other SARS-CoV-2 proteins (Bueno et al. 2021) . This characterization will also be important for emerging SARS-CoV-2 vaccines (Hwang et al. 2021; Sohail et al. 2021) that incorporate domains or peptides derived from S as well as other SARS-CoV-2 proteins (e.g., EpiVacCorona involves peptides derived from both S and nucleocapsid (Aleksandr B. Ryzhikov et al. 2021; A. B. Ryzhikov et al. 2021) ) . The peptide pools provided by the developed platform can help with assessing T cell responses in above scenarios.
Analysis of the predictions of our in silico approach revealed that some of the experimentally-determined epitopes may be highly promiscuous and thus may cover a large percentage of the global population. Specifically, we identified 12 epitopes which have been experimentally-determined to be associated with at least 2 HLA alleles, while they are predicted to be associated with at least eight or more additional HLA alleles (Table 14) . For example, in silico predictions suggest that the epitope YLQPRTFLL-the most immunoprevalent SARS-CoV- 2 epitope determined so far (Quadeer et al. 2021) -is associated with a total of 22 HLA alleles, two of which (HLA-A*02: 01 and HLA-B*08: 01) have been determined experimentally thus far. While further experiments are required to validate these predicted associations, such promiscuous epitopes, recognized potentially by a large population, can be of interest in the context of designing robust next generation COVID-19 vaccines.
Table 14. List of epitopes having experimentally-determined association with at least 2 HLA alleles and in silico predicted association with at least 8 additional HLA alleles.
Figure PCTCN2022070948-appb-000104
Figure PCTCN2022070948-appb-000105
Compared to T cell response assessment using overlapping peptide pools, peptide pools optimized for a specific region comprise of a limited set of peptides in the context of cognate HLA alleles, and would thus be less susceptible to peptide competition (Pala et al. 1988) . This is of practical importance particularly for a virus like SARS-CoV-2 which has a large (~10k residues) proteome. In addition to overlapping peptide pools, generalized peptide pools have also been proposed to assess SARS-CoV-2 T cell responses. Such pools comprise of peptides associated with a few globally prevalent HLA alleles (e.g., the pool proposed in (Grifoni et al. 2020) comprises of peptides associated with 12 most-prevalent HLA-A and -B alleles) . Compared to such generalized peptide pools, a peptide pool optimized for the HLA alleles prevalent in a specific region would be expected to measure T cell responses in a population more comprehensively.
Recent studies have demonstrated that COVID-19 vaccines elicit strong CD4 + T cell responses, along with CD8 + responses (Sahin et al. 2020; Tauzin et al. 2021) . The pools for measuring CD4 + T cell responses, provided on the platform at present, comprise only of experimentally-determined CD4 + T cell epitopes. In future when sufficient experimentally-determined HLA-resolved CD4 + T cell epitope data becomes available, CD4 + pools provided by SARS2TPools may then be supplemented by in silico predictions following an approach similar to the one employed for CD8 + T cell epitopes. SARS2TPools will be periodically updated to incorporate more experimental CD4 + and CD8 + T cell epitope data as it becomes available.
Methods and Materials
Data collection. We downloaded experimentally-determined HLA class I and class II restricted SARS-CoV-2 T cell epitope data (CD8 + and CD4 +, respectively) from the immune epitope database (IEDB) (Vita et al. 2019) on March 10, 2021. We included all epitopes that were reported in positive T cell assays with associated HLA information available. The data consisted of 768 and 445 unique class I and class II epitope-HLA pairs, respectively. Majority of the HLA class I restricted epitopes (474/768) were nine residues long, which is the canonical length of epitopes restricted by HLA class I alleles. The epitope data was found to be biased towards a handful of HLA alleles, with only 10 HLA class I alleles (HLA-A*02: 01, HLA-A*03: 01, HLA-A*11: 01, HLA-A*24: 02, HLA-A*29: 02, HLA-A*68: 01, HLA-B*07: 02, HLA-B*35: 01, HLA-B*51: 01, HLA-B*57: 01) having 20 or more nine-residue-long epitopes. Collectively, the epitopes restricted by these 10 HLA alleles corresponded to ~62% (295/474) of nine-residue-long epitopes in the data. In the case of HLA class II restricted epitopes, all the available epitopes were 15 resides long, and only 3 HLA alleles had more than 20 epitopes in the data.
In silico prediction methods. Performance of several in silico epitope prediction methods were benchmarked against the set of experimentally-determined SARS-CoV-2 epitopes associated with the 10 HLA class I alleles having the most data. The considered methods included the current state-of-the-art methods such as MHCflurry (O’ Donnell et al. 2020) , NetMHCpan4.1 (Reynisson et al. 2020) , HLAthena (Sarkizova et al. 2020) , NetMHCpan4.0 (Jurtz et al. 2017) , NetMHC4.0 (Andreatta and Nielsen 2016) , along with other common prediction methods that have been employed for predicting SARS-CoV-2 epitopes (Sohail et al. 2021) such as NetMHCpan3.0 (Nielsen and Andreatta 2016) , SMM (Peters and Sette 2005) , SMMPMBEC (Kim et al. 2009) , and IEDB consensus (Moutaftsi et al. 2006) . We considered the eluted ligand and binding affinity predictions (denoted by suffix BA and EL respectively) of NetMHCpan4.1 and NetMHCpan4.0 as separate methods, as was done for the latter method in (Sarkizova et al. 2020) and (Paul et al. 2020) . Similarly, we considered the binding affinity and the presentation score predictions of MHCflurry as two separate methods, referred to as MHCflurry2.0BA and MHCflurry2.0P. In cases where a method required an input other than the protein sequence, HLA allele, and length of the predicted peptides, we used the default parameter settings for that method.
Union approach. In this work, we have proposed a union approach based on combining the top-ranked predictions of MHCflurry2.0P and NetMHCpan4.1BA to obtain a set of peptides  restricted by a given HLA. This approach was motivated by performance comparison analysis of the 12 in silico epitope prediction methods listed above. Briefly, we ranked peptides in ascending order of their predicted score using each method and compared the histograms of ranks of experimentally-determined SARS-CoV-2 CD8 + T cell epitopes associated with the 10 HLA class I alleles with the most data. We found that these histograms were bi-modal for all 12 methods. That is, while the top predictions of each method contained a large number of experimentally-determined epitopes, a good number of epitopes were also ranked quite low by each method (FIG. 10) . Exploring the relationships among the set of top 20 ranked peptides per HLA allele predicted by these methods revealed that the predictions of MHCflurry2.0P were most distinct from those of other methods (FIGS. 11A-D) . Consistent results were obtained when this set was constructed by pooling the top 10 to top 25 ranked peptides restricted by each HLA allele (FIGS. 11A-D) . Predictions of MHCflurry2.0P also contained a large number of experimentally-determined SARS-CoV-2 epitopes that were not present in the set of top-ranked peptides predicted by any other method (FIG 7A) . Given the uniqueness of the predictions of MHCflurry2.0P, we asked if a strategy that combines the predictions of MHCflurry2.0P with any of the other 11 methods would work better than any individual method. The union strategy combines the top x predictions of any two methods and provides a set of peptides whose size can vary between x and 2x depending on the number of common peptides predicted by each method. We fixed one of the methods as MHCflurry2.0P and predicted 11 peptide pools by combining predictions of MHCflurry2.0P with those of the other 11 methods. Our analysis showed that the pool predicted by the union approach always had a higher hit-rate (the fraction of experimentally-determined SARS-CoV-2 epitopes present in the set of top-ranked peptides) than those predicted by the individual methods (FIG. 12) . While comparison among the various union approaches did not readily reveal a clear winner, the unions of MHCflurry2.0P with the in silico methods NetMHC4.0, NetMHCpan4.0BA, NetMHCpan4.1BA, and NetMHCpan4.1EL ranked among the top (FIG. 13 Table 13. ) 
Table 13. Performance comparison of the MHCflurry2.0P-based union methods based on the rank-sum metric, R i.
Figure PCTCN2022070948-appb-000106
Figure PCTCN2022070948-appb-000107
*Rank-sum metric R i is defined as
Figure PCTCN2022070948-appb-000108
where r i (x) is the i-th union method’s rank (assigned based on hit-rate) among the 11 union methods for the set of top x ranked predicted peptides (FIG. 13) . Hit-rate represents the fraction of experimentally known epitopes present in the set of top x ranked predicted peptides.
The union method implemented in SARS2TPools combines the predictions of MHCflurry2.0P with NetMHCpan4.1BA. SARS2TPools provides optimized peptide pools by supplementing experimentally-determined epitopes with small or large sized group of in silico predicted epitopes corresponding respectively to top 10 and top 20 predictions of each method being combined. We used the default thresholds of MHCflurry2.0P and NetMHCpan4.1BA to assess whether or not a peptide is predicted to be an epitope. However, the platform also provides a relaxed threshold which can be particularly useful for specific proteins with very limited number of predicted epitopes.
Statistical analysis. Statistical analyses were performed using the R language (version 3.6) on the RStudio server (version 1.3) . The software platform was developed using the open source R Shiny (version 1.5) development framework.
EXAMPLE 3 –COVIDEP: A WEB-BASED PLATFORM FOR REAL-TIME RESPORTING OF VACCINE TARGET RECOMMENDATIONS FOR SARS-COV-2
Introduction
The COVID-19 pandemic, caused by the novel coronavirus SARS-CoV-2, has brought much of the world to a virtual lockdown. As the virus continues to spread rapidly and the pandemic intensifies, the need for an effective vaccine is becoming increasingly apparent. A critical part of vaccine design is to identify targets, or epitopes, that can induce an effective immune response against SARS-CoV-2. This problem is challenged by our limited understanding of this novel coronavirus and of its interplay with the human immune system.
In response to this challenge, we have developed COVIDep (COVIDep. ust. hk) , a first-of-its-kind web-based platform that pools genetic data for SARS-CoV-2 and immunological data for the 2003 SARS virus, SARS-CoV, to identify B-cell and T-cell epitopes to serve as vaccine target recommendations for SARS-CoV-2 (FIG. 14A) . The identified epitopes are experimentally-derived from SARS-CoV and have a close genetic match with the available SARS-CoV-2 sequences (see FIG. 15 for a detailed protocol description) . Briefly, COVIDep periodically pools SARS-CoV-2 sequence data from the GISAID database (gisaid. org) and compares with experimentally-determined T cell and B cell epitopes of SARS-CoV, obtained from the ViPR database (www. viprbrc. org) . The T cell epitopes were determined based on either positive T cell assays or positive MHC binding assays for SARS-CoV. For the B cell epitopes, both linear and discontinuous epitopes were considered. The system outputs those epitopes that are genetically similar in SARS-CoV-2, based on an epitope screening parameter. This user-defined parameter allows the user to select epitopes based on their conservation in the SARS-CoV-2 sequence data, where conservation is defined as the fraction of SARS-CoV-2 sequences with the exact epitope sequence. The value of this parameter is set to 0.95 as default; however, the user may change this value to adjust the stringency of the screening criterion. For example, reducing the value of the parameter will allow for the consideration of epitopes with greater genetic variation, potentially  increasing the set of recommended SARS-CoV-2 vaccine targets. For the identified T cell epitopes, the population coverage analysis tool available at IEDB (www. iedb. org) is used to estimate the percentage of a specified population that can elicit a response against them. For T-cell epitopes, it provides estimates of population coverage, globally and for specific regions. COVIDep is flexible and user-friendly, comprising an intuitive graphical interface and interactive visualizations. In addition to producing formatted, exportable lists of the identified B-cell and T-cell epitopes and their basic characteristics, COVIDep includes displays for each of the SARS-CoV-2 proteins, showing the locations of the identified epitopes on the primary structure. Further graphical displays are provided to aid interpretation of the data, including a temporal and geographical breakdown of the analysed sequences, and a display of the observed genetic variation (amino acid mutation frequencies) for each of the SARS-CoV-2 proteins. The platform is updated daily, based on the latest SARS-CoV-2 sequence data in the GISAID database (gisaid. org) . Periodic updates are important since SARS-CoV-2 sequences are being made available at an increasing rate through international data sharing efforts, and the identification of vaccine targets is influenced by newly observed genetic variation.
The vaccine targets recommended by COVIDep exploit the genetic similarities between SARS-CoV-2 and SARS-CoV, along with known immune targets for SARS-CoV that have been determined experimentally (available at the ViPR database; viprbrc. org) . The system implements a protocol that identifies, from among the SARS epitopes that can induce a human immune response, those that are genetically similar in SARS-CoV-2. This approach, proposed and tested in Example 1 [see also 56] based on limited early data, identified known SARS-CoV epitopes that had an identical genetic match in SARS-CoV-2. These epitopes presented initial vaccine target recommendations for potentially eliciting a protective, cross-reactive immune response against SARS-CoV-2. Similar results were reported in a subsequent independent study [57] , where a related approach exploiting genetic similarity between SARS-CoV and SARS-CoV-2 was used to identify potential SARS-CoV-2 vaccine targets.
The use of SARS-CoV immunological data to inform vaccine targets for SARS-CoV-2 is being supported by experimental results. There is evidence of cross-neutralization by SARS-CoV-derived antibodies binding to genetically similar regions of SARS-CoV-2’s spike protein [58-60] . Conversely, studies have demonstrated that specific SARS-CoV-derived antibodies  binding to the spike’s receptor binding domain, which has significant genetic differences in SARS-CoV-2, have limited cross-reactivity [61] . T cell responses against spike protein epitopes that are genetically similar in SARS-CoV and SARS-CoV-2 have also been reported in COVID-19 infected patients [62, 63] , and in a preclinical vaccine trial [64] (FIG. 14B) . For instance, FIG. 14B illustrates the T-cell epitopes reported by COVIDep (as of 20 May 2020) for the spike protein of SARS-CoV-2. Here, the Search box (in the top right) was used to select only the HLA-A*02: 01-restricted epitopes. (An explanation of all interactive COVIDep visualizations is incorporated in the “How to use COVIDep page” of the platform. ) . Of the 14 epitopes listed in the display, 9 of them ( IEDB IDs  36724, 54507, 54725, 69657, 71663, 2801, 54680, 16156, and 37289) overlap with epitopes against which cytotoxic CD8+ T cell responses have been observed in peripheral blood mononuclear cells isolated from COVID-19 patients [62, 63] . T cell responses were also recorded against protein regions overlapping with the epitope with IEDB ID 71663 in a pre-clinical trial of a DNA vaccine candidate [64] . Epitopes recommended by COVIDep have notable overlap with the findings in these and other [65, 66] experimental studies [67] (see Figures 2 and 3 in [67] ) .
The recommendations provided by COVIDep may be used to broadly guide vaccine designs and associated experimental studies, and may help to expedite the discovery of an effective vaccine for COVID-19.
Methods and Materials
Data availability. The SARS-CoV-2 full genome sequence data was periodically downloaded from the Global Initiative on Sharing Avian Influenza Database (GISAID; www. gisaid. org) . The SARS-CoV epitope sequence data was downloaded from the Virus Pathogen Database and Analysis Resource (ViPR; viprbrc. org) . The population coverage statistics of HLA alleles were obtained from the Immune Epitope Database and Analysis Resource (IEDB; iedb. org) .
Code availability. The source code for the developed platform is available at the COVIDep GitHub repository github. com/COVIDep) .
All patents, patent applications, and other publications, including GenBank Accession Numbers and equivalents, cited in this application are incorporated by reference in the entirety for all purposes.
REFERENCES
1. Wang, C.; Horby, P.W.; Hayden, F.G.; Gao, G.F. A novel coronavirus outbreak of global health concern. Lancet 2020, 395, 470–473.
2. Centers-of-Disease-Control-and-Prevention Confirmed 2019-nCoV cases globally. Available online: https: //www. cdc. gov/coronavirus/2019-ncov/locations-confirmed-cases. html (accessed on Feb 24, 2020) .
3. World-Health-Organization Statement on the second meeting of the International Health Regulations (2005) Emergency Committee regarding the outbreak of novel coronavirus (2019-nCoV) . Available online: https: //www. who. int/news-room/detail/30-01-2020-statement-on-the-second-meeting-of-the-international-health-regulations- (2005) -emergency-committee-regarding-the-outbreak-of-novel-coronavirus- (2019-ncov) (accessed on Feb 24, 2020) .
4. World-Health-Organization Coronavirus disease (COVID-19) outbreak. Available online: https: //www. who. int/emergencies/diseases/novel-coronavirus-2019 (accessed on Feb 23, 2020) .
5. World-Health-Organization Statement on the meeting of the International Health Regulations (2005) Emergency Committee regarding the outbreak of novel coronavirus (2019-nCoV) . Available online: https: //www. who. int/news-room/detail/23-01-2020-statement-on-the-meeting-of-the-international-health-regulations- (2005) -emergency-committee-regarding-the-outbreak-of-novel-coronavirus- (2019-ncov) (accessed on Feb 24, 2020) .
6. Huang, C.; Wang, Y.; Li, X.; Ren, L.; Zhao, J.; Hu, Y.; Zhang, L.; Fan, G.; Xu, J.; Gu, X.; et al. Clinical features of patients infected with 2019 novel coronavirus in Wuhan, China. Lancet 2020, 395, 497–506.
7. Heymann, D.L. Data sharing and outbreaks: best practice exemplified. Lancet 2020, 395, 469–470.
8. Liu, X.; Wang, X. -J. Potential inhibitors for 2019-nCoV coronavirus M protease from clinically approved medicines. bioRxiv 2020.01.29.924100 2020.
9. Zhou, P.; Yang, X. -L.; Wang, X. -G.; Hu, B.; Zhang, L.; Zhang, W.; Si, H. -R.; Zhu, Y.; Li, B.; Huang, C. -L.; et al. A pneumonia outbreak associated with a new coronavirus of probable bat origin. Nature 2020.
10. World-Health-Organization Update 49 -SARS case fatality ratio, incubation period. Available online: https: //www. who. int/csr/sars/archive/2003_05_07a/en/ (accessed on Feb 23, 2020) .
11. World-Health-Organization Middle East respiratory syndrome coronavirus (MERS-CoV) . Available online: https: //www. who. int/emergencies/mers-cov/en/ (accessed on Feb 23, 2020) .
12. Lu, R.; Zhao, X.; Li, J.; Niu, P.; Yang, B.; Wu, H.; Wang, W.; Song, H.; Huang, B.; Zhu, N.; et al. Genomic characterisation and epidemiology of 2019 novel coronavirus: implications for virus origins and receptor binding.  Lancet  2020, 6736, 1–10.
13. Letko, M.; Munster, V. Functional assessment of cell entry and receptor usage for lineage B β-coronaviruses, including 2019-nCoV. bioRxiv 2020.01.22.915660 2020.
14. Hoffmann, M.; Kleine-Weber, H.; Kruger, N.; Muller, M.; Drosten, C.; Pohlmann, S. The novel coronavirus 2019 (2019-nCoV) uses the SARS-coronavirus receptor ACE2 and the cellular protease TMPRSS2 for entry into target cells. bioRxiv 2020.01.31.929042 2020.
15. Yang, Z. -Y.; Kong, W. -P.; Huang, Y.; Roberts, A.; Murphy, B.R.; Subbarao, K.; Nabel, G.J. A DNA vaccine induces SARS coronavirus neutralization and protective immunity in mice. Nature 2004, 428, 561–564.
16. Deming, D.; Sheahan, T.; Heise, M.; Yount, B.; Davis, N.; Sims, A.; Suthar, M.; Harkema, J.; Whitmore, A.; Pickles, R.; et al. Vaccine efficacy in senescent mice challenged with recombinant SARS-CoV bearing epidemic and zoonotic spike variants. PLoS Med. 2006, 3, e525.
17. Graham, R.L.; Becker, M.M.; Eckerle, L.D.; Bolles, M.; Denison, M.R.; Baric, R.S. A live, impaired-fidelity coronavirus vaccine protects in an aged, immunocompromised mouse model of lethal disease. Nat. Med. 2012, 18, 1820–1826.
18. Lin, Y.; Shen, X.; Yang, R.F.; Li, Y.X.; Ji, Y.Y.; He, Y.Y.; Shi, M. De; Lu, W.; Shi, T.L.; Wang, J.; et al. Identification of an epitope of SARS-coronavirus nucleocapsid protein. Cell Res. 2003, 13, 141–145.
19. Wang, J.; Wen, J.; Li, J.; Yin, J.; Zhu, Q.; Wang, H.; Yang, Y.; Qin, E.; You, B.; Li, W.; et al. Assessment of immunoreactive synthetic peptides from the structural proteins of severe acute respiratory syndrome coronavirus. Clin. Chem. 2003, 49, 1989–1996.
20. Liu, X.; Shi, Y.; Li, P.; Li, L.; Yi, Y.; Ma, Q.; Cao, C. Profile of antibodies to the nucleocapsid protein of the severe acute respiratory syndrome (SARS) -associated coronavirus in probable SARS patients. Clin. Vaccine Immunol. 2004, 11, 227–228.
21. Tang, F.; Quan, Y.; Xin, Z. -T.; Wrammert, J.; Ma, M. -J.; Lv, H.; Wang, T. -B.; Yang, H.; Richardus, J.H.; Liu, W.; et al. Lack of peripheral memory B cell responses in recovered patients with severe acute respiratory syndrome: A six-year follow-up study. J. Immunol. 2011, 186, 7264–7268.
22. Peng, H.; Yang, L. -T.; Wang, L. -Y.; Li, J.; Huang, J.; Lu, Z. -Q.; Koup, R.A.; Bailer, R.T.; Wu, C. -Y. Long-lived memory T lymphocyte responses against SARS coronavirus nucleocapsid protein in SARS-recovered patients. Virology 2006, 351, 466–475.
23. Fan, Y. -Y.; Huang, Z. -T.; Li, L.; Wu, M. -H.; Yu, T.; Koup, R.A.; Bailer, R.T.; Wu, C. -Y. Characterization of SARS-CoV-specific memory T cells from recovered individuals 4 years after infection. Arch. Virol. 2009, 154, 1093–1099.
24. Ng, O. -W.; Chia, A.; Tan, A.T.; Jadi, R.S.; Leong, H.N.; Bertoletti, A.; Tan, Y. -J. Memory T cell responses targeting the SARS coronavirus persist up to 11 years post-infection. Vaccine 2016, 34, 2008–2014.
25. Liu, W.J.; Zhao, M.; Liu, K.; Xu, K.; Wong, G.; Tan, W.; Gao, G.F. T-cell immunity of SARS-CoV: Implications for vaccine development against MERS-CoV. Antiviral Res. 2017, 137, 82–92.
26. Li, C.K. -F.; Wu, H.; Yan, H.; Ma, S.; Wang, L.; Zhang, M.; Tang, X.; Temperton, N.J.; Weiss, R.A.; Brenchley, J.M.; et al. T cell responses to whole SARS coronavirus in humans. J. Immunol. 2008, 181, 5490–5500.
27. Channappanavar, R.; Fett, C.; Zhao, J.; Meyerholz, D.K.; Perlman, S. Virus-specific memory CD8 T cells provide substantial protection from lethal severe acute respiratory syndrome coronavirus infection. J. Virol. 2014, 88, 11034–11044.
28. Katoh, K.; Standley, D.M. MAFFT multiple sequence alignment software version 7: Improvements in performance and usability. Mol. Biol. Evol. 2013, 30, 772–780.
29. Pickett, B.E.; Sadat, E.L.; Zhang, Y.; Noronha, J.M.; Squires, R.B.; Hunt, V.; Liu, M.; Kumar, S.; Zaremba, S.;Gu, Z.; et al. ViPR: An open bioinformatics database and analysis resource for virology research. Nucleic Acids Res. 2012, 40, D593–D598.
30. Vita, R.; Mahajan, S.; Overton, J.A.; Dhanda, S.K.; Martini, S.; Cantrell, J.R.; Wheeler, D.K.; Sette, A.; Peters, B. The immune epitope database (IEDB) : 2018 update. Nucleic Acids Res. 2019, 47, D339–D343.
31. Mirarab, S.; Nguyen, N.; Guo, S.; Wang, L. -S.; Kim, J.; Warnow, T. PASTA: Ultra-large multiple sequence alignment for nucleotide and amino-acid sequences. J. Comput. Biol. 2015, 22, 377–386.
32. Huson, D.H.; Scornavacca, C. Dendroscope 3: An interactive tool for rooted phylogenetic trees and networks. Syst. Biol. 2012, 61, 1061–1067.
33. Ahmed, S.F. Data and software code for reproducing results of this paper. Available online: https: //github. com/faraz107/2019-nCoV-T-Cell-Vaccine-Candidates (accessed on Feb 25, 2020) .
34. Li, F. Structure of SARS coronavirus spike receptor-binding domain complexed with receptor. Science. 2005, 309, 1864–1868.
35. Dahirel, V.; Shekhar, K.; Pereyra, F.; Miura, T.; Artyomov, M.; Talsania, S.; Allen, T.M.; Altfeld, M.; Carrington, M.; Irvine, D.J.; et al. Coordinate linkage of HIV evolution reveals regions of immunological vulnerability. Proc. Natl. Acad. Sci. 2011, 108, 11530–11535.
36. Quadeer, A.A.; Louie, R.H.Y.; Shekhar, K.; Chakraborty, A.K.; Hsing, I. -M.; McKay, M.R. Statistical linkage analysis of substitutions in patient-derived sequences of genotype 1a hepatitis C virus nonstructural protein 3 exposes targets for immunogen design. J. Virol. 2014, 88, 7628–7644.
37. Ahmed, S.F.; Quadeer, A.A.; Morales-Jimenez, D.; McKay, M.R. Sub-dominant principal components inform new vaccine targets for HIV Gag. Bioinformatics 2019, 35, 3884–3889.
38. Quadeer, A.A.; Morales-Jimenez, D.; McKay, M.R. Co-evolution networks of HIV/HCV are modular with direct association to structure and function. PLOS Comput. Biol. 2018, 14, e1006409.
39. Prabakaran, P.; Gan, J.; Feng, Y.; Zhu, Z.; Choudhry, V.; Xiao, X.; Ji, X.; Dimitrov, D. S. Structure of severe acute respiratory syndrome coronavirus receptor-binding domain complexed with neutralizing antibody. J. Biol. Chem. 2006, 281, 15829–15836.
40. Zhu, Z.; Chakraborti, S.; He, Y.; Roberts, A.; Sheahan, T.; Xiao, X.; Hensley, L.E.; Prabakaran, P.; Rockx, B.; Sidorov, I.A.; et al. Potent cross-reactive neutralization of SARS coronavirus isolates by human monoclonal antibodies. Proc. Natl. Acad. Sci. 2007, 104, 12123–12128.
41. Hwang, W.C.; Lin, Y.; Santelli, E.; Sui, J.; Jaroszewski, L.; Stec, B.; Farzan, M.; Marasco, W.A.; Liddington, R.C. Structural basis of neutralization by a human anti-severe acute respiratory syndrome spike protein antibody, 80R. J. Biol. Chem. 2006, 281, 34610–34616.
42. UniProt UniProtKB -P59594 (SPIKE_CVHSA) . Available online: https: //www. uniprot. org/uniprot/P59594 (accessed on Feb 23, 2020) .
43. Walls, A.C.; Park, Y. -J.; Tortorici, M.A.; Wall, A.; McGuire, A.T.; Veesler, D. Structure, function and antigenicity of the SARS-CoV-2 spike glycoprotein. bioRxiv 2020, 2020.02.19.956581.
44. Walls, A.C.; Xiong, X.; Park, Y. -J.; Tortorici, M.A.; Snijder, J.; Quispe, J.; Cameroni, E.; Gopal, R.; Dai, M.;Lanzavecchia, A.; et al. Unexpected receptor functional mimicry elucidates activation of coronavirus fusion. Cell 2019, 176, 1026-1039. e15.
45. Walls, A.C.; Tortorici, M.A.; Snijder, J.; Xiong, X.; Bosch, B. -J.; Rey, F.A.; Veesler, D. Tectonic conformational changes of a coronavirus spike glycoprotein promote membrane fusion. Proc. Natl. Acad. Sci. 2017, 114, 11157–11162.
46. Song, W.; Gui, M.; Wang, X.; Xiang, Y. Cryo-EM structure of the SARS coronavirus spike glycoprotein in complex with its host cell receptor ACE2. PLOS Pathog. 2018, 14, e1007236.
47. Wrapp, D.; Wang, N.; Corbett, K.S.; Goldsmith, J.A.; Hsieh, C. -L.; Abiona, O.; Graham, B.S.; McLellan, J.S. Cryo-EM structure of the 2019-nCoV spike in the prefusion conformation. Science. 2020, 2011, eabb2507.
48. Tian, X.; Li, C.; Huang, A.; Xia, S.; Lu, S.; Shi, Z.; Lu, L.; Jiang, S.; Yang, Z.; Wu, Y.; et al. Potent binding of 2019 novel coronavirus spike protein by a SARS coronavirus-specific human monoclonal antibody. Emerg. Microbes Infect. 2020, 9, 382–385.
49. Ferguson, A.L.; Mann, J.K.; Omarjee, S.; Ndung’u, T.; Walker, B.D.; Chakraborty, A.K. Translating HIV sequences into quantitative fitness landscapes predicts viral vulnerabilities for rational immunogen design. Immunity 2013, 38, 606–617.
50. Chakraborty, A.K.; Barton, J.P. Rational design of vaccine targets and strategies for HIV: a crossroad of statistical physics, biology, and medicine. Reports Prog. Phys. 2017, 80, 032601.
51. Quadeer, A.A.; Louie, R.H.Y.; McKay, M.R. Identifying immunologically-vulnerable regions of the HCV E2 glycoprotein and broadly neutralizing antibodies that target them. Nat. Commun. 2019, 10, 2073.
52. Louie, R.H.Y.; Kaczorowski, K.J.; Barton, J.P.; Chakraborty, A.K.; McKay, M.R. Fitness landscape of the human immunodeficiency virus envelope protein that is targeted by antibodies. Proc. Natl. Acad. Sci. 2018, 115, E564–E573.
53. Quadeer, A.A.; Barton, J.P.; Chakraborty, A.K.; McKay, M.R. Deconvolving mutational patterns of poliovirus outbreaks reveals its intrinsic fitness landscape. Nat. Commun. 2020, 11, 377.
54. Mann, J.K.; Barton, J.P.; Ferguson, A.L.; Omarjee, S.; Walker, B.D.; Chakraborty, A.; Ndung’u, T. The fitness landscape of HIV-1 Gag: Advanced modeling approaches and validation of model predictions by in vitro testing. PLoS Comput. Biol. 2014, 10, e1003776.
55. Ramaiah, A.; Arumugaswami, V. Insights into cross-species evolution of novel human coronavirus 2019-nCoV and defining immune determinants for vaccine development. bioRxiv 2020.01.29.925867 2020.
56. Ahmed, S.F., Quadeer, A.A. &McKay, M.R. Preliminary identification of potential vaccine targets for the COVID-19 coronavirus (SARS-CoV-2) based on SARS-CoV immunological studies. Viruses 12, 254 (2020) .
57. Grifoni, A. et al. A sequence homology and bioinformatic approach can predict candidate targets for immune responses to SARS-CoV-2.  Cell Host Microbe  27, 1–10 (2020) .
58. Walls, A.C. et al. Structure, function, and antigenicity of the SARS-CoV-2 spike glycoprotein. Cell 180, 1–12 (2020) .
59. Wang, C. et al. A human monoclonal antibody blocking SARS-CoV-2 infection. Nat. Commun. 11, 2251 (2020) .
60. Pinto, D. et al. Cross-neutralization of SARS-CoV-2 by a human monoclonal SARS-CoV antibody. Nature (2020) . doi: 10.1038/s41586-020-2349-y
61. Wrapp, D. et al. Cryo-EM structure of the 2019-nCoV spike in the prefusion conformation. Science 367, 1260–1263 (2020) .
62. Chour, W. et al. Shared antigen-specific CD8+ T cell responses against the SARS-COV-2 spike protein in HLA A*02: 01 COVID-19 participants. medRxiv 2020.05.04.20085779 (2020) . doi: 10.1101/2020.05.04.20085779
63. Shomuradova, A.S. et al. SARS-CoV-2 epitopes are recognized by a public and diverse repertoire of human T-cell receptors. bioRxiv 2020.05.20.20107813 (2020) . doi: 10.1101/2020.05.20.20107813
64. Smith, T.R.F. et al. Immunogenicity of a DNA vaccine candidate for COVID-19. Nat. Commun. 11, 2601  (2020) .
65. Poh, C.M. et al. Two linear epitopes on the SARS-CoV-2 spike protein that elicit neutralising antibodies in COVID-19 patients. Nat. Commun. 11, 2806 (2020) .
66. Yin, D. et al. A single dose SARS-CoV-2 simulating particle vaccine induces potent neutralizing activities. bioRxiv 2020.05.14.093054 (2020) . doi: 10.1101/2020.05.14.093054
67. Ahmed, S.F., Quadeer, A.A. &McKay, M.R. COVIDep platform for real-time reporting of vaccine target recommendations for SARS-CoV-2: Description and connections with COVID-19 immune responses and preclinical vaccine trials. bioRxiv 2020.05.23.111385 (2020) . doi: 10.1101/2020.05.23.111385.
68. Abdool Karim SS, de Oliveira T. 2021. New SARS-CoV-2 variants -clinical, public health, and vaccine implications. N. Engl. J. Med. [Internet] : NEJMc2100362. Available from: www. nejm. org/doi/10.1056/NEJMc2100362.
69. Abu-Raddad LJ, Chemaitelly H, Butt AA. 2021. Effectiveness of the BNT162b2 COVID-19 vaccine against the B. 1.1.7 and B. 1.351 variants. N. Engl. J. Med. [Internet] : NEJMc2104974. Available from: www. nejm. org/doi/10.1056/NEJMc2104974.
70. Altmann DM, Boyton RJ. 2020. SARS-CoV-2 T cell immunity: Specificity, function, durability, and role in protection. Sci. Immunol. [Internet] 5: 2–7. Available from: 10.0.4.102/sciimmunol. abd6160.
71. Anderson EJ, Rouphael NG, Widge AT, Jackson LA, Roberts PC, Makhene M, Chappell JD, Denison MR, Stevens LJ, Pruijssers AJ, et al. 2020. Safety and immunogenicity of SARS-CoV-2 mRNA-1273 vaccine in older adults. N. Engl. J. Med. 383: 2427–2438.
72. Andreatta M, Nielsen M. 2016. Gapped sequence alignment using artificial neural networks: Application to the MHC class I system. Bioinformatics [Internet] 32: 511–517. Available from: https: //academic. oup. com/bioinformatics/article/32/4/511/1744469.
73. Bergamaschi L, Mescia F, Turner L, Hanson A, Kotagiri P, Dunmore BJ, Ruffieux H, De Sa A, Huhn O, Morgan MD, et al. 2021. Delayed bystander CD8+ T cell activation, early immune pathology and persistent dysregulation characterise severe COVID-19. medRxiv [Internet] : 2021.01.11.20248765. Available from: medrxiv. org/content/early/2021/03/26/2021.01.11.20248765. abstract.
74. Bertoletti A, Tan AT, Le Bert N. 2021. The T-cell response to SARS-CoV-2: Kinetic and quantitative aspects and the case for their protective role. Oxford Open Immunol. [Internet] 2: 1–9. Available from: https: //academic. oup. com/ooim/article/doi/10.1093/oxfimm/iqab006/6146940.
75. Bueno SM, Abarca K, González PA, Gálvez NMS, Soto JA, Duarte LF, Schultz BM, Pacheco GA, González LA, Ríos M, et al. 2021. Interim report: Safety and immunogenicity of an inactivated vaccine against SARS-CoV-2 in healthy Chilean adults in a phase 3 clinical trial. medRxiv [Internet] : 2021.03.31.21254494. Available from: medrxiv. org/content/early/2021/04/01/2021.03.31.21254494. abstract.
76. Chen Z, John Wherry E. 2020. T cell responses in patients with COVID-19. Nat. Rev. Immunol. [Internet] 20:529–536. Available from: www. nature. com/articles/s41577-020-0402-6.
77. Cohen KW, Linderman SL, Moodie Z, Czartoski J, Lai L, Mantus G, Norwood C, Nyhoff LE, Edara VV, Floyd K, et al. 2021. Longitudinal analysis shows durable and broad immune memory after SARS-CoV-2 infection with persisting antibody responses and memory B and T cells. medRxiv [Internet] : 2021.04.19.21255739. Available from: medrxiv. org/content/early/2021/04/27/2021.04.19.21255739. abstract.
78. Collier DA, De Marco A, Ferreira IATM, Meng B, Datir R, Walls AC, Kemp S SA, Bassi J, Pinto D, Fregni CS, et al. 2021. Sensitivity of SARS-CoV-2 B. 1.1.7 to mRNA vaccine-elicited antibodies. Nature.
79. Ella R, Vadrevu KM, Jogdand H, Prasad S, Reddy S, Sarangi V, Ganneru B, Sapkal G, Yadav P, Abraham P, et al. 2021. Safety and immunogenicity of an inactivated SARS-CoV-2 vaccine, BBV152: A double-blind, randomised, phase 1 trial. Lancet Infect. Dis. [Internet] 3099: 2020.12.21.20248643. Available from: https: //doi. org/10.1016/S1473-3099 (20) 30942-7.
80. Emary KRW, Golubchik T, Aley PK, Ariani C V., Angus B, Bibi S, Blane B, Bonsall D, Cicconi P, Charlton S, et al. 2021. Efficacy of ChAdOx1 nCoV-19 (AZD1222) vaccine against SARS-CoV-2 variant of concern 202012/01 (B. 1.1.7) : An exploratory analysis of a randomised controlled trial. Lancet 397: 1351–1362.
81. Fedry J, Hurdiss DL, Wang C, Li W, Obal G, Drulyte I, Du W, Howes SC, van Kuppeveld FJM, 
Figure PCTCN2022070948-appb-000109
F, et al. 2021. Structural insights into the cross-neutralization of SARS-CoV and SARS-CoV-2 by the human monoclonal antibody 47D11. Sci. Adv. [Internet] 7: eabf5632. Available from: https: //advances. sciencemag. org/lookup/doi/10.1126/sciadv. abf5632.
82. Folegatti PM, Ewer KJ, Aley PK, Angus B, Becker S, Belij-Rammerstorfer S, Bellamy D, Bibi S, Bittaye  M, Clutterbuck EA, et al. 2020. Safety and immunogenicity of the ChAdOx1 nCoV-19 vaccine against SARS-CoV-2: A preliminary report of a phase 1/2, single-blind, randomised controlled trial. Lancet [Internet] 396: 467–478. Available from: https: //linkinghub. elsevier. com/retrieve/pii/S0140673620316044.
83. Gonzalez-Galarza FF, McCabe A, Santos EJM dos, Jones J, Takeshita L, Ortega-Rivera ND, Cid-Pavon GM Del, Ramsbottom K, Ghattaoraya G, Alfirevic A, et al. 2019. Allele frequency net database (AFND) 2020 update: Gold-standard data classification, open access genotype data and new query tools. Nucleic Acids Res. [Internet] . Available from: https: //academic. oup. com/nar/advance-article/doi/10.1093/nar/gkz1029/5624967.
84. Grifoni A, Weiskopf D, Ramirez SI, Mateus J, Dan JM, Moderbacher CR, Rawlings SA, Sutherland A, Premkumar L, Jadi RS, et al. 2020. Targets of T cell responses to SARS-CoV-2 coronavirus in humans with COVID-19 disease and unexposed individuals. Cell [Internet] 181: 1489-1501. e15. Available from: https: //doi. org/10.1016/j. cell. 2020.05.015.
85. Hall VJ, Foulkes S, Saei A, Andrews N, Oguti B, Charlett A, Wellington E, Stowe J, Gillson N, Atti A, et al.2021. COVID-19 vaccine coverage in health-care workers in England and effectiveness of BNT162b2 mRNA vaccine against infection (SIREN) : A prospective, multicentre, cohort study. Lancet [Internet] 6736: 1–11. Available from: https: //linkinghub. elsevier. com/retrieve/pii/S014067362100790X.
86. Hwang W, Lei W, Katritsis NM, MacMahon M, Chapman K, Han N. 2021. Current and prospective computational approaches and challenges for developing COVID-19 vaccines. Adv. Drug Deliv. Rev. [Internet] 172: 249–274. Available from: https: //linkinghub. elsevier. com/retrieve/pii/S0169409X21000387
87. Jin Y, Wang J, Bachtiar M, Chong SS, Lee CGL. 2018. Architecture of polymorphisms in the human genome reveals functionally important and positively selected variants in immune response and drug transporter genes. Hum. Genomics [Internet] 12: 43. Available from: https: //humgenomics. biomedcentral. com/articles/10.1186/s40246-018-0175-1
88. Jurtz V, Paul S, Andreatta M, Marcatili P, Peters B, Nielsen M. 2017. NetMHCpan-4.0: Improved peptide–MHC class I interaction predictions integrating eluted ligand and peptide binding affinity data. J. Immunol. [Internet] 199: 3360–3368. Available from: www. jimmunol. org/lookup/doi/10.4049/jimmunol. 1700893
89. Keech C, Albert G, Cho I, Robertson A, Reed P, Neal S, Plested JS, Zhu M, Cloney-Clark S, Zhou H, et al. 2020. Phase 1–2 trial of a SARS-CoV-2 recombinant spike protein nanoparticle vaccine. N. Engl. J. Med. [Internet] : NEJMoa2026920. Available from: www. nejm. org/doi/10.1056/NEJMoa2026920
90. Kim Y, Sidney J, Pinilla C, Sette A, Peters B. 2009. Derivation of an amino acid similarity matrix for peptide: MHC binding and its application as a Bayesian prior. BMC Bioinformatics [Internet] 10: 394. Available from: www. ncbi. nlm. nih. gov/pubmed/19948066
91. Klasse PJ, Nixon DF, Moore JP. 2021. Immunogenicity of clinically relevant SARS-CoV-2 vaccines in nonhuman primates and humans. Sci. Adv. 7: 1–23.
92. Liao M, Liu Y, Yuan J, Wen Y, Xu G, Zhao J, Cheng L, Li J, Wang X, Wang F, et al. 2020. Single-cell landscape of bronchoalveolar immune cells in patients with COVID-19. Nat. Med. [Internet] 26: 842–844. Available from: dx. doi. org/10.1038/s41591-020-0901-9.
93. Liu Y, Liu J, Xia H, Zhang X, Fontes-Garfias CR, Swanson KA, Cai H, Sarkar R, Chen W, Cutler M, et al. 2021. Neutralizing activity of BNT162b2-elicited serum. N. Engl. J. Med. [Internet] 384: 1466–1468. Available from: www. nejm. org/doi/10.1056/NEJMc2102017.
94. Logunov DY, Dolzhikova I V., Zubkova O V., Tukhvatullin AI, Shcheblyakov D V., Dzharullaeva AS, Grousova DM, Erokhova AS, Kovyrshina A V., Botikov AG, et al. 2020. Safety and immunogenicity of an rAd26 and rAd5 vector-based heterologous prime-boost COVID-19 vaccine in two formulations: Two open, non-randomised phase 1/2 studies from Russia. Lancet 396: 887–897.
95. Madhi SA, Baillie V, Cutland CL, Voysey M, Koen AL, Fairlie L, Padayachee SD, Dheda K, Barnabas SL, Bhorat QE, et al. 2021. Efficacy of the ChAdOx1 nCoV-19 COVID-19 vaccine against the B. 1.351 variant. N.Engl. J. Med. [Internet] : NEJMoa2102214. Available from: www. nejm. org/doi/10.1056/NEJMoa2102214.
96. Mazzoni A, Maggi L, Capone M, Spinicci M, Salvati L, Colao MG, Vanni A, Kiros ST, Mencarini J, Zammarchi L, et al. 2020. Cell-mediated and humoral adaptive immune responses to SARS-CoV-2 are lower in asymptomatic than symptomatic COVID-19 patients. Eur. J. Immunol. 50: 2013–2024.
97. Mor V, Gutman R, Yang X, White EM, McConeghy KW, Feifer RA, Blackman CR, Kosar CM, Bardenheier BH, Gravenstein SA. 2021. Short‐term impact of nursing home SARS‐CoV‐2 vaccinations on new infections, hospitalizations, and deaths. J. Am. Geriatr. Soc. [Internet] : jgs. 17176. Available from: https: //onlinelibrary. wiley. com/doi/10.1111/jgs. 17176
98. Moutaftsi M, Peters B, Pasquetto V, Tscharke DC, Sidney J, Bui HH, Grey H, Sette A. 2006. A consensus  epitope prediction approach identifies the breadth of murine TCD8+-cell responses to vaccinia virus. Nat. Biotechnol. 24: 817–819.
99. Nielsen M, Andreatta M. 2016. NetMHCpan-3.0; improved prediction of binding to MHC class I molecules integrating information from multiple receptor and peptide length datasets. Genome Med. [Internet] 8: 33. Available from: https: //genomemedicine. biomedcentral. com/articles/10.1186/s13073-016-0288-x
100. O’ Donnell TJ, Rubinsteyn A, Laserson U. 2020. MHCflurry 2.0: Improved pan-allele prediction of MHC class I-presented peptides by incorporating antigen processing. Cell Syst. [Internet] 11: 42-48. e7. Available from: https: //doi. org/10.1016/j. cels. 2020.06.010
101. Pala P, Bodmer HC, Pemberton RM, Cerottini J-C, Maryanski JL, Askonas BA. 1988. Competition between unrelated peptides recognized by H-2-Kd restricted T cells. J. Immunol. 141: 2289–2294.
102. Paul S, Croft NP, Purcell AW, Tscharke DC, Sette A, Nielsen M, Peters B. 2020. Benchmarking predictions of MHC class I restricted T cell epitopes in a comprehensively studied model system. PLoS Comput. Biol. [Internet] 16: 1–18. Available from: dx. doi. org/10.1371/journal. pcbi. 1007757
103. Peters B, Sette A. 2005. Generating quantitative models describing the sequence specificity of biological processes with the stabilized matrix method. BMC Bioinformatics [Internet] 6. Available from: www. ncbi. nlm. nih. gov/pubmed/15927070
104. Piccoli L, Park Y-J, Tortorici MA, Czudnochowski N, Walls AC, Beltramello M, Silacci-Fregni C, Pinto D, Rosen LE, Bowen JE, et al. 2020. Mapping neutralizing and immunodominant sites on the SARS-CoV-2 spike receptor-binding domain by structure-guided high-resolution serology. Cell [Internet] 183: 1024-1042. e21. Available from: https: //linkinghub. elsevier. com/retrieve/pii/S0092867420312344
105. Pinto D, Park Y, Beltramello M, Walls AC, Tortorici MA, Bianchi S, Jaconi S, Culap K, Zatta F, De Marco A, et al. 2020. Cross-neutralization of SARS-CoV-2 by a human monoclonal SARS-CoV antibody. Nature [Internet] 583: 290–295. Available from: dx. doi. org/10.1038/s41586-020-2349-y
106. Planas D, Bruel T, Grzelak L, Guivel-Benhassine F, Staropoli I, Porrot F, Planchais C, Buchrieser J, Rajah MM, Bishop E, et al. 2021. Sensitivity of infectious SARS-CoV-2 B. 1.1.7 and B. 1.351 variants to neutralizing antibodies. Nat. Med. [Internet] . Available from: www. nature. com/articles/s41591-021-01318-5
107. Quadeer AA, Ahmed SF, McKay MR. 2021. Landscape of epitopes targeted by T cells in 852 convalescent COVID-19 patients: Meta-analysis, immunoprevalence and web platform. Cell Reports Med. [Internet] : 100312. Available from: https: //linkinghub. elsevier. com/retrieve/pii/S2666379121001555
108. Ramasamy MN, Minassian AM, Ewer KJ, Flaxman AL, Folegatti PM, Owens DR, Voysey M, Aley PK, Angus B, Babbage G, et al. 2020. Safety and immunogenicity of ChAdOx1 nCoV-19 vaccine administered in a prime-boost regimen in young and old adults (COV002) : A single-blind, randomised, controlled, phase 2/3 trial. Lancet 396: 1979–1993.
109. Reynisson B, Alvarez B, Paul S, Peters B, Nielsen M. 2020. NetMHCpan-4.1 and NetMHCIIpan-4.0: Improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Res. 48: W449–W454.
110. Reynolds CJ, Swadling L, Gibbons JM, Pade C, Jensen MP, Diniz MO, Schmidt NM, Butler DK, Amin OE, Bailey SNL, et al. 2020. Discordant neutralizing antibody and T cell responses in asymptomatic and mild SARS-CoV-2 infection. Sci. Immunol. [Internet] 5: eabf3698. Available from: https: //immunology. sciencemag. org/lookup/doi/10.1126/sciimmunol. abf3698
111. Robinson J, Halliwell JA, Hayhurst JD, Flicek P, Parham P, Marsh SGE. 2015. The IPD and IMGT/HLA database: Allele variant databases. Nucleic Acids Res. [Internet] 43: D423–D431. Available from: academic. oup. com/nar/article/43/D1/D423/2438496/The-IPD-and-IMGTHLA-database-allele-variant
112. Rydyznski Moderbacher C, Ramirez SI, Dan JM, Grifoni A, Hastie KM, Weiskopf D, Belanger S, Abbott RK, Kim Christina, Choi J, et al. 2020. Antigen-specific adaptive immunity to SARS-CoV-2 in acute COVID-19 and associations with age and disease severity. Cell [Internet] : 1–17. Available from: https: //doi. org/10.1016/j. cell. 2020.09.038
113. Ryzhikov Aleksandr B., Ryzhikov EА, Bogryantseva MP, Danilenko ED, Imatdinov IR, Nechaeva EA, Pyankov O V., Pyankova OG, Susloparov IM, Taranov OS, et al. 2021. Immunogenicity and protectivity of the peptide candidate vaccine against SARS-CoV-2. Ann. Russ. Acad. Med. Sci. [Internet] 76: 5–19. Available from: https: //vestnikramn. spr-journal. ru/jour/article/view/1528
114. Ryzhikov A. B., Ryzhikov ЕА, Bogryantseva MP, Usova S V., Danilenko ED, Nechaeva EA, Pyankov O V., Pyankova OG, Gudymo AS, Bodnev SA, et al. 2021. A single blind, placebo-controlled randomized study of the safety, reactogenicity and immunogenicity of the “EpiVacCorona” vaccine for the prevention of COVID-19, in volunteers aged 18–60 years (phase I–II) . Russ. J. Infect. Immun. [Internet] 11: 283–296.  Available from: https: //www. iimmun. ru/iimm/article/view/1699
115. Sadoff J, Le Gars M, Shukarev G, Heerwegh D, Truyers C, de Groot AM, Stoop J, Tete S, Van Damme W, Leroux-Roels I, et al. 2021. Interim results of a phase 1–2a trial of Ad26. COV2. S COVID-19 vaccine. N. Engl. J. Med.: 1–12.
116. Sahin U, Muik A, Derhovanessian E, Vogler I, Kranz LM, Vormehr M, Baum A, Pascal K, Quandt J, Maurus D, et al. 2020. COVID-19 vaccine BNT162b1 elicits human antibody and TH1 T cell responses. Nature [Internet] 586: 594–599. Available from: dx. doi. org/10.1038/s41586-020-2814-7
117. Sahin U, Muik A, Vogler I, Derhovanessian E, Kranz LM, Vormehr M, Quandt J, Bidmon N, Ulges A, Baum A, et al. 2021. BNT162b2 vaccine induces neutralizing antibodies and poly-specific T cells in humans. Nature [Internet] . Available from: www. nature. com/articles/s41586-021-03653-6
118. Sarkizova S, Klaeger S, Le PM, Li LW, Oliveira G, Keshishian H, Hartigan CR, Zhang W, Braun DA, Ligon KL, et al. 2020. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nat. Biotechnol. [Internet] 38: 199–209. Available from: dx. doi. org/10.1038/s41587-019-0322-9
119. Shrotri M, Swinnen T, Kampmann B, Parker EPK. 2021. An interactive website tracking COVID-19 vaccine development. Lancet Glob. Heal. [Internet] 9: e590–e592. Available from: dx. doi. org/10.1016/S2214-109X (21) 00043-7
120. Sohail MS, Ahmed SF, Quadeer AA, McKay MR. 2021. In silico T cell epitope identification for SARS-CoV-2: Progress and perspectives. Adv. Drug Deliv. Rev. [Internet] 171: 29–47. Available from: https: //doi. org/10.1016/j. addr. 2021.01.007
121. Tarke A, Sidney J, Methot N, Zhang Y, Dan JM, Goodwin B, Rubiro P, Sutherland A, da Silva Antunes R, Frazier A, et al. 2021. Negligible impact of SARS-CoV-2 variants on CD4+ and CD8+ T cell reactivity in COVID-19 exposed donors and vaccinees. bioRxiv [Internet] : 2021.02.27.433180. Available from: www. ncbi. nlm. nih. gov/pubmed/33688655%0Awww. pubmedcentral. nih. gov/articlerender. fcgi? artid=PMC7 941626
122. Tauzin A, Nayrac M, Benlarbi M, Gong SY, Gasser R, Beaudoin-Bussières G, Brassard N, Laumaea A, Vézina D, Prévost J, et al. 2021. A single dose of the SARS-CoV-2 vaccine BNT162b2 elicits Fc-mediated antibody effector functions and T-cell responses. Cell Host Microbe [Internet] . Available from: https: //doi. org/10.1016/j. chom. 2021.06.001
123. Vita R, Mahajan S, Overton JA, Dhanda SK, Martini S, Cantrell JR, Wheeler DK, Sette A, Peters B. 2019. The immune epitope database (IEDB) : 2018 update. Nucleic Acids Res. [Internet] 47: D339–D343. Available from: https: //doi. org/10.1093/nar/gky1006
124. Wall EC, Wu M, Harvey R, Kelly G, Warchal S, Sawyer C, Daniels R, Hobson P, Hatipoglu E, Ngai Y, et al. 2021. Neutralising antibody activity against SARS-CoV-2 VOCs B. 1.617.2 and B. 1.351 by BNT162b2 vaccination. Lancet [Internet] 6736: 3–5. Available from: dx. doi. org/10.1016/S0140-6736 (21) 01290-3
125. Woldemeskel BA, Garliss CC, Blankson JN. 2021. SARS-CoV-2 mRNA vaccines induce broad CD4+ T cell responses that recognize SARS-CoV-2 variants and HCoV-NL63. J. Clin. Invest. 131.
128. Wyllie D, Mulchandani R, Jones HE, Taylor-phillips S, Brooks T, Charlett A, Ades AE, Makin A, Oliver I. 2020. SARS-CoV-2 responsive T cell numbers are associated with protection from COVID-19: A prospective cohort study in keyworkers. medRxiv.
127. Zhou D, Dejnirattisai W, Supasa P, Liu C, Mentzer AJ, Ginn HM, Zhao Y, Duyvesteyn HME, Tuekprakhon A, Nutalai R, et al. 2021. Evidence of escape of SARS-CoV-2 variant B. 1.351 from natural and vaccine-induced sera. Cell [Internet] 184: 2348-2361. e6. Available from: https: //doi. org/10.1016/j. cell. 2021.02.037
128. Zhu FC, Guan XH, Li YH, Huang JY, Jiang T, Hou LH, Li JX, Yang BF, Wang L, Wang WJ, et al. 2020. Immunogenicity and safety of a recombinant adenovirus type-5-vectored COVID-19 vaccine in healthy adults aged 18 years or older: A randomised, double-blind, placebo-controlled, phase 2 trial. Lancet [Internet] 396: 479–488. Available from: dx. doi. org/10.1016/S0140-6736 (20) 31605-6
129. Jim Boonyaratanakornkit and Justin J. Taylor. 2019. Techniques to Study Antigen-Specific B Cell Responses. Front. Immunol., 24 July 2019. doi. org/10.3389/fimmu. 2019.01694
130. David D. Chaplin, 2010. Overview of the immune response. J Allergy Clin Immunol. 2010 Feb; 125 (2 Suppl 2) : S3–23. doi: 10.1016/j. jaci. 2009.12.980
131. Bercovici et al., 2000. New Methods for Assessing T-Cell Responses. Clin Diagn Lab Immunol. 2000 Nov; 7 (6) : 859–864. doi: 10.1128/cdli. 7.6.859-864.2000
132. Powell &Newman, eds., Vaccine Design (the subunit and adjuvant approach) (1995) .
133. Rolland, Crit. Rev. Therap. Drug Carrier Systems 15: 143-198 (1998)
134. Fisher-Hoch et al., Proc. Natl. Acad. Sci. USA 86: 317-321 (1989)
135. Flexner et al., Ann. N.Y. Acad. Sci. 569: 86-103 (1989)
136. Flexner et al., Vaccine 8: 17-21 (1990)
137. Berkner, Biotechniques 6: 616-627 (1988)
138. Rosenfeld et al., Science 252: 431-434 (1991)
139. Kolls et al., Proc. Natl. Acad. Sci. USA 91: 215-219 (1994)
140. Kass-Eisler et al., Proc. Natl. Acad. Sci. USA 90: 11498-11502 (1993)
141. Guzman et al., Circulation 88: 2838-2848 (1993)
142. Guzman et al., Cir. Res. 73: 1202-1207 (1993)
143. Ulmer et al., Science 259: 1745-1749 (1993) and reviewed by Cohen, Science 259: 1691-1692 (1993)
144. Kobiyama, et al Vaccines, 2013, 1 (3) , 278-292
145. Mosmann &Coffman, Ann. Rev. Immunol. 7: 145-173 (1989)
146. Sato et al., Science 273: 352 (1996) .
147. Pharmaceutical Dosage Forms (vols. 1-3, 1992)
148. Lloyd, The Art, Science and Technology of Pharmaceutical Compounding (1999)
149. Pickar, Dosage Calculations (1999)
150. Bray, B.L., 2003. Large-scale manufacture of peptide therapeutics by chemical synthesis. Nature ReviewsDrug Discovery, 2 (7) , pp. 587-593.

Claims (21)

  1. A peptide of no more than 500 amino acids, comprising (1) at least one T cell epitope set forth in Table 3 or Tables 15-47 or (2) at least one B cell epitope set forth in Table 4, the peptide optionally further comprising at least one heterologous amino acid sequence.
  2. The peptide of claim 1, comprising or consisting of at least one of the T cell epitopes.
  3. The peptide of claim 1, comprising or consisting of at least one of the B cell epitopes.
  4. The peptide of any one of claims 1-3, comprising or consisting of at least one of the T cell epitopes and at least one heterologous amino acid sequence.
  5. The peptide of any one of claims 1-3, comprising or consisting of at least one of the B cell epitopes and at least one heterologous amino acid sequence.
  6. A nucleic acid comprising a polynucleotide sequence encoding the peptide of any one of claims 1-5.
  7. An expression cassette comprising a polynucleotide sequence encoding the peptide of any one of claims 1-5, operably linked to a promoter.
  8. A vector comprising the expression cassette of claim 7.
  9. A host cell comprising the vector of claim 8.
  10. A composition comprising (1) the peptide of any one of claims 1-5, the nucleic acid of claim 6, the expression cassette of claim 7, the vector of claim 8, or the host cell of claim 9; and (2) a pharmaceutically acceptable excipient.
  11. The composition of claim 10, further comprising an adjuvant.
  12. The composition of claim 10, comprising a plurality of peptides each comprising a T cell epitope set forth in Table 3 or Tables 15-47.
  13. A method of eliciting an immune response in a subject in need thereof, the method comprising administering to the subject an effective amount of a composition comprising (1) the peptide of any one of claims 1-5, the nucleic acid of claim 6, the expression cassette of claim 7, or the vector of claim 8.
  14. The method of claim 13, wherein the composition is administered to the subject by a route selected from the group consisting of, subcutaneous, intramuscular, and oral.
  15. The method of claim 13 or 14, wherein the subject is at risk of exposure to SARS-CoV or SARS-CoV-2 infection.
  16. A kit for eliciting an immune response in a subject in need thereof, the kit comprising a first container containing the composition comprising (1) the peptide of any one of claims 1-5, the nucleic acid of claim 6, the expression cassette of claim 7, or the vector of claim 8, optionally an additional container containing a therapeutic agent against SARS-CoV-2.
  17. The kit of claim 16, further comprising at least a second container each containing at least one different composition.
  18. A method for detecting T cell immunity against SARS-CoV-2 in a subject, comprising:
    (1) contacting T cells obtained from the subject with a T cell epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having an HLA allele associated with the epitope; and
    (2) detecting activation of the T cells, thereby detecting presence of T cell immunity against SARS-CoV-2 in the subject.
  19. The method of claim 18, wherein step (2) comprises detection of T cell proliferation or T cell secretion of one or more cytokines.
  20. The method of claim 18, wherein step (2) comprises T cell proliferation assay, flow cytometry, ELISPOT, or ELISA.
  21. The method of claim 18, wherein step (1) comprises contacting T cells obtained from the subject with a composition comprising a plurality of peptides each comprising a T cell  epitope set forth in Table 3 or Tables 15-47 and antigen-presenting cells having HLA alleles associated with each of the plurality of epitopes.
PCT/CN2022/070948 2021-01-11 2022-01-10 Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes Ceased WO2022148455A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/260,767 US20250000965A1 (en) 2021-01-11 2022-01-10 Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202163136175P 2021-01-11 2021-01-11
US63/136,175 2021-01-11
US202163229063P 2021-08-03 2021-08-03
US63/229,063 2021-08-03

Publications (1)

Publication Number Publication Date
WO2022148455A1 true WO2022148455A1 (en) 2022-07-14

Family

ID=82357945

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/070948 Ceased WO2022148455A1 (en) 2021-01-11 2022-01-10 Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes

Country Status (2)

Country Link
US (1) US20250000965A1 (en)
WO (1) WO2022148455A1 (en)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022268916A3 (en) * 2021-06-23 2023-03-02 Ose Immunotherapeutics Pan-coronavirus peptide vaccine
WO2024038157A1 (en) * 2022-08-17 2024-02-22 PMCR GmbH Immunization against coronavirus
WO2024038155A1 (en) * 2022-08-17 2024-02-22 PMCR GmbH Immunization against viral infections disease(s)
CN117659140A (en) * 2023-10-26 2024-03-08 中国人民解放军海军军医大学 New coronavirus HLA-A2 restricted epitope peptide and its application
WO2024085143A1 (en) * 2022-10-17 2024-04-25 国立大学法人 熊本大学 Nucleocapsid-derived antigen peptide, nucleic acid, vector, pharmaceutical composition, hla/antigen peptide complex, and method for detecting t cell
LU103078B1 (en) * 2023-02-28 2024-08-28 PMCR GmbH IMMUNIZATION AGAINST CORONAVIRUS

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119798385A (en) * 2025-01-03 2025-04-11 中国人民解放军军事科学院军事医学研究院 A HLA-B restricted SARS-CoV-2 lymphocyte antigen epitope peptide and its application
CN120081914A (en) * 2025-01-03 2025-06-03 中国人民解放军军事科学院军事医学研究院 A HLA-C restricted SARS-CoV-2 lymphocyte antigen epitope peptide and its application

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111228483A (en) * 2020-03-19 2020-06-05 四川大学 Broad-spectrum antibody spray for novel coronavirus and SARS virus
CN111358943A (en) * 2020-03-03 2020-07-03 重庆医科大学附属永川医院 Double-targeting immune enhancement type multivalent vaccine of novel coronavirus and preparation method thereof
CN111440229A (en) * 2020-04-13 2020-07-24 中国人民解放军军事科学院军事医学研究院 Novel coronavirus T cell epitope and application thereof
CN111939250A (en) * 2020-08-17 2020-11-17 郑州大学 Novel vaccine for preventing COVID-19 and preparation method thereof
CN112194711A (en) * 2020-10-15 2021-01-08 深圳市疾病预防控制中心(深圳市卫生检验中心、深圳市预防医学研究所) B cell linear epitope of novel coronavirus S protein, antibody, identification method and application

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20230160890A1 (en) * 2018-12-19 2023-05-25 The Board Of Trustees Of The Leland Stanford Junior University Compositions and methods for diagnosing narcolepsy

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111358943A (en) * 2020-03-03 2020-07-03 重庆医科大学附属永川医院 Double-targeting immune enhancement type multivalent vaccine of novel coronavirus and preparation method thereof
CN111228483A (en) * 2020-03-19 2020-06-05 四川大学 Broad-spectrum antibody spray for novel coronavirus and SARS virus
CN111440229A (en) * 2020-04-13 2020-07-24 中国人民解放军军事科学院军事医学研究院 Novel coronavirus T cell epitope and application thereof
CN111939250A (en) * 2020-08-17 2020-11-17 郑州大学 Novel vaccine for preventing COVID-19 and preparation method thereof
CN112194711A (en) * 2020-10-15 2021-01-08 深圳市疾病预防控制中心(深圳市卫生检验中心、深圳市预防医学研究所) B cell linear epitope of novel coronavirus S protein, antibody, identification method and application

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022268916A3 (en) * 2021-06-23 2023-03-02 Ose Immunotherapeutics Pan-coronavirus peptide vaccine
WO2024038157A1 (en) * 2022-08-17 2024-02-22 PMCR GmbH Immunization against coronavirus
WO2024038155A1 (en) * 2022-08-17 2024-02-22 PMCR GmbH Immunization against viral infections disease(s)
WO2024085143A1 (en) * 2022-10-17 2024-04-25 国立大学法人 熊本大学 Nucleocapsid-derived antigen peptide, nucleic acid, vector, pharmaceutical composition, hla/antigen peptide complex, and method for detecting t cell
LU103078B1 (en) * 2023-02-28 2024-08-28 PMCR GmbH IMMUNIZATION AGAINST CORONAVIRUS
CN117659140A (en) * 2023-10-26 2024-03-08 中国人民解放军海军军医大学 New coronavirus HLA-A2 restricted epitope peptide and its application

Also Published As

Publication number Publication date
US20250000965A1 (en) 2025-01-02

Similar Documents

Publication Publication Date Title
WO2022148455A1 (en) Identification and uses of peptide sequences of sars-cov-2 t cell and b cell epitopes
CN111892648B (en) Novel coronavirus polypeptide vaccine coupled with TLR7 agonist and application thereof
JP6395855B2 (en) Porcine epidemic diarrhea virus vaccine
JP6178336B2 (en) Synthetic peptide-based marker vaccine and diagnostic system for effective control of porcine genital respiratory syndrome (PRRS)
Hoque et al. Genomic diversity and evolution, diagnosis, prevention, and therapeutics of the pandemic COVID-19 disease
Zhang et al. Current advancements and potential strategies in the development of MERS-CoV vaccines
CN117957016A (en) SARS-COV-2 and influenza combined vaccine
CN114096675A (en) Coronavirus immunogenic compositions and uses thereof
US9932372B2 (en) Designer peptide-based PCV2 vaccine
CN113801207B (en) Tandem epitope peptide vaccine for novel coronavirus and its application
Zhou et al. Development of variant‐proof severe acute respiratory syndrome coronavirus 2, pan‐sarbecovirus, and pan‐β‐coronavirus vaccines
WO2003087129A2 (en) Immunogenic peptides, and method of identifying same
WO2022003119A1 (en) Cross-reactive coronavirus vaccine
Yashvardhini et al. Immunoinformatics Identification of B‐and T‐Cell Epitopes in the RNA‐Dependent RNA Polymerase of SARS‐CoV‐2
WO2022090679A1 (en) Coronavirus polypeptide
Zhang et al. A novel linker-immunodominant site (LIS) vaccine targeting the SARS-CoV-2 spike protein protects against severe COVID-19 in Syrian hamsters
TW200538153A (en) Feline calicivirus vaccines
Jia et al. Effective preparation and immunogenicity analysis of antigenic proteins for prevention of porcine enteropathogenic coronaviruses PEDV/TGEV/PDCoV
TWI329515B (en)
KR101845571B1 (en) Marker vaccine for classical swine fever
Dong et al. Immune responses of mice immunized by DNA plasmids encoding PCV2 ORF 2 gene, porcine IL-15 or the both
Bhatnagar et al. Molecular targets for diagnostics and therapeutics of severe acute respiratory syndrome (SARS-CoV)
CN118159288A (en) Virus-like particle vaccines for coronavirus
US20140024017A1 (en) IDENTIFICATION OF A NOVEL HUMAN POLYOMAVIRUS (IPPyV) AND APPLICATIONS
CA3161222A1 (en) Epitopic vaccine for african swine fever virus

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22736616

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 22736616

Country of ref document: EP

Kind code of ref document: A1