EP4670167A1 - SYSTEMS AND METHODS FOR MANIPULATING SYNTHETIC ANTIGENS TO PROMOTE TAILORED IMMUNE RESPONSES - Google Patents
SYSTEMS AND METHODS FOR MANIPULATING SYNTHETIC ANTIGENS TO PROMOTE TAILORED IMMUNE RESPONSESInfo
- Publication number
- EP4670167A1 EP4670167A1 EP24713301.0A EP24713301A EP4670167A1 EP 4670167 A1 EP4670167 A1 EP 4670167A1 EP 24713301 A EP24713301 A EP 24713301A EP 4670167 A1 EP4670167 A1 EP 4670167A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- antigen
- rna
- mutations
- engineered
- sars
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K39/00—Medicinal preparations containing antigens or antibodies
- A61K39/12—Viral antigens
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61P—SPECIFIC THERAPEUTIC ACTIVITY OF CHEMICAL COMPOUNDS OR MEDICINAL PREPARATIONS
- A61P31/00—Antiinfectives, i.e. antibiotics, antiseptics, chemotherapeutics
- A61P31/12—Antivirals
- A61P31/14—Antivirals for RNA viruses
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2770/00—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssRNA viruses positive-sense
- C12N2770/00011—Details
- C12N2770/20011—Coronaviridae
- C12N2770/20034—Use of virus or viral component as vaccine, e.g. live-attenuated or inactivated virus, VLP, viral protein
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
Definitions
- Vaccination can play a pivotal role in managing and ensuring public health.
- individuals When vaccinated with sufficiently effective immunogenic compositions against a particular infectious agent, individuals experience reduced risk of infection and/or reduced severity of disease in the event of an in infection. Accordingly, development of highly effective vaccination techniques that can keep pace with ever evolving valiants of circulating pathogens and newly emergent diseases is a critical challenge.
- methods and system of the present disclosure provide for engineering antigens to reduce their activation of a memory immune response, such as B cell and/or T cell based response, when introduced into a subject.
- Designing antigens in this manner may, for example, improve their performance as immunogenic compositions for purposes of vaccination.
- reducing the extent to which a memory immune response is triggered can lead to improved production of new antibodies that are selectively tailored by the subject’s immune system to neutralize particular (e.g., arisen) epitopes of a reference antigen.
- systems and methods of the present disclosure identify, within computer representations of a reference antigen, conserved regions that are similar to regions of other, for example previously circulating, variants of the reference antigen and thus likely to trigger a memory response. Approaches described herein then disrupt these conserved region(s), for example by introducing amino acid modifications within I across them. In this manner, an engineered antigen can be generated that retains certain portions of the reference antigen, such particular target epitopes, but replaces a conserved region with disrupted versions thereof.
- engineered antigens with disrupted conserved region(s) when manufactured and introduced into a subject arc less likely to trigger a memory immune response, e.g., from memory B or T cells, and, instead, encourage naive responses and, accordingly, production of new neutralizing antibodies that are expressly tailored to the retained target epitopes.
- Such engineered antigens may, accordingly, offer improved efficacy when used as immunogenic compositions, particularly for viral infectious agents that are prone to mutation.
- the present disclosure provides methods for in-silico design of an engineered antigen [e.g., for (e.g., characterized in that) eliciting an immune response directed to one or more target epitopes of a reference antigen of an infectious agent while reducing the engineered antigen’s activation of a (e.g., B cell and/or T cell) memory immune response (e.g., relative to) to the reference antigen], the method comprising: (a) receiving and/or accessing, by a processor of a computing device, a polypeptide model representing (e.g., as a sequence of and/or 3D structural model of) a reference antigen of an infectious agent; (b) identifying, by the processor, within the polypeptide model, one or more memory-triggering conserved region(s) representing conserved portions of the reference antigen that are determined likely to trigger a memory immune response [e.g., portions of the reference antigen that are determined to (i) correspond to known epitopes and
- a reference antigen is or comprises at least a portion of a naturally occurring variant of a viral protein [e.g., wherein the infectious agent is a variant of a particular virus (e.g., an influenza virus, a coronavirus, a respiratory syncytial virus, a filovirus) and wherein the reference antigen is or comprises at least a portion of a protein thereof],
- a particular virus e.g., an influenza virus, a coronavirus, a respiratory syncytial virus, a filovirus
- a reference antigen (e.g., and/or viral protein) is or comprises at least a portion of a SARS-Cov2 Spike polypeptide [e.g., a Receptor Binding Domain (RBD); e.g., an N-Terminal region; e.g., substantially all of a (e.g., an entire) Spike protein] [e.g., a portion selected to focus on a minimal relevant vaccine antigen, e.g., to facilitate removal of as many conserved epitopes as possible without e.g., resorting to introducing point mutations (e.g., thereby limiting number of epitopes in which point mutations are to be introduced)] .
- SARS-Cov2 Spike polypeptide e.g., a Receptor Binding Domain (RBD); e.g., an N-Terminal region; e.g., substantially all of a (e.g., an entire) Spike protein] [e.g.
- a reference antigen is or comprises at least a portion of a particular SARS-CoV-2 variant Spike polypeptide.
- a particular SARS-CoV-2 variant is a member of an Omicron and/or XBB lineage classification (e.g., according to a WHO, Pango, Nextstrain, etc. classification) (e.g., wherein the particular SARS-CoV-2 variant is XBB.1.5; e.g., wherein the particular SARS-CoV-2 variant is JN.l).
- Omicron and/or XBB lineage classification e.g., according to a WHO, Pango, Nextstrain, etc. classification
- an infectious agent is or comprises (e.g., a particular variant of) an RNA virus (e.g., a virus that encodes its genetic information with RNA) and the reference antigen is or comprises at least a portion of a protein thereof.
- an RNA virus e.g., a virus that encodes its genetic information with RNA
- the reference antigen is or comprises at least a portion of a protein thereof.
- a reference antigen is or comprises a bacterial protein (e.g., wherein the infectious agent is a bacteria and the reference antigen is or comprises at least a portion of a particular protein thereof).
- a reference antigen is or comprises an antigen (e.g., a surface antigen; e.g., a protein) of a parasite (e.g., wherein the infectious agent is a parasite and the reference antigen is or comprises at least a portion of a particular protein thereof) (e.g., wherein the infectious agent is a malaria parasite and the reference antigen is an antigen thereof).
- one or more memory -triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar’ to (i) one or more (e.g., pre-existing) variants thereof and/or (ii) an initial/wild-type strain (e.g., a first-observed strain) ⁇ e.g., wherein the one or more memory-triggering conserved region(s) represent portion(s) of the reference antigen having sufficient sequence similarity [e.g., at least 80% (including, e.g., at least 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher] identical; e.g., identical] to corresponding portions of (i) one or more (e.g., pre-existing) variants thereof and/or (ii) an initial/wild-type strain (e.g., a first-observed strain) ⁇ .
- sufficient sequence similarity e.g., at least 80% (including,
- a reference antigen is a particular target SARS-CoV-2 variant (e.g., XBB.1.5; e.g., JN.l) S polypeptide or portion thereof (e.g., wherein the reference antigen is an RBD of the target SARS-CoV-2S protein) and wherein the one or more memorytriggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar to corresponding portion(s) of (i) one or more (e.g., pre-existing) other SARS-CoV-2 variant polypeptides and/or (ii) Wuhan SARS-CoV-2 polypeptide [e.g., wherein the one or more memory-triggering conserved region(s) represent unmutated portion(s) of the reference antigen that are common to the reference antigen and corresponding portions of (i) the one or more (e.g., pre-existing) SARS-CoV-2 variant polypeptide) s) and/or (ii) Wuhan SARS-CoV-2 poly
- one or more memory -triggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and unmutated on the reference antigen (e.g., known epitopes without any hallmark mutations) [e.g., wherein the reference antigen is a particular sub-region (e.g., an RBD) of a target SARS-CoV-2 variant S protein (e.g., XBB.1.5; e.g., JN.l) and wherein the set of conserved epitope regions represent known epitopes that are present and un-muted on the particular sub-region of the target SARS-CoV-2 variant S protein relative corresponding sub-regions of SARS-CoV-2 S proteins of to one or more pre-existing variants and/or Wuhan strain].
- the reference antigen is a particular sub-region (e.g., an RBD) of a target SARS-CoV-2 variant S protein (e.g., XBB.1.5;
- step (b) comprises: obtaining, by the processor, data corresponding to a set of known epitopes and identifying, within the reference antigen, each of one or more particular known epitopes of the set; obtaining, by the processor, an identification of a set of hallmark mutations of the reference antigen; and identifying, by the processor, as the set of conserved epitope regions, those particular known epitopes that correspond to portions of the reference antigen without any hallmark mutations.
- a set of known epitopes comprise one or more of the epitopes listed in Table 2A.
- a set of known epitopes comprise one or more of the epitopes listed in Table 2B.
- one or more memory -triggering conserved region(s) are or comprise a conserved surface representing a (e.g., continuous; e.g., contiguous) un-mutated (e.g., lacking any hallmark mutations) surface of the reference antigen [e.g., wherein the reference antigen is or comprises an XBB.1.5 variant of an RBD of a SARS-Cov-2 S protein and the conserved surface is or comprises at least a portion (e.g., up to all) of the following positions: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y38O, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412,
- provided methods comprise identifying, and/or accessing an identification of (e.g., positions of), a set of one or more hallmark mutations of the reference antigen [e.g., wherein the reference antigen is a protein of a particular viral variant classification and the set of hallmark mutations comprises those mutations of the reference antigen that are present at or above a particular threshold rate (e.g., appearing in sequences identified as belonging to the viral variant classification at or above the threshold rate (e.g., a fraction, percentage, etc.))].
- a particular threshold rate e.g., appearing in sequences identified as belonging to the viral variant classification at or above the threshold rate (e.g., a fraction, percentage, etc.)
- a reference antigen is or comprises a SARS-Cov2 XBB.1.5 Spike protein ⁇ e.g., a SARS-CoV-2 S protein with XBB.1.5 hallmark mutations [e.g., comprising at least a portion (e.g., all) of the following mutations: T191, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N6
- step (c) comprises selecting, by the processor the one or more amino acid modifications from a set of allowed mutations [e.g., selecting a location and/or particular modification (e.g., substitution, deletion, insertion) from the set of allowed mutations].
- a set of allowed mutations is or comprises (e.g., a list, table, etc. representing) a plurality of mutations observed as occurring (e.g., present at or above a particular rate) within a set of related antigens.
- a reference antigen is a SARS-CoV-2 (e.g., Spike) protein of a particular variant and the set of related antigens comprises corresponding proteins of other, related, variants (e.g., within a particular set of lineages and/or sub-lineages including the reference antigen).
- SARS-CoV-2 e.g., Spike
- the set of related antigens comprises corresponding proteins of other, related, variants (e.g., within a particular set of lineages and/or sub-lineages including the reference antigen).
- a reference antigen is a member of an Omicron lineage and the set of related antigens comprises observed variants belonging to the Omicron lineage.
- a set of allowed mutations comprises at least a portion (e.g., a subset; e.g., all) of mutations listed in Table 3 (e.g., excluding those removed, as marked with strikethrough in Table 3) (e.g., and, optionally, one or more additional mutations; e.g., and not any other mutations).
- a set of allowed mutations comprises at least a portion (e.g., a subset) of mutations listed in Table 3, excluding one or both of L335F and L390R.
- a reference antigen is a SARS-CoV2 (e.g., Spike) protein and the set of related antigens comprises corresponding proteins of other (e.g., SARS-CoV 1, MERS) coronavirus.
- SARS-CoV2 e.g., Spike
- MERS MERS
- one or more memory -triggering conserved regions are or comprise a set of conserved epitope regions and step (c) comprises introducing at least one amino acid modification within each of at least a portion (e.g., all; e.g., a particular subset) of the conserved epitope regions.
- one or more memory -triggering conserved regions are or comprise a conserved surface and step (c) comprises generating the one or more amino acid modifications at positions distributed (e.g., approximately evenly, as opposed to clustered) throughout / across the conserved surface [e.g., distributed in a substantially equidistant manner over a 3D representation of the conserved surface (e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
- a 3D representation of the conserved surface e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
- provided methods comprise performing steps (b) and (c) repeatedly to generate a plurality of candidate polypeptide models, each representing a candidate engineered variant (e.g., each candidate polypeptide model comprising a distinct combination of amino acid modifications and representing a unique artificially engineered version of the reference antigen).
- provided methods comprise determining, by the processor, values of one or more performance scores for each of the candidate polypeptide models and selecting a subset of the candidate polypeptide models based at least in part on the determined performance score values.
- one or more performance scores comprise one or both of: (a) an immune escape score indicative of a likelihood and/or relative capability of a particular candidate engineered variant to be recognized and neutralized by antibodies, and (b) a fitness score indicative of a likelihood and/or viability of a particular candidate engineered variant.
- determining one or both of (a) an immune escape score and (b) a fitness score comprises using a machine learning model [e.g., a language model that receives, as input, a representation of an amino acid sequence of the particular candidate engineered variant (e.g., wherein the input does not comprise a 3D structural representation of the reference)] [e.g., to compute a likelihood score representing a predicted likelihood of the particular candidate variant occurring; e.g., to compute a semantic change score indicative of a distance between an embedding representation of (i) the particular candidate variant (e.g., generated from one or more hidden layers of the machine learning model) and (ii) one or more reference variants (c.g.., a WT variant; c.g., a variant of the reference antigen with which a subject has previously been infected, e.g., via vaccination and/or natural exposure; e.g., the reference antigen)].
- a machine learning model e.g., a language model that receive
- determining one or both of (a) an immune escape score and (b) a fitness score comprises using a 3D structural model of at least a portion of the particular candidate variant [e.g., computing a viral polypeptide receptor binding score (e.g., an ACE2 binding score); e.g., computing an epitope alteration score],
- one or more performance scores comprise(s) a position spread score that measures an extent to which amino acid modifications are evenly distributed across a surface (e.g., the conserved surface) of the candidate engineered variant.
- one or more performance scores comprise(s) a mutation co-occurrence score (e.g., that measures a degree to which one or more amino acid modifications of a particular candidate variant are aligned with a natural co-occurrence rates).
- a mutation co-occurrence score e.g., that measures a degree to which one or more amino acid modifications of a particular candidate variant are aligned with a natural co-occurrence rates.
- provided methods comprise identifying, by the processor, within the polypeptide model, one or more target regions representing portions (e.g., target epitopes) of the reference antigen to be retained and excluding the one or more target regions from the one or more memory-triggering conserved region(s) (e.g., wherein the reference antigen is a SARS-CoV-2 S protein and/or portion thereof and the one or more target regions are or comprise an ACE2 binding interface, e.g., thereby preserving ACE2 binding interface regions).
- portions e.g., target epitopes
- the reference antigen is a SARS-CoV-2 S protein and/or portion thereof and the one or more target regions are or comprise an ACE2 binding interface, e.g., thereby preserving ACE2 binding interface regions.
- provided methods comprise causing, by the processor, rendering of the disrupted polypeptide model for graphical display.
- provided methods comprise generating, from the disrupted polypeptide model, a corresponding RNA sequence.
- provided methods comprise producing (e.g., as a vaccine) a composition comprising a polypeptide based on (e.g., having a substantially same amino acid sequence as represented by) the disrupted polypeptide model.
- provided methods comprise assessing the biological activity of the engineered antigen in vitro.
- biological activity of the engineered antigen is characterized in that: the engineered antigen is expressed and folded properly [e.g., based on a binding assay, such as a flow-cytometry-based binding assay (e.g., with hACE2)]; and/or the engineered antigen does not bind to antibodies that bind to the reference antigen [e.g., shows reduced/abrogated binding (in comparison with the reference antigen) for a panel of one or more neutralizing antibodies (e.g., antibodies that bind to distinct epitope classes; e.g., one or more (e.g., each of) classes A, B, C, D, E, and F)]; and/or pseudoviruses loaded with the engineered antigen are able to enter cells; and/or the engineered antigen is immunogenic, and/or the engineered antigen reduces the engineered antigen
- provided methods comprise producing (e.g., as a vaccine) a composition comprising a nucleic acid encoding the amino acid sequence represented by the disrupted polypeptide model.
- the present disclosure provides vaccine compositions comprising a polypeptide and/or nucleic acid of one or more aspects or embodiments described herein (e.g., in paragraphs above).
- the present disclosure provides methods of vaccination comprising administering to a subject or a population of subjects provided vaccines according to one or more aspects or embodiments described herein (e.g., in paragraphs above).
- the present disclosure provides systems comprising a processor of a computing device and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method of one or more aspects or embodiments described herein (e.g., in paragraphs above).
- the present disclosure provides methods of manufacturing an immunogenic composition
- methods of manufacturing an immunogenic composition comprising: comparing sequences of viral proteins from different variants of an infectious disease agent (e.g., influenza; e.g., SARS-CoV-2) to identify (residual) conserved sites in an antigen of interest (e.g., via one or more systems and/or methods, including any of those recited in the preceding claims); replacing at least one or more of the residual conserved sites with a sequence that is characterized by [features/characteristics of such amino acid substitutions, e.g., providing new immunogenic epitopes from different variants] to generate a new sequence (e.g., via one or more systems and/or methods, including any of those recited in the preceding claims); and producing a vaccine that delivers at least a portion of the new sequence including at least one of the residual conserved sites replaced.
- an infectious disease agent e.g., influenza; e.g., SARS-CoV-2
- RNA comprising a nucleotide sequence that encodes an engineered antigen [e.g., for (e.g., characterized in that) eliciting an immune response directed to one or more target epitopes of a reference antigen of an infectious agent while reducing the engineered antigen’s activation of a (e.g., B cell and/or T cell) memory immune response to the reference antigen], wherein the engineered antigen corresponds to (e.g., has a sequence of) a particular reference antigen having been (e.g., artificially) altered to introduce one or more amino acid modifications within at least a portion of one or more memory-triggering conserved regions having been identified as portions of the reference antigen that are determined likely to trigger a memory immune response [e.g., portions that are determined to (i) correspond to known epitopes and/or potential epitopes and/or (ii) do not comprise hallmark mutations of the reference antigen] [e.g., such
- one or more memory -triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar (e.g., having sufficient sequence similarity; e.g., common) to one or more (e.g., pre-existing) variants thereof.
- one or more memory -triggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and unmutated on the reference antigen (e.g., known epitopes without any hallmark mutations).
- a set of known epitopes comprise one or more of the epitopes listed in Tables 2A and/or 2B.
- one or more memory -triggering conserved region(s) are or comprise a conserved surface representing a (e.g., continuous; e.g., contiguous) un-mutated (e.g., lacking any hallmark mutations) surface of the reference antigen [e.g., wherein the target polypeptide is or comprises an XBB.1.5 variant of an RBD of a SARS-CoV-2 Spike (S) protein and the conserved surface is or comprises the following positions: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y38O, G381 , S383, P384, T385, K386, N388, D389, L390, C391 , F392, T393, N394, Y396, P412, G413, Q414, T
- a reference antigen is or comprises at least a portion of (e.g., an RBD of) an XBB.1.5 variant of a SARS-Cov2 Spike protein [e.g., and the hallmark mutations comprise at least a portion (e.g., all) of the following mutations: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679
- one or more amino acid modifications are selected from a set of allowed mutations [e.g., selecting a location and/or particular modification (e.g., substitution, deletion, insertion) from the set of allowed mutations].
- a set of allowed mutations is or comprises a plurality of mutations observed as occurring (e.g., present at or above a particular rate) within a set of related polypeptides.
- a set of related polypeptides comprises corresponding polypeptides of other, related, SARS-CoV 2 variants (e.g., within a particular set of lineages and/or sub-lineages including the particular SARS-CoV 2 variant).
- a target (SARS-CoV 2 variant) polypeptide is a member of an Omicron lineage and the set of related polypeptides comprises corresponding polypeptides (e.g., Spike protein sequences) of observed variants belonging to the Omicron lineage.
- one or more memory -triggering conserved regions are or comprise a set of conserved epitope regions and the engineered antigen has at least one amino acid modification within each of at least a portion (e.g., all; e.g., a particular subset) of the conserved epitope regions.
- one or more memory-triggering conserved regions are or comprise a conserved surface and the one or more amino acid modifications occur at positions distributed (e.g., approximately evenly, as opposed to clustered) throughout / across the conserved surface [e.g., distributed in a substantially equidistant manner over a 3D representation of the conserved surface (e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
- a 3D representation of the conserved surface e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, second or third column) in Table 5A [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in the third column of Table 5A)][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the sequences listed in (e.g., each row of) Table 5B [e.g., wherein the engineered antigen is or comprises a SARS- CoV-2 S protein RBD having any one of the sequences listed in Table 5B (e.g., any one of SEQ ID NOs: 7 through 34)].
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, first or second column) in Table 7A and/or Table 7B
- the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 7A and/or Table 7B (e.g., identified as LI
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, second or third column) in Table 8A [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in third column of Table 8A][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the sequences listed in (e.g., each row of) Table 8B [e.g., wherein the engineered antigen is or comprises a SARS- CoV-2 S protein RBD having any one of the sequences listed in Table 8B (e.g., any one of SEQ ID NOs: 39 through 64 and 100)].
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, second column) in Table 9A [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 9A].
- the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S43 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 S protein RBD with additional mutations according to S43 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.l .5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with additional mutations N360D P384S L390R T430I F464Y H519N].
- S43 in Table 9A e.g., wherein the engineered antigen is or comprises an XBB.1.5 S protein RBD with additional mutations according to S43 in Table 9A
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: N360D, P384S, L390R, T430I, F464Y, H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G3
- a particular domain e
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise the mutation combinations identified as S48 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 RBD with additional mutations according to S48 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the XBB.1.5 hallmark mutations identified in Table IB and) with additional mutations L335F K356T P384S L390R T430I F464Y H519N].
- S48 in Table 9A e.g., wherein the engineered antigen is or comprises an XBB.1.5 RBD with additional mutations according to S48 in Table 9A
- the engineered antigen comprises a SARS-CoV
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: L335F, K356T, P384S, L390R, T4301, F464Y, and H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E,
- the reference antigen is a particular target valiant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigen of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion of a SARS-CoV-2 S protein (e.g., a RBD) with mutations P384S L390R T430I F464Y H519N.
- a SARS-CoV-2 S protein e.g., a RBD
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion of a SARS-CoV-2 S protein (e.g., a RBD) with one or more of (e.g., a subset of; e.g., all of) mutations K356T P384S L390R T430I F464Y H519N.
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] provided methods, systems, vaccine compositions, method of manufacturing, or RNAs do not include one or both of mutations L335F and L390R.
- the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and]
- provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (c.g., in each row of) Table 15A and/or Table 15B [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in Table 15A][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g.,
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 23 in Table 15 A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: I332V L335F K356T P384S T430I L452Q F464Y H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E,
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: I332V L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452Q N460K F464Y S477N T478K E484A F486P F490S Q498R N501 Y Y505H E516Q H519N.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 22 in Table 15A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: L335F K.356T P384S T430I L452R F464Y H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.l .5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E,
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H E516Q H519N T523S.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 09 in Table 15A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T N360S P384S N388K T430I N450D F464Y H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.l.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E,
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S N388K D405N R408S K417N T4301 N440K V445P G446S N450D N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N T523S.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S129 in Table 15 A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T L335F P384S D389G T430I N450D F464Y H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 56 in Table 15 A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T P384S L390R T430I N450D F464Y I472V H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S L390R D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y I472V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 12 in Table 15A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T P384S D389G T430I N450D F464Y I468V H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (c.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G D405N R408S K417N T430I N440K V445P G446S N450D N460K F464Y I468V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S125 in Table 15A and/or Table 15B.
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T L335F P384S T430I F464Y I468V H519N ⁇ e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252
- engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S N460K F464Y I468V S477N T478K E484A F486P F490S Q498R N501 Y Y505H E516Q H519N T523S.
- the present disclosure provides methods of manufacturing an RNA comprising a nucleotide sequence that encodes an engineered antigen corresponding to an engineered version of a reference antigen, the method comprising producing an RNA whose nucleotide sequence, when compared with (e.g., a nucleotide sequence of) the reference antigen shows difference(s) (e.g., comprises nucleotides encoding for one or more mutations) relative to the reference antigen in one or more memory-triggering conserved regions [e.g., one or more conserved epitopes and/or one or more conserved surface(s)] that are common to the reference antigen and (i) one or more pre-existing variants of the reference antigen and/or (ii) a wild-type strain of the reference antigen.
- difference(s) e.g., comprises nucleotides encoding for one or more mutations
- a reference antigen is or comprises a particular target variant SARS-CoV-2 S protein RBD [e.g., an XBB.1.5 RBD (e.g., having an amino acid sequence according to any one of SEQ ID NOs: 35, 36, and 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the XBB.1.5 hallmark mutations identified in Table IB)].
- a particular target variant SARS-CoV-2 S protein RBD e.g., an XBB.1.5 RBD (e.g., having an amino acid sequence according to any one of SEQ ID NOs: 35, 36, and 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the XBB.1.5 hallmark mutations identified in Table IB)].
- RNA manufactured according to provided methods is or comprises the RNA of any one the aspects and embodiments described herein (e.g., wherein the differences in comparison with the reference antigen correspond/comprise nucleotide differences encoding for any of the mutation combinations of various aspects and embodiments described herein, e.g., in paragraphs above).
- provided methods of manufacturing comprise producing RNA via in vitro transcription (IVT).
- IVTT in vitro transcription
- FIG. 1 is a block-flow diagram and schematic illustrating an example process for design of an engineered antigen, according to an illustrative embodiment.
- FIG. 2 is a block-flow diagram of an example process for inserting amino acid modifications, according to an illustrative embodiment.
- FIG. 3 is a block-flow diagram of an example process for generating and scoring multiple candidate antigen designs, according to an illustrative embodiment.
- FIG. 4A is a set of views of a 3D model of a receptor-binding domain (RBD) of a SARS-Cov2 Spike (.S') protein of an XBB variant, colorized to show XBB hallmark mutations and an identified conserved region, according to an illustrative embodiment.
- RBD receptor-binding domain
- S' SARS-Cov2 Spike
- FIG. 4B is another set of views of the XBB RBD .S' protein model shown in FIG. 6A, colorized to show XBB hallmark mutations, an identified conserved region, and an ACE2 binding interface, according to an illustrative embodiment.
- FIG. 4C is another set of views of the XBB RBD .S' protein model shown in FIGs. 6 A and 6B, colorized to show a relative frequency of epitope involvement at various amino acid sites within the protein, according to an illustrative embodiment.
- FIG. 4D is a set of views of a 3D model of a first engineered antigen design, generated using the XBB RBD 5 protein as a reference antigen.
- FIG. 4E is a set of views of a 3D model of a second engineered antigen design, generated using the XBB RBD ,S' protein as a reference antigen.
- FIG. 5A is a schematic of SARS CoV 2 and its relevant proteins, adapted from Jain el al., Vaccines 2020, 8(4), 649 (www.mdpi.com/2076-393X/8/4/649.).
- FIG. 5B is a set of views of a 3D model of a receptor-binding domain (RBD) of a SARS-Cov2 Spike (.S') protein of an XBB. 1 .5 variant, colorized to show XBB hallmark mutations and an identified conserved region, according to an illustrative embodiment.
- RBD receptor-binding domain
- S' SARS-Cov2 Spike
- FIG. 5C is a set of views of a 3D model of an XBB.1.5 spike protein (full spike protein), according to an illustrative embodiment.
- FIG. 5D is another set of views of the XBB RBD 5 protein model shown in FIG. 7B, colorized to show XBB hallmark mutations, an identified conserved region, and an ACE2 binding interface, according to an illustrative embodiment.
- FIG. 5E is another set of views of the XBB RBD 5 protein model shown in FIGs. 5B and 5D, colorized to show a relative frequency of epitope involvement at various amino acid sites within the protein, according to an illustrative embodiment.
- FIG. 5F is another set of views of the XBB RBD S protein model shown in FIGs. 5B and 5D, colorized to show a relative frequency of epitope involvement at various amino acid sites within a conserved surface of the protein, according to an illustrative embodiment.
- FIG. 5G is a block flow diagram of an example process for generating candidate variant designs, according to an illustrative embodiment.
- FIGs. 6A shows a set of views of a 3D model of candidate engineered antigen design, generated using an XBB.1.5 RBD 5 protein as a reference antigen, according to an illustrative embodiment.
- Colorization/shading in FIG. 6A identifies XBB.1.5 hallmark mutations, an identified conserved region, and amino acid sites mutated based on various in- silico antigen design approaches described herein.
- FIG. 6B shows a set of views of a 3D model of a SARS-CoV-2 spike protein with XBB.1.5 hallmark mutations and additional mutations corresponding to those shown in FIG. 6A highlighted.
- FIG. 6C shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
- FIG. 6D shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
- FIG. 6E shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
- FIG. 6F shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 .S' protein RBD as a reference antigen, according to an illustrative embodiment.
- FIG. 7 shows radar plots of certain performance scores for candidate engineered antigen designs, according to an illustrative embodiment.
- FIG. 8 shows a plot of predicted versus experimental ACE2 binding change, according to an illustrative embodiment.
- FIG. 9A shows a view of a 3D model for an engineered RBD construct, according to an illustrative embodiment .
- FIG. 9B shows a view of a 3D model for an engineered RBD construct, according to an illustrative embodiment.
- FIG. 9C is a schematic of a SARS-CoV-2 trimer, adapted from Starr et al., SARS- CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021 ). https://doi.org/! 0.1038/s41586-021 -03807-6.
- FIG. 9D shows a view of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes on an XBB.1.5 backbone, according to an illustrative embodiment.
- FIG. 9E shows a view of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes on an XBB.1.5 backbone, according to an illustrative embodiment.
- FIG. 10A is a diagram of an example procedure for testing engineered antigens, according to an illustrative embodiment.
- FIG. 10B is a block-flow diagram of an example protocol for generating and evaluating engineered constructs as described herein, according to an illustrative embodiment.
- FIG. 10C is a set of graphs evaluating antibody binding to WT and XBB.1.5 S protein for different epitope classes, according to an illustrative embodiment.
- FIG. 10D is a graph of a dilution series for a hACE2-mFc binding agent.
- FIG. 10E is a graph of a dilution series of human BNT162b2 3 (triple-vax) polyclonal serum.
- FIG. 11 A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 11B shows binding dose response curves for certain engineered SARS-CoV
- FIG. 11C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 12A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 12B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 13A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 13B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 13C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 13D shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 14A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 14B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 14C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 14D shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
- FIG. 15A is a schematic illustrating example groups and a dosage / sampling schedule for evaluating immune imprinting and performance of candidate vaccines in vaccine- cxpcricnccd mice. Yellow-filled cells indicate days on which sera sample will be collected, gray-filled cells indicate days on which vaccines will be administered, and green-filled cells indicate days on which mice are sacrificed and final samples collected, according to an illustrative embodiment.
- FIG. 15B is a schematic illustrating example groups and a dosage I sampling schedule for evaluating performance of candidate vaccines in vaccine naive mice. Yellow-filled cells indicate days on which sera sample will be collected, gray-filled cells indicate days on which vaccines will be administered, and green-filled cells indicate days on which mice are sacrificed and final samples collected, according to an illustrative embodiment.
- FIG. 16A is a schematic illustrating FACS-based assay, e.g., using approaches described in Powell and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety, according to an illustrative embodiment.
- FIG. 16B is a schematic illustrating depletion assay, e.g., using approaches described in Trocet and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety, according to an illustrative embodiment.
- FIG. 17 is a schematic illustrating approaches for collecting and analyzing spleen sample, lymph nodes as well as analysis of blood samples, according to certain embodiments.
- FIG. 18 is a diagram illustrating mutations included in various engineered antigen construct designs, according to an illustrative embodiment.
- FIG. 19A is a schematic illustrating an example construct comprising a membrane anchored RBD+fibritin domain with viral signal peptide, according to an illustrative embodiment.
- FIG. 19B is a graph showing ACE2 binding assay results for a set of engineered construct designs, according to an illustrative embodiment.
- FIG. 19C is a graph showing serum (from vaccinated patients) binding assay results for a set of engineered construct designs, according to an illustrative embodiment.
- FIG. 20A is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
- FIG. 20B is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
- FIG. 20C is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
- FIG. 21 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.
- FIG. 22 is a diagram of an example computing device and an example mobile computing device used in certain embodiments.
- Administration typically refers to the administration of a composition to a subject or system.
- routes that may, in appropriate circumstances, be utilized for administration to a subject, for example a human.
- administration may be ocular, oral, parenteral, topical, etc.
- administration may be bronchial (e.g., by bronchial instillation), buccal, dermal (which may be or comprise, for example, one or more of topical to the dermis, intradermal, interdermal, transdermal, etc.), enteral, intra-arterial, intradermal, intragastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intraventricular, within a specific organ (e. g. intrahepatic), mucosal, nasal, oral, rectal, subcutaneous, sublingual, topical, tracheal (e.g., by intratracheal instillation), vaginal, vitreal, etc.
- bronchial e.g., by bronchial instillation
- buccal which may be or comprise, for example, one or more of topical to the dermis, intradermal, interdermal, transdermal, etc.
- enteral intra-arterial, intradermal, intragas
- administration may involve dosing that is intermittent (e.g., a plurality of doses separated in time) and/or periodic (e.g., individual doses separated by a common period of time) dosing. In some embodiments, administration may involve continuous dosing (e.g., perfusion) for at least a selected period of time.
- adult refers to a human eighteen years of age or older. In some embodiments, a human adult has a weight within the range of about 90 pounds to about 250 pounds.
- affinity is a measure of the tightness with which two or more binding partners associate with one another. Those skilled in the art are aware of a variety of assays that can be used to assess affinity, and will furthermore be aware of appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity is assessed over a plurality of concentrations (e.g., of one binding partner at a time). In some embodiments, affinity is assessed in the presence of one or more potential competitor entities (e.g., that might be present in a relevant - e.g., physiological - setting).
- affinity is assessed relative to a reference (e.g., that has a known affinity above a particular threshold [a “positive control” reference] or that has a known affinity below a particular threshold [ a “negative control” reference”].
- affinity may be assessed relative to a contemporaneous reference; in some embodiments, affinity may be assessed relative to a historical reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
- agent is used to refer to an entity (e.g., for example, a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or complex, combination, mixture or system [e.g., cell, tissue, organism] thereof), or phenomenon (e.g., heat, electric current or field, magnetic force or field, etc.).
- entity e.g., for example, a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or complex, combination, mixture or system [e.g., cell, tissue, organism] thereof), or phenomenon (e.g., heat, electric current or field, magnetic force or field, etc.).
- the term may be utilized to refer to an entity that is or comprises a cell or organism, or a fraction, extract, or component thereof.
- the term may be used to refer to a natural product in that it is found in and/or is obtained from nature.
- the term may be used to refer to one or more entities that is man-made in that it is designed, engineered, and/or produced through action of the hand of man and/or is not found in nature.
- an agent may be utilized in isolated or pure form; in some embodiments, an agent may be utilized in crude form.
- potential agents may be provided as collections or libraries, for example that may be screened to identify or characterize active agents within them.
- the term “agent” may refer to a compound or entity that is or comprises a polymer; in some cases, the term may refer to a compound or entity that comprises one or more polymeric moieties. In some embodiments, the term “agent” may refer to a compound or entity that is not a polymer and/or is substantially free of any polymer and/or of one or more particular polymeric moieties. In some embodiments, the term may refer to a compound or entity that lacks or is substantially free of any polymeric moiety.
- Amelioration- refers to the prevention, reduction or palliation of a state, or improvement of the state of a subject. Amelioration includes, but does not require complete recovery or complete prevention of a disease, disorder or condition (e.g., radiation injury).
- amino acid refers to a compound and/or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds.
- an amino acid has the general structure H2N-C(H)(R)-COOH.
- an amino acid is a naturally-occurring amino acid.
- an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid.
- Standard amino acid refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides.
- Nonstandard amino acid refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source.
- an amino acid including a carboxy-and/or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structure above.
- an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and/or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and/or the hydroxyl group) as compared with the general structure.
- such modification may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid.
- amino acid may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.
- animal refers to any member of the animal kingdom. In some embodiments, “animal” refers to humans, of either sex and at any stage of development. In some embodiments, “animal” refers to non-human animals, at any stage of development. In certain embodiments, the non-human animal is a mammal (e.g., a rodent, a mouse, a rat, a rabbit, a monkey, a dog, a cat, a sheep, cattle, a primate, and/or a pig). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and/or worms. In some embodiments, an animal may be a transgenic animal, genetically engineered animal, and/or a clone.
- mammal e.g., a rodent, a mouse, a rat, a rabbit, a monkey, a dog, a cat, a sheep, cattle, a primate, and/or
- Antibody refers to a polypeptide that includes canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular target antigen.
- intact antibodies as produced in nature are approximately 150 kD tetrameric agents comprised of two identical heavy chain polypeptides (about 50 kD each) and two identical light chain polypeptides (about 25 kD each) that associate with each other into what is commonly referred to as a “Y-shaped” structure.
- Each heavy chain is comprised of at least four domains (each about 110 amino acids long)- an amino-terminal variable (VH) domain (located at the tips of the Y structure), followed by three constant domains: CHI, CH2, and the carboxy -terminal CH3 (located at the base of the Y’s stem).
- VH amino-terminal variable
- CHI amino-terminal variable
- CH2 amino-terminal variable
- CH3 located at the base of the Y’s stem
- a short region known as the “switch” connects the heavy chain variable and constant regions.
- the “hinge” connects CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect the two heavy chain polypeptides to one another in an intact antibody.
- Each light chain is comprised of two domains - an amino-terminal variable (VL) domain, followed by a carboxy -terminal constant (CL) domain, separated from one another by another “switch”.
- Intact antibody tetramers are comprised of two heavy chain-light chain dimers in which the heavy and light chains are linked to one another by a single disulfide bond; two other disulfide bonds connect the heavy chain hinge regions to one another, so that the dimers arc connected to one another and the tetramer is formed.
- Naturally-produced antibodies are also glycosylated, typically on the CH2 domain.
- Each domain in a natural antibody has a structure characterized by an “immunoglobulin fold” formed from two beta sheets (e.g., 3-, 4-, or 5- stranded sheets) packed against each other in a compressed antiparallel beta barrel.
- Each variable domain contains three hypervariable loops known as “complement determining regions” (CDR1, CDR2, and CDR3) and four somewhat invariant “framework” regions (FR1, FR2, FR3, and FR4).
- CDR1, CDR2, and CDR3 three hypervariable loops known as “complement determining regions” (CDR1, CDR2, and CDR3) and four somewhat invariant “framework” regions (FR1, FR2, FR3, and FR4).
- the Fc region of naturally-occurring antibodies binds to elements of the complement system, and also to receptors on effector cells, including for example effector cells that mediate cytotoxicity.
- affinity and/or other binding attributes of Fc regions for Fc receptors can be modulated through glycosylation or other modification.
- antibodies produced and/or utilized in accordance with the present disclosure include glycosylated Fc domains, including Fc domains with modified or engineered such glycosylation.
- any polypeptide or complex of polypeptides that includes sufficient immunoglobulin domain sequences as found in natural antibodies can be referred to and/or used as an “antibody”, whether such polypeptide is naturally produced (e.g., generated by an organism reacting to an antigen), or produced by recombinant engineering, chemical synthesis, or other artificial system or methodology.
- an antibody is polyclonal; in some embodiments, an antibody is monoclonal.
- an antibody has constant region sequences that are characteristic of mouse, rabbit, primate, or human antibodies.
- antibody sequence elements are humanized, primatized, chimeric, etc., as is known in the ail.
- an antibody utilized in accordance with the present disclosure is in a format selected from, but not limited to, intact IgA, IgG, IgE or IgM antibodies; bi- or multi- specific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fvs; polypeptide-Fc fusions; single domain antibodies, alternative scaffolds or antibody mimetics (e.g., anticalins, FN3 monobodies, DARPins, Affibodies, Affilins, Affimers, Affitins, Alphabodies
- SMIPsTM Small Modular ImmunoPharmaceuticals
- an antibody may lack a covalent modification (e.g., attachment of a glycan) that it would have if produced naturally.
- an antibody may contain a covalent modification (e.g., attachment of a glycan, a payload [e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.], or other pendant group [e.g., poly-ethylene glycol, etc.])
- Antigen refers to an agent that elicits an immune response; and/or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody.
- an antigen elicits a humoral response (e.g., including production of antigen- specific antibodies); in some embodiments, an antigen elicits a cellular response (e.g., involving T-cells whose receptors specifically interact with the antigen).
- an antigen binds to an antibody and may or may not induce a particular physiological response in an organism.
- an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer (in some embodiments other than a biologic polymer [e.g., other than a nucleic acid or amino acid polymer) etc.
- an antigen is or comprises a polypeptide.
- an antigen is or comprises a glycan.
- an antigen may he provided in isolated or pure form, or alternatively may be provided in crude form (c.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source).
- antigens utilized in accordance with the present invention are provided in a crude form.
- an antigen is a recombinant antigen.
- Arisen epitope As used herein, the term “arisen epitope(s)” is used to refer an epitope, e.g., a portion of an antigen, which a subject’s immune system has not encountered previously. For example, as circulating viral pathogens mutate, mutations in various epitopes found on viral proteins can result arisen epitopes that are versions of epitopes present on proteins of previously circulating variants, with one or more mutations introduced.
- composition may be used to refer to a discrete physical entity that comprises one or more specified components.
- a composition may be of any form - e.g., gas, gel, liquid, solid, etc.
- composition or method described herein as “comprising” one or more named elements or steps is open-ended, meaning that the named elements or steps are essential, but other elements or steps may be added within the scope of the composition or method.
- any composition or method described as “comprising” (or which "comprises") one or more named elements or steps also describes the corresponding, more limited composition or method “consisting essentially of (or which "consists essentially of) the same named elements or steps, meaning that the composition or method includes the named essential elements or steps and may also include additional elements or steps that do not materially affect the basic and novel characteristic(s) of the composition or method.
- any composition or method described herein as “comprising” or “consisting essentially of one or more named elements or steps also describes the corresponding, more limited, and closed-ended composition or method “consisting of (or “consists of) the named elements or steps to the exclusion of any other unnamed element or step.
- any composition or method disclosed herein known or disclosed equivalents of any named essential element or step may be substituted for that element or step. [0177] Determine'.
- the methodologies described herein include a step of “determining”. Those of ordinary skill in the art, reading the present specification, will appreciate that such “determining” can utilize or be accomplished through use of any of a variety of techniques available to those skilled in the art, including for example specific techniques explicitly referred to herein.
- determining involves manipulation of a physical sample. In some embodiments, determining involves consideration and/or manipulation of data or information, for example utilizing a computer or other processing unit adapted to perform a relevant analysis. In some embodiments, determining involves receiving relevant information and/or materials from a source. In some embodiments, determining involves comparing one or more features of a sample or entity to a comparable reference.
- Engineered Antigen As used herein, the term “engineered antigen,” is used to refer to an antigen that is or is intended to be artificially created and intentionally introduced to a subject, e.g., in order to generate an immune response (e.g., via vaccination), as opposed to a naturally occurring antigen, evolved through natural processes.
- engineered antigens are or comprise polypeptides.
- Engineered polypeptide antigens may be designed to match, be similar to, or based on other, reference antigens, which may themselves be engineered antigens or may be naturally occurring antigens.
- an engineered antigen is created by beginning with a structure of a reference antigen and introducing one or more amino acid modifications therein, e.g., to achieve a desired behavior / result upon planned introduction to a subject.
- engineered antigens may be designed in-silico - i.e., via computer- implemented systems and methods, using polypeptide models that represent various physical antigens and other computer representations.
- engineered antigens may be encoded by ribonucleic acid (RNA), which, in turn, may be used to manufacture an engineered antigen (e.g., in-vitro), or which may be directly administered to a subject, e.g., as an RNA vaccine.
- RNA ribonucleic acid
- Epitope As used herein, the term “epitope,” is used to refer to used herein, includes any moiety that is specifically recognized by an immunoglobulin (e.g., antibody or receptor) binding component.
- an epitope is comprised of a plurality of chemical atoms or groups on an antigen.
- such chemical atoms or groups are surface-exposed when the antigen adopts a relevant three-dimensional conformation.
- such chemical atoms or groups arc physically near to each other in space when the antigen adopts such a conformation.
- at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).
- Epitope Alteration Score As used interchangeably herein, the terms “epitope alteration score” and “epitope score” both refer to a measure of alteration to a viral polypeptide at epitope positions. In some embodiments, such alteration can be characterized by the impact of mutation(s) in one or more epitopes of a viral variant on recognition by antibodies (e.g., neutralizing antibodies). For example, in some embodiments, such alteration can be characterized by determining the number of antibodies potentially escaped. In some embodiments, antibodies for characterization have been isolated from patients who have been vaccinated against a disease or who have previously been infected with a disease (e.g., SARS- CoV-2).
- antibodies for characterization have previously been shown to bind a reference sequence.
- an epitope alteration score can be determined by comparison of mutations in a variant candidate to one or more regions of a reference sequence that have previously been shown to bind antibodies (e.g., through structural data).
- an epitope alteration score can be determined by enumerating the number of unique epitopes involving altered positions, as measured across one or more known antibody- viral polypeptide complex structures (e.g., all known antibody -viral polypeptide complex structures).
- an epitope alteration score is a measure of how many distinct epitopes are evaded by a variant candidate as compared to a reference sequence (e.g., as compared to a wild type sequence).
- an epitope alteration score is computed based on known binding sites of antibodies, e.g., as reported in Protein Data Bank.
- an epitope alteration score can change over time with identification of new epitope positions and/or discoveries of epitope-binding antibodies.
- an epitope alteration score can be used to characterize degree of alteration of a SARS-CoV-2 Spike polypeptide at epitope positions, for example, in some embodiments by counting the number or percentage of antibodies potentially escaped. In various embodiments described herein, an epitope alteration score can be normalized such that it ranks between 0 and 100%.
- Growth score refers to a measure of the rate at which a given variant is growing in a subject population (e.g., at a given time).
- a growth score refers to lineagelevel growth.
- a growth score of a given variant can be determined by referencing growth of a parent species or a known variant of substantially the same lineage, or a known variant having a similar sequence (e.g., a sequence that is at least 90% identical to the given variant).
- growth of a given variant is a function of the change in the number of subjects within a subject population who are reported as being infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a defined period of time). In some embodiments, growth of a given variant is a function of the change in the proportion of a subject population infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a defined period of time).
- a growth score of a given variant can be an empirically determined by considering sequences associated with a given variant (e.g., in some embodiments including sequences associated with a lineage) that have been observed within a defined period and computing its proportion among all observed sequences at a given time relative to a reference level (e.g., its proportion determined over a defined period of time).
- a growth score can be normalized such that it ranks between 0 and 100%.
- a human is an embryo, a fetus, an infant, a child, a teenager, an adult, or a senior citizen.
- Infectivity Score or “Fitness Prior Score” is a measure of a viral variant’s evolutionary fitness, and is a function of the efficiency with which a virus replicates and/or the efficiency with which a virus infects host cells.
- calculation of a fitness prior score comprises determining one or more of a log-likelihood score, a viral polypeptide receptor binding score, and/or a growth score.
- a fitness prior score is determined by referencing each of a log-likelihood score, a viral polypeptide receptor binding score, and a growth score.
- an appropriate reference measurement may be or comprise a measurement in a particular system (e.g., in a single individual) under otherwise comparable conditions absent presence of (e.g., prior to and/or after) a particular agent or treatment, or in presence of an appropriate comparable reference agent.
- an appropriate reference measurement may be or comprise a measurement in comparable system known or expected to respond in a particular way, in presence of the relevant agent or treatment.
- Log-likelihood refers to a measure of the existence probability of a variant polypeptide sequence, which has been determined using natural language processing algorithms.
- log-likelihood can be determined using a transformer model.
- log-likelihood can be determined without a reference sequence.
- log-likelihood is a transformer-derived log-likelihood without reference. The higher the log-likelihood of a variant, the more probable the variant is to occur from a language model perspective. In various embodiments described herein, loglikelihood can be normalized such that it ranks between 0 and 100%.
- a log-likelihood measures how log-likelihood of a variant polypeptide sequence compares to the entire population of known variants. In some embodiments, a log-likelihood measures how loglikelihood of a variant polypeptide sequence compares to other variants with similar mutational loads (“conditional log-likelihood”). Such conditional log-likelihood is particularly useful for assessing variants with high mutation counts (e.g., at least 30 or more, including, e.g., at least 40, at least 50, at least 60, at least 70, or more mutation counts).
- high mutation counts e.g., at least 30 or more, including, e.g., at least 40, at least 50, at least 60, at least 70, or more mutation counts.
- Machine learning module As used herein, the terms “machine learning module” and “machine learning model” arc used interchangeably and refer to a computer implemented process (e.g., a software function) that implements one or more particular machine learning algorithms, such as an artificial neural networks (ANN), random forest, decision trees, support vector machines, and the like, in order to determine, for a given input, one or more output values.
- machine learning models are deep learning models or deep neural networks - ANNs that comprise, in addition to an input layer and an output layer, one or more hidden layers (e.g., in between).
- Examples of deep learning models include, without limitation, recurrent neural networks (RNNs) (e.g., long short-term memory networks (LSTMs), bi-directional LSTMs (biLSTMs)), attention-based networks, such as transformer models, and convolutional neural networks (CNNs).
- RNNs recurrent neural networks
- LSTMs long short-term memory networks
- biLSTMs bi-directional LSTMs
- attention-based networks such as transformer models
- CNNs convolutional neural networks
- machine learning modules implementing machine learning techniques are trained in a supervised manner, for example using curated and/or manually annotated datasets.
- machine learning models may be trained in an unsupervised manner, using unlabeled data.
- a machine learning model may be trained via a reinforcement approach, for example wherein a reward / penalty system is used to train a machine learning model to learn strategies for accomplishing specified tasks.
- Training a machine learning model may be used to determine various parameters of a model, such as weights associated with layers in neural networks.
- a machine learning module is trained, e.g., to accomplish a specific task such as predicting types of hidden amino acids within of polypeptide sequences based on their context, values of determined parameters are fixed and the (e.g., unchanging, static) machine learning module is used to process new data (e.g., different from the training data), such as a new amino acid sequence, referred to as an inference.
- machine learning modules may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, for example to dynamically update the machine learning module.
- a trained machine learning module is a classification algorithm with adjustable and/or fixed (e.g., locked) parameters, e.g., a random forest classifier.
- two or more machine learning modules may be combined and implemented as a single module and/or a single software application.
- two or more machine learning modules may also be implemented separately, e.g., as separate software applications.
- a machine learning module may be software and/or hardware.
- a machine learning module may be implemented entirely as software, or certain functions of an ANN module may be carried out via specialized hardware (e.g. , via an application specific integrated circuit (ASIC), field programmable gate arrays (FPGAs), and the like).
- ASIC application specific integrated circuit
- FPGAs field programmable gate arrays
- Model As used herein, the term “model” is used to identify a computer representation of a particular physical object or quantity, e.g., that is accessed, displayed by, used as input to, generated as output of, etc., computer-implemented methods and systems and/or one or more steps and/or modules or functions thereof. For example, as described in further detail herein, various computer implemented systems and methods may operate on, process, and generate polypeptide models that represent physical polypeptides, such as particular proteins or portions thereof.
- polypeptides may be implemented in a variety of formats, such as a string of characters (e.g., letters, each representing a particular amino acid type) representing an amino acid sequence (e.g., a FASTA file), or a 3D structural model that includes information about a (e.g., relative) 3D location of amino acids and/or atoms thereof, such as a Protein Data Bank (PDB) file format which may be used to describe a 3D structure of a particular protein and includes, among other things, atomic coordinates of atoms of particular protein.
- PDB Protein Data Bank
- nucleic acid in its broadest sense, refers to any compound and/or substance that is or can be incorporated into an oligonucleotide chain.
- a nucleic acid is a compound and/or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester linkage.
- nucleic acid refers to an individual nucleic acid residue (e.g., a nucleotide and/or nucleoside); in some embodiments, "nucleic acid” refers to an oligonucleotide chain comprising individual nucleic acid residues.
- a "nucleic acid” is or comprises RNA; in some embodiments, a “nucleic acid” is or comprises DNA.
- a nucleic acid is, comprises, or consists of one or more natural nucleic acid residues.
- a nucleic acid is, comprises, or consists of one or more nucleic acid analogs.
- a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone.
- a nucleic acid is, comprises, or consists of one or more "peptide nucleic acids", which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, are considered within the scope of the present disclosure.
- a nucleic acid has one or more phosphorothioate and/or 5'-N-phosphoramidite linkages rather than phosphodicstcr bonds.
- a nucleic acid is, comprises, or consists of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxy guanosine, and deoxy cytidine).
- adenosine thymidine, guanosine, cytidine
- uridine deoxyadenosine
- deoxythymidine deoxy guanosine
- deoxy cytidine deoxy cytidine
- a nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2- aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5- methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5- bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5 -methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8- oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated
- a nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared with those in natural nucleic acids.
- a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or protein.
- a nucleic acid includes one or more introns.
- nucleic acids are prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), reproduction in a recombinant cell or system, and chemical synthesis.
- a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues long.
- a nucleic acid is partly or wholly single stranded; in some embodiments, a nucleic acid is partly or wholly double stranded.
- a nucleic acid has a nucleotide sequence comprising at least one element that encodes, or is the complement of a sequence that encodes, a polypeptide. In some embodiments, a nucleic acid has enzymatic activity.
- Pareto score refers to a measure of a variant’s performance with respect to / as evaluated via one or more scoring metrics, such as various scoring metrics and/or combinations thereof, described herein. These may include, but are not necessarily limited to, scores that evaluate fitness and ability to escape an immune response.
- a Pareto score comprises a combination of an immune escape score (e.g., as described herein) and a fitness prior score (e.g., as described herein).
- a Pareto score captures the relative evolutionary advantage of a given strain. In some embodiments, such a Pareto score can be determined as described in the Examples.
- a Pareto score is an optimality score, which, for example in some embodiments ranks a variant relative to other sequences, e.g., ones that are observed in a population.
- a high Pareto score at a given time for a specific lineage indicates that fewer variants have higher scores for fitness prior and immune escape at that time.
- a Pareto score in some embodiments is a ranking system, and fitness prior and immune escape scores incorporated therein can change as new data are acquired, the Pareto score for a given variant can change over time.
- Pareto optimality is defined over a set of lineages.
- lineages are Pareto optimal within a set if there are no lineages in the set with higher immune escape and higher fitness prior scores.
- a Pareto score is a measure of the degree of Pareto optimality. For example, in some embodiments, lineages with the highest Pareto score are Pareto optimal; and lineages with the second-best Pareto score would be Pareto optimal, if the Pareto optimal lineages were removed from the set, and so on.
- a patient refers to any organism to which a provided composition is or may be administered, e.g., for experimental, diagnostic, prophylactic, cosmetic, and/or therapeutic purposes. Typical patients include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and/or humans). In some embodiments, a patient is a human. In some embodiments, a patient is suffering from or susceptible to one or more disorders or conditions. In some embodiments, a patient displays one or more symptoms of a disorder or condition. In some embodiments, a patient has been diagnosed with one or more disorders or conditions. In some embodiments, the disorder or condition is or includes a viral infection (e.g., a SARS-CoV-2 infection). In some embodiments, the patient is receiving or has received certain therapy to diagnose and/or to treat a disease, disorder, or condition.
- animals e.g., mammals such as mice, rats, rabbits, non-human primates, and/or humans.
- a patient is a human.
- Peptide' refers to a polypeptide that is typically relatively short, for example having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
- Pharmaceutical composition ' refers to an active agent, formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, active agent is present in a unit dose amount that is appropriate for administration in a therapeutic regimen that shows a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population.
- compositions may be specially formulated for administration in solid or liquid form, including those adapted for the following: oral administration, for example, drenches (aqueous or non-aqueous solutions or suspensions), tablets, e.g., those targeted for buccal, sublingual, and systemic absorption, boluses, powders, granules, pastes for application to the tongue; parenteral administration, for example, by subcutaneous, intramuscular, intravenous or epidural injection as, for example, a sterile solution or suspension, or sustained-release formulation; topical application, for example, as a cream, ointment, or a controlled-release patch or spray applied to the skin, lungs, or oral cavity; intravaginally or intrarectally, for example, as a pessary, cream, or foam; sublingually; ocularly; transdermally; or nasally, pulmonary, and to other mucosal surfaces.
- oral administration for example, drenches (aqueous or non-aqueous solutions or suspension
- composition as disclosed herein means that the carrier, diluent, or excipient must be compatible with the other ingredients of the composition and not deleterious to the recipient thereof.
- composition or vehicle such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, involved in carrying or transporting the subject compound from one organ, or portion of the body, to another organ, or portion of the body.
- Each carrier must be “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient.
- materials which can serve as pharmaceutically-acceptable carriers include: sugars, such as lactose, glucose and sucrose; starches, such as corn starch and potato starch; cellulose, and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean oil; glycols, such as propylene glycol; polyols, such as glycerin, sorhitol, mannitol and polyethylene glycol; esters, such as ethyl oleate and ethyl laurate; agar; buffering agents, such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline;
- composition grade refers to standards for chemical and biological drug substances, drug products, dosage forms, compounded preparations, excipients, medical devices, and dietary supplements, established by a recognized national or regional pharmacopeia (e.g., The United States Pharmacopeia and The Formulary (USP-NF)).
- a recognized national or regional pharmacopeia e.g., The United States Pharmacopeia and The Formulary (USP-NF)
- Polypeptide' refers to a polymeric chain of amino acids.
- a polypeptide has an amino acid sequence that occurs in nature.
- a polypeptide has an amino acid sequence that does not occur in nature.
- a polypeptide has an amino acid sequence that is engineered in that it is designed and/or produced through action of the hand of man.
- a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both.
- a polypeptide may comprise or consist of only natural amino acids or only nonnatural amino acids.
- a polypeptide may comprise D-amino acids, L- amino acids, or both.
- a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and/or may comprise a cyclic portion.
- a polypeptide is not cyclic and/or does not comprise any cyclic portion.
- a polypeptide is linear.
- a polypeptide may be or comprise a stapled polypeptide.
- the term “polypeptide” may be appended to a name of a reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides.
- exemplary polypeptides within the class whose amino acid sequences and/or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family.
- a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and/or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class).
- a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and/or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%.
- a conserved region that may in some embodiments be or comprise a characteristic sequence element
- Such a conserved region usually encompasses at least 3-4 and often up to 20 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more contiguous amino acids.
- a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide.
- a useful polypeptide as may comprise or consist of a plurality of fragments, each of which is found in the same parent polypeptide in a different spatial arrangement relative to one another than is found in the polypeptide of interest (e.g., fragments that are directly linked in the parent may be spatially separated in the polypeptide of interest or vice versa, and/or fragments may be present in a different order in the polypeptide of interest than in the parent), so that the polypeptide of interest is a derivative of its parent polypeptide.
- Prevent or prevention refers to reducing the risk of developing the disease, disorder and/or condition and/or to delaying onset of one or more characteristics or symptoms of the disease, disorder or condition. Prevention may be considered complete when onset of a disease, disorder or condition has been delayed for a predefined period of time.
- Ribonucleotide encompasses unmodified ribonucleotides and modified ribonucleotides.
- unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U).
- Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications ( e.g.
- ribonucleotide also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.
- RNA Ribonucleic acid
- an RNA refers to a polymer of ribonucleotides.
- an RNA is single stranded.
- an RNA is double stranded.
- an RNA comprises both single and double stranded portions.
- an RNA can comprise a backbone structure as described in the definition of “Nucleic acid I Polynucleotide” above.
- An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments where an RNA is an mRNA.
- RNA typically comprises at its 3’ end a poly(A) region.
- an RNA typically comprises at its 5’ end an art-recognized cap structure, e.g., for recognizing and attachment of an mRNA to a ribosome to initiate translation.
- an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and/or by chemical synthesis methods).
- an RNA is a single-stranded RNA.
- a single- stranded RNA may comprise self-complementary elements and/or may establish a secondary and/or tertiary structure.
- encoding it can mean that it comprises a nucleic acid sequence that itself encodes or that it comprises a complement of the nucleic acid sequence that encodes.
- a single- stranded RNA can be a self-amplifying RNA (also known as selfreplicating RNA).
- Semantic Change refers to a measure of a functional change of a viral polypeptide of a variant (e.g., in some embodiments a viral polypeptide that interacts with a host cell receptor and/or is otherwise involved in host cell entry) with respect to at least one or a plurality of (e.g., at least two, at least three, at least four, or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species and/or known variants, e.g., of the same lineage) from the language model perspective.
- a semantic change is a measure of a functional change of a viral polypeptide of a variant (e.g., in some embodiments a viral polypeptide that interacts with a host cell receptor and/or otherwise involved in host cell entry) with respect to a plurality of (e.g., at least two or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species and/or known variants, e.g., of the same lineage) from the language model perspective.
- a relevant language model can comprise Transformer-derived embedding differences (e.g., as described herein) with respect to at least one or a plurality of (e.g., at least two, at least three, at least four, or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species or known variants, e.g., of the same lineage).
- a semantic change score can be computed using LI norm.
- a sematic change score can be computed using L2 norm (also known as Euclidean norm).
- semantic change describes how different a variant is with regard to an underlying statistical model (e.g., in some embodiments a large machine learning model fine-tuned on viral protein sequences observed until a given time point).
- semantic change score depends on sequences observed, and thus the semantic change score may change over time, as an underlying model is trained on new variant sequences and/or reference sequences.
- a semantic change score is determined for a variant Spike polypeptide from SARS-Co-V-2 as described herein.
- a semantic change score can be normalized such that it ranks between 0 and 100%.
- Subject refers an organism, typically a mammal (e.g., a human, in some embodiments including prenatal human forms).
- a subject is suffering from a relevant disease, disorder or condition.
- a subject is susceptible to a disease, disorder, or condition.
- a subject displays one or more symptoms or characteristics of a disease, disorder or condition.
- a subject does not display any symptom or characteristic of a disease, disorder, or condition.
- a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition.
- a subject is a patient.
- a subject is an individual to whom diagnosis and/or therapy is and/or has been administered.
- the term “substantially” refers to the qualitative condition of exhibiting total or near-total extent or degree of a characteristic or property of interest.
- One of ordinary skill in the biological arts will understand that biological and chemical phenomena rarely, if ever, go to completion and/or proceed to completeness or achieve or avoid an absolute result.
- the term “substantially” is therefore used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
- Variant refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In general, whether a particular molecule is properly considered to be a “variant” of a reference molecule is based on its degree of structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements.
- a variant by definition, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule.
- a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and/or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g., that are attached to the polypeptide or nucleic acid backbone).
- moieties e.g., carbohydrates, lipids, phosphate groups
- a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%.
- a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid.
- a reference polypeptide or nucleic acid has one or more biological activities.
- a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid.
- a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions.
- a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference.
- a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference.
- a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference.
- a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions or deletions as compared to the reference.
- a reference polypeptide or nucleic acid is one found in nature.
- Vaccination' refers to the administration of a composition intended to generate an immune response, for example to a disease (c.g., to a viral epitope).
- vaccination can be administered before, during, and/or after development of a disease.
- vaccination includes multiple administrations, appropriately spaced in time, of a vaccinating composition.
- Viral polypeptide receptor binding score refers to a measure of binding affinity between a viral polypeptide that plays a role in host recognition and/or host cell entry, and a corresponding host protein with which the viral polypeptide interacts to recognize and/or enter a host cell.
- a viral polypeptide receptor binding score is determined in silico.
- a viral polypeptide receptor binding score can be determined using a conformational sampling algorithm.
- a viral polypeptide receptor binding score can be determined using structures that have been optimized using a probabilistic optimization algorithm (for example, in some embodiments a variant of simulated annealing, aiming to overcome local energy barriers and follow a kinetically accessible path toward an attainable deep energy minimum with respect to a knowledge-based, protein-oriented potential).
- a viral polypeptide receptor binding score can be calculated using the change in solvent accessible surface area (SAS A) of a viral polypeptide in a complexed state (e.g., a bound state) and a non-complexed state (e.g., a non-bound state).
- SAS A solvent accessible surface area
- a viral polypeptide receptor binding score can be determined by calculating the change in energy of the complexed (e.g., bound) and non-complexed (e.g., non-bound) structures of a viral polypeptide and its cognate host receptor.
- change in binding energy can be estimated by differences in Gibbs free energy between bound and unbound states.
- a viral polypeptide receptor binding score can be normalized such that it ranks between 0 and 100%.
- a viral polypeptide receptor binding score can be calculated in silico, e.g., by calculating the change in Gibbs Free Energy, or the change in solvent accessible surface area in the bound and unbound states.
- a viral polypeptide receptor binding score can be calculated using in vitro binding data (e.g., using a dissociation constant, KD, or an association rate, kOn).
- in vitro binding data can be determined methods known in the art, including, e.g., but not limited to biolayer interferometry (BLI) and/or surface plasmon resonance (SPR).
- BLI biolayer interferometry
- SPR surface plasmon resonance
- an “ACE2 binding score” is a measure of binding affinity between an S protein of a coronavirus (e.g., SARS-CoV-2) or an immunogenic fragment of the S protein (e.g., the RBD domain) and the ACE2 protein.
- an ACE2 binding score can be calculated in silico, e.g., by calculating the change in Gibbs Free Energy, or the change in solvent accessible surface area in the bound and unbound states.
- an ACE2 binding score can be calculated using in vitro binding data (e.g., using a dissociation constant, KD, or an association rate, k on ). In some embodiments, such in vitro binding data can be determined methods known in the art, including, e.g., but not limited to biolayer interferometry (BLI) and/or surface plasmon resonance (SPR).
- BLI biolayer interferometry
- SPR surface plasmon resonance
- Wild-type As used herein, the term “wild-type” has its art-understood meaning that refers to an entity having a structure and/or activity as found in nature in a “normal” (as contrasted with mutant, diseased, altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). In some embodiments, in the context of SARS-CoV-2, “wild-type” refers to the Wuhan variant.
- Headers are provided for the convenience of the reader - the presence and/or placement of a header is not intended to limit the scope of the subject matter described herein.
- engineered antigens created via the technologies described herein are computer engineered versions of reference antigens, designed to interact with a subject’s immune system in particular, e.g., desired, ways.
- a reference antigen may be a naturally occurring protein, such as a variant of a particular viral protein, or portion thereof.
- Antigen engineering techniques described in further detail herein may be used to design a custom, tailored version of such reference antigens to encourage production of new antibodies and reduce likelihood of triggering and/or extent of a memory immune response that may result in generation of antibodies from memory B-cells that result prior exposure to a variant of reference antigen.
- the present disclosure exemplifies certain aspects of provided technologies via approaches for designing engineered versions of a recently evolved XBB SARS-Cov2 Spike protein variant. While certain examples provided herein are described with reference to particular proteins and viral variants, one of skill in the art, having read the present disclosure, will appreciate that approaches described herein may be applied and adopted, e.g., to other pathogens (e.g., virus types), proteins, sub-regions, and the like, to encourage production of desired immune response.
- pathogens e.g., virus types
- proteins e.g., sub-regions, and the like
- Immune imprinting is a phenomenon whereby initial exposure to a particular antigen can limit (e.g., subsequent) development of immune responses against epitopes that are unique to new variants of the antigen.
- immune systems respond, among other things, by generating antibodies that bind to and neutralize portions of antigen(s) of the agent, in a highly specific fashion. Subsequently, the immune system retains a ‘memory’ of the antigen(s), along with the ability to produce the particular antibodies that target it, in the form of memory B and T cells.
- RNA viruses include, without limitation, influenza, coronavirus (e.g., severe acute respiratory syndrome-related coronavirus), human immunodeficiency virus (HIV), Respiratory syncytial vims (RSV), and the like.
- XBB includes mutations across several epitopes. Many of these mutations as shared with other, earlier variants, while a subset are unique to XBB. When these shared, pre-existing, epitopes trigger memory immune responses, antibodies that are tailored to earlier variants to which an individual was exposed are produced in lieu of new antibodies that are expressly designed to neutralize XBB.
- systems and methods described herein include approaches for designing engineered antigens that reduce an extent to which a memory response is triggered by shared epitopes of a new antigen and encourage the immune system to generate novel responses to particular target epitopes that are unique to the new antigen.
- the present disclosure provides systems and methods for in-silico design of engineered antigens that are designed to encourage an immune response to particular epitopes of a reference antigen of an infectious agent.
- a reference antigen may be a protein (e.g., of or produced by) the infectious agent and/or a portion thereof.
- a reference antigen is a particular viral protein and/or portion thereof, such as a surface protein.
- a reference antigen is a protein or portion of a protein that has been determined to be utilized by the virus to infect host cells, for example implicated in binding to particular host cell proteins and/or facilitating fusion with a host cell membrane.
- a reference antigen is or comprises a particular portion - e.g., a sub-unit, domain, etc. of a viral protein.
- an infectious agent is or comprises a coronavirus, such as SARS-CoV 2 and a reference antigen thereof is a SARS-CoV 2 Spike protein, or a portion of the SARS-CoV 2 Spike protein.
- a reference antigen may be an entire Spike protein, or may be a particular portion, such as an /V-tcrminal region (e.g., an N- Terminal Domain (NTD)) or a Receptor Binding Domain (RBD).
- NTD N- Terminal Domain
- RBD Receptor Binding Domain
- a particular portion is selected to focus on a minimal relevant vaccine antigen, e.g., to facilitate removal of as many conserved epitopes as possible without, e.g., resorting to introducing point mutations (e.g., thereby limiting number of epitopes in which point mutations are to be introduced).
- antigen engineering technologies of the present disclosure aim to design engineered versions of a reference antigen that, when introduced (e.g., administered) to a subject, encourage their immune system to generate new antibodies that are expressly tailored to particular epitopes of the reference antigen.
- this involves reducing a likelihood and/or an extent to which an engineered version of a reference antigen will trigger a memory immune response, thereby ameliorating certain obstacles immune imprinting phenomena can, as described herein, cause in regards to e.g., vaccination.
- systems and methods described herein identify those portions of a particular reference antigen that remain are likely to trigger memory immune response, and generate, in-silico, one or more engineered variants in which these memory triggering sub-region(s) are disrupted.
- FIG. 1 shows an example process 100 for generating engineered antigens according to certain embodiments.
- example process 100 begins via accessing, generating, or otherwise obtaining, a computer representation of at least a portion of a reference antigen of interest, referred to herein as a polypeptide model 102.
- a variety of formats may be used for computer representation of reference antigens such as particular proteins and/or portions thereof.
- a polypeptide model 102 be or comprise an amino acid sequence, such as ordered string of letters, each representing a particular amino acid (c.g., FASTA a file).
- a polypeptide model may be or comprise a structural model, such as a 3D model that includes an indication of a location of each amino acid (or atom thereof) of a reference antigen in 3D space.
- protein data bank (PDB) format files may be used and/or accessed to obtain 3D structural information about particular proteins, for example based on derived crystallographic structures.
- a reference antigen may be a protein of a particular target viral valiant, such as a particular target SARS-CoV-2 variant's spike (S) protein or a portion thereof.
- S SARS-CoV-2 variant's spike
- a full-length SARS-CoV-2 S protein comprising a “Wild-Type” or
- “Wuhan” sequence has a sequence corresponding to that of the first detected SARS-CoV-2 strain, consisting of 1273 amino acids and having an amino acid sequence according to SEQ ID NO: 1
- position numberings in a SARS-CoV-2 S protein and/or portion thereof given herein are in relation to the amino acid sequence of SEQ ID NO: 1.
- One of skill in the art reading the present disclosure will understand and be able to determine corresponding positions in a SARS-CoV-2 S protein variant sequence from locations of positions provided relative to the amino acid sequence of SEQ ID NO: 1 (i.e. , a person of skill in the art provided positions relative to SEQ ID NO: 1, or another variant, will be able to determine corresponding positions in the S protein sequence of another SARS-CoV-2 variant or a fragment thereof).
- a portion of a SARS-CoV-2 S protein is described as having certain mutations, unless otherwise indicated, it should be understood that position numberings of the mutations are given, and identify locations, in relation to the amino acid sequence of SEQ ID NO: 1.
- a spike (S) protein described herein can be modified in such a way that the prototypical prefusion conformation is stabilized.
- Certain mutations that stabilize a prefusion confirmation are known in the art, e.g., as disclosed in WO 2021243122 A2 and Hsieh, Ching-Lin, et al. (“Structure-based design of prefusion- stabilized SARS-CoV-2 spikes,” Science 369.6510 (2020): 1501-1505), the contents of each which are incorporated by reference herein in their entirety.
- a SARS-CoV-2 S protein may be stabilized by introducing one or more proline mutations.
- a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to residues 986 and/or 987 of SEQ ”D NO: 1. In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at one or more positions corresponding to residues 817, 892, 899, and 942 of SEQ ID NO: 1 . In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to each of residues 817, 892, 899, and 942 of SEQ ID NO: 1. In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to each of residues 817, 892, 899, 942, 986, and 987 of SEQ ID NO: 1.
- stabilization of the prototypical prefusion conformation of a SARS-CoV-2 S protein may be obtained by introducing two consecutive proline substitutions at residues 986 and 987.
- spike (S) protein stabilized protein variants are obtained in a way that the amino acid residue at position 986 is exchanged to proline and the amino acid residue at position 987 is also exchanged to proline.
- a SARS-CoV-2 S protein variant wherein the prototypical prefusion conformation is stabilized comprises the amino acid sequence shown in SEQ ID NO: 2:
- CoV-2 S protein amino acid sequence e.g., as compared to SEQ ID NO: 1, are useful herein.
- B.1.1.7 (“Variant of Concern 202012/01” (VOC-202012/01))
- B.l .1 .7 (“alpha variant”) is a SARS-CoV-2 variant that was first detected in October 2020 in the United Kingdom from a sample taken the previous month, and quickly began to spread by mid-December. It is correlated with a significant increase in the rate of COVID-19 infection; this increase is thought to be at least partly due to a change of N501 Y inside the spike glycoprotein’s receptor-binding domain, which is needed for binding to ACE2 in human cells.
- B.l.1.7 is defined by 23 mutations: 13 non-synonymous mutations, 4 deletions, and 6 synonymous mutations (i.e., there are 17 mutations that change proteins and six that do not).
- Spike protein changes in B.l.1.7 include deletion 69-70, deletion 144, N501Y, A570D, D614G, P681H, T716I, S982A, and D1118H.
- the B.1.351 variant is defined by multiple spike protein changes including: L18F, D80A, D215G, deletion 242-244, R246I, K417N, E484K, N501Y, D614G and A701V. There are three mutations of particular interest in the spike region of the B.l.351 genome: K417N, E484K, N501Y.
- B .1.1.298 was discovered in North Jutland, Denmark, and is believed to have been spread from minks to humans via mink farms. Several different mutations in the spike protein of the virus have been confirmed. The specific mutations include deletion 69-70, Y453F, D614G, I692V, M1229I, and optionally S1147L.
- Lineage B.l.1.248 (the “gamma variant”), known as the Brazil(ian) variant, is one of the variants of SARS-CoV-2 which has been named P.l lineage.
- P.l has a number of S- protein modifications (L18F, T20N, P26S, D138Y, R190S, K417T, E484K, N501Y, D614G, H655Y, T1027I, VI 176F) and is similar in certain key RBD positions (K417, E484, N501) to variant B.1.351 from South Africa.
- Lineage B.1.427/B.1.429 (the “epsilon variant”), also known as CAL.20C, is defined by the following modifications in the S-protein: S 131, W152C, L452R, and D614G, of which the L452R modification is of particular concern.
- CDC has listed B.1 ,427/B.1 .429 as a “variant of concern”.
- B.1.525 ( “eta variant”) carries the same E484K modification as found in the P.l, and B.1.351 variants, and also carries the same AH69/AV70 deletion as found in B.1.1.7, and B.1.1.298. It also carries the modifications D614G, Q677H and F888L.
- B.1.526 ( “iota variant”) was detected as an emerging lineage of viral isolates in the New York region that shares mutations with previously reported variants. The most common sets of spike mutations in this lineage are L5F, T95I, D253G, E484K, D614G, and A701V.
- B .1.529 (“Omicron variant”) was first detected in South Africa in November 2021. Omicron multiplies around 70 times faster than Delta variants, and quickly became the dominant strain of SARS-CoV-2 worldwide. Since its initial detection, a number of Omicron sublineages have arisen. Listed below are the current Omicron variants of concern, along with certain characteristic mutations associated with the S protein of each.
- the S protein of BA.4 and BA.5 have the same set of characteristic mutations, which is why the below table has a single row for “BA.4 or BA.5”, and why the present disclosure refers to a “BA.4/5” S protein in some embodiments.
- the S proteins of the BA.4.6 and BF.7 Omicron variants have the same set of characteristic mutations, which is why the below table has a single row for “BA.4.6 or BF.7”).
- Table 1A Certain Omicron Variants of Concern and Their Characteristic Mutations.
- SARS-CoV-2 S proteins described herein comprise one or more mutations (including, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) characteristic of a certain Omicron variant (e.g., one or more mutations of an Omicron variant listed in Table 1A, e.g., each of the mutations associated with a given XBB variant in the above Table 1A).
- a particular immunogenic portion e.g., sub-region
- a full-length coronavirus S protein e.g., SARS-CoV-2 S protein
- a reference antigen for creation of engineered antigens as described herein.
- an immunogenic portion of a coronavirus e.g., SARS- CoV-2 S protein lacks certain features that are in the full-length polypeptide (e.g., features that have been shown or predicted to interfere with induction of a naive immune response).
- an immunogenic portion of a coronavirus e.g., SARS-CoV-2 S protein lacks regions that have (i) a low number or density of B cell neutralization epitopes and/or (ii) a high number or density of B cell epitopes not associated with neutralization.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a full S2 domain.
- a coronavirus (e.g., SARS-CoV-2) S protein lacking a full S2 domain lacks regions of S2 that have (i) a low number or density of B cell epitopes associated with neutralization or (ii) a high number of B cell epitopes not associated with neutralization, but retains other portions of S2.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a full S2 domain, but comprises certain sequences that can improve immunogenicity and/or stability of an immunogenic portion (e.g., in some embodiments, an immunogenic portion lacks a full S2 domain but retains a TM sequence).
- a person of skill in the art reading the present disclosure will be able to identify B cell epitopes in a coronavirus (e.g., SARS-CoV-2) S protein and determine which epitopes arc or are not associated with neutralization. For example, a person of skill in the ail will be aware of numerous studies that have identified such regions using antibody binding studies (e.g., studies characterizing antibodies produced in subjects infected with or vaccinated against SARS-CoV- 2).
- an immunogenic portion of a coronavirus (e.g., SARS- CoV-2) S protein comprises certain regions that have been determined to have a high number or density of neutralization epitopes and optionally a high mutation rate.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises an N-terminal domain (NTD) of the S protein.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) protein comprises a receptor binding domain (RED) of the S protein.
- an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises an SI domain of the S protein.
- an immunogenic portion of a coronavirus e.g., SARS-Coactivated coronavirus-
- CoV-2 S protein comprises an RED and an NTD and omits other features of the SI domain.
- Coronavirus e.g., SARS-CoV-2
- SARS-CoV-2 S proteins are well characterized, and a person of skill in the art will be able to determine which portions of an S protein sequence correspond to immunogenic portions discussed herein (e.g., which portions of an S protein sequence correspond to the NTD, the RED, the SI, and the S2 domains).
- an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises residues 327 to 528 of SEQ ID NO: 1 or a corresponding region.
- an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises the amino acid sequence: or a corresponding region.
- an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises the amino acid sequence: or a corresponding region.
- an SI domain of a coronavirus comprises amino acids 1 to 678 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant.
- an SI domain of a SARS-CoV-2 S protein comprises amino acids 1 to 683 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant.
- an SI domain of a SARS-CoV-2 S protein comprises amino acids 1 to 685 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant.
- an SI domain of a SARS-CoV-2 S protein comprises the amino acid sequence:
- an SI domain of a SARS-CoV-2 S protein comprises the amino acid sequence:
- an S2 domain of a SARS-CoV-2 S protein comprises amino acids 679 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant.
- an SI domain of a SARS-CoV-2 S protein comprises amino acids 684 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant.
- an SI domain of a SARS-CoV-2 S protein comprises amino acids 686 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant.
- compositions described herein deliver an immunogenic portion of an S protein of a SARS-CoV-2 variant.
- the variant is a variant of concern (e.g., a variant that has been predicted to and/or has been shown to spread rapidly in a relevant jurisdiction, e.g., as identified by certain public health agencies, e.g., the Center for Disease Control and Prevention (CDC), Public Health England and the CO VID- 19 Genomics UK Consortium for the UK, the Canadian COVID Genomics Network (CanCOGeN), and/or the
- a variant has been predicted to have a highly likelihood of becoming a variant of concern (e.g., using sequence-based algorithms that predict the ability of a variant to escape previously developed immune responses and/or measure the “fitness” of a given variant, such as described, e.g., in WO2022/235847 and WO2022/235853, the contents of each of which are incorporated by reference herein in their entirety).
- an RBD comprises mutations associated with a variant described herein. A person of skill in the art will be able to identify which portions of a given variant correspond to immunogenic portions described herein.
- a polypeptide comprises two or more SARS-CoV-2 subdomains (e.g., two or more SI domains or RBDs).
- a polypeptide comprises two or more receptor binding domains linked in tandem, e.g., as described in Dai, Lianpan, et al. “A universal design of betacoronavirus vaccines against COVID- 19, MERS, and SARS,” Cell 182.3 (2020): 722-733, and Han, Yuxuan, et al.
- the two or more subdomains are from the same SARS-CoV-2 variant (e.g., a variant described herein). In some embodiments, at least two of the two or more subdomains are from different SARS-CoV-2 variants (e.g., from different variants of concern, different Omicron variants, an Omicron variant and a non-Omicron variant, or a Wuhan strain and an Omicron variant).
- approaches described herein leverage information describing unique features of particular reference antigens, such as identifications of hallmark mutations.
- Hallmark mutations are those mutations that serve to distinguish particular reference antigens (e.g., proteins) from other, similar, antigens.
- a reference antigen is a particular protein or portion thereof of a particular viral variant
- hallmark mutations are those mutations that serve to distinguish the particular viral variant, as well as closely related variants, from other, for example variants.
- hallmark mutations may be identified based on one or more classification schemes, such as World Health Organization (WHO) classification(s), GISAID, Pango Lineage, Nextstrain clade, etc.
- WHO World Health Organization
- a reference antigen may be a SARS CoV 2 Spike protein of a particular XBB variant. Hallmark mutations may then be identified as those mutations that appear in greater than a particular hallmark threshold percentage of all sequences (c.g., within a particular database) classified as members of an XBB lineage, according to the WHO designations.
- a hallmark threshold is fifty-percent or greater.
- a hallmark threshold is sixty percent or greater.
- a hallmark threshold is seventy percent or greater.
- a hallmark threshold is seventy-five percent or greater.
- Exemplary XBB.1.5 hallmark mutations e.g., mutations relative to a Wuhan strain, e.g., identified via SEQ ID NO: 1 and/or SEQ ID NO. 2 identified as those mutations occurring in greater than 50% of sequences for variants classified as XBB.1.5 by the WHO are as shown in Table IB, below:
- one or more memory-triggering conserved region(s) 122 of polypeptide model 102 are identified 120.
- Identified (memory triggering) conserved region(s) 122 represent those particular portions of a reference antigen (that is represented by polypeptide model 102) that are determined to be likely to trigger a memory immune response.
- a memorytriggering conserved region may represent these particular portions of the reference antigen in a variety of manners.
- a memory -triggering conserved region may be or comprise a list of amino-acid positions, or, additionally or alternatively, be or comprise an identified 3D region on a surface of a 3D polypeptide model.
- memory triggering conserved regions may be identified and determined to be implicated in triggering of memory immune rcsponsc(s) in a variety of manners, using, for example and without limitation, a-priori known or determined data, such as lists of epitopes, structural models, in combination with experimental techniques (e.g., binding assays), and computational methods, individually or in combination.
- memory-triggering portions of a reference antigen determined by identifying a set of conserved epitopes, within the reference antigen, that are likely to trigger a memory immune response.
- Sets of shared epitopes of a reference antigen may be identified by comparing an initial set of known epitopes with mutations present in the reference antigen.
- a reference antigen is a version of a particular protein, e.g., a particular viral variant’s version of the protein.
- a set of known epitopes comprises epitopes that satisfy particular criteria, such as having been identified as targets of antibody binding (e.g., binding epitopes) and/or neutralizing antibodies (e.g., neutralizing epitopes).
- Known epitopes may be, for example, previously characterized epitopes of the particular virus of which the reference antigen is a variant, such as epitopes to which antibodies have been determined to bind.
- a known epitope is a site to which a neutralizing epitope binds (e.g., a neutralizing B-cell epitope).
- a set of known epitopes may include binding and/or neutralizing epitopes.
- Data identifying a set of known epitopes may be or comprise, for each known epitope of the set, a list of amino acid positions. Such data may be obtained from sources such as results of targeted experiments, proprietary datasets, and public information such as literature and publicly available datasets.
- data on epitopes of SARS-CoV 2 Spike protein epitopes can be found via CoV-AbDab, IEDB, etc.
- a known epitope data set may then be screened analyzed to determine which of the known epitopes are mutated in the reference antigen and which are un-mutated and, accordingly, conserved. For example, data identifying a set of known epitopes can be compared against a list of hallmark mutations of a particular reference antigen.
- Epitopes of the set of known epitopes that comprise / overlap with one or more hallmark mutations of reference antigen may be identified as mutated epitopes, and other epitopes of the set, e.g., which do not comprise a hallmark mutation, identified as conserved epitopes
- a set of known epitopes may be filtered using additional criteria, such as identifying and removing epitopes that are subsets of other known epitopes or known epitopes that are or are not located within particular regions such as, e.g., within the context of SARS-CoV2, an RBD and/or N- terminal region.
- identifying a memory-triggering conserved region comprises identifying a conserved surface of a polypeptide model representing a particular reference antigen.
- a conserved surface represents portions of a reference antigen that are available and/or have potential to bind with antibodies created from a memory immune response (e.g., memory B-cells), but, for example, may not necessarily have been previously identified as known epitopes.
- a conserved surface may represent those portions (e.g., amino acid sites) of the particular reference antigen that are (i) sufficiently surface accessible and (ii) determined to be sufficiently similar to other (e.g., related) antigens and/or unaffected by mutations present in the particular reference antigen such that they could be a binding target of antibodies associated with a memory immune response.
- amino acid positions that are determined to be located at a surface may be identified (e.g., using polypeptide model 102).
- a set of surface amino acids may be compared with hallmark mutations of the reference antigen to, for example, filter out (e.g., remove from the set) amino acid positions that are sites of hallmark mutations and/or within a particular distance [e.g., in 3D space, ‘as the crow flies;’ or traversed over a 3D surface (e.g., a geodesic distance)] of one or more hallmark mutations.
- the set of remaining amino acid sites can then be used to define a conserved surface.
- a conserved surface may be further refined via evaluation of a spatial distribution of amino acid sites, e.g., to exclude (from the conserved surface) portions that are determined too small for an antibody to bind to. For example, isolated patches below a particular threshold size (e.g., surface area) may be removed from the conserve surface.
- a particular threshold size e.g., surface area
- one or more memory triggering conserved regions are or comprise a set of conserved epitopes.
- one or more memory triggering sub-regions are or comprise a conserved surface.
- both a set of conserved epitopes and a conserved surface are included/identified as one or more memory triggering sub-regions.
- systems and methods of the present disclosure aim to disrupt these regions, e.g., in order to lessen or avoid (e.g., reduce a likelihood of and/or extent of) triggering a memory immune response.
- approaches described herein introduce (e.g., in-silico) amino acid modifications, such as substitutions, insertions, deletions, etc., into at least a portion of one or more identified memory triggering conserved regions of a reference antigen.
- approaches for introducing amino acid modifications into conserved regions aim to distribute amino acid modifications throughout conserved regions.
- one or more amino acid modifications are introduced into at least a portion of one or more conserved epitope regions.
- one or more amino acid modifications are introduced into each of one or more conserved epitope region(s), e.g., to disrupt each identified conserved epitope.
- one or more amino acid modifications are introduced into (e.g., across) a conserved surface.
- a plurality of amino acid modifications are generated.
- a quantity and location of amino acid modifications are introduced to create a particular desired spatial distribution of mutations.
- Design criteria may include, without limitation, a density across a 3D surface representing the conserved region, a minimum and/or maximal separation between amino acid modifications, spatially across a 3D conserved surface.
- desirable distribution of mutations across a conserved surface may be achieved by computing a position spread score that measures an extent to which locations of amino acid modifications are, e.g., evenly, distributed across a conserved surface and, e.g., evaluating designs using, at least in part, the computed position spread score.
- criteria such as density, minimum/maximum separation, position spread score, etc. is elevated using introduced amino acid modifications (e.g., to ensure introduced amino acid modifications are introduced in a manner to distribute them across a conserved region).
- criteria such as density, minimum/maximum separation, position spread score, etc.
- introduced amino acid modifications and existing, e.g., hallmark, mutations of a particular antigen such as a viral protein variant e.g., to ensure introduced amino acid modifications are introduced in a manner to distribute them across a conserved region and also to avoid mutating existing unique epitopes, e.g., which are desired to be retained, e.g., in order to promote production of antibodies directed to (e.g., having an affinity to) these regions).
- amino acid modifications introduced into memorytriggering sub-regions are selected from a set of allowed mutations.
- a set of allowed mutations comprises, for each of at least a portion of amino acid positions within a particular sequence, a set of one or more mutations having been identified as allowable.
- Allowable mutations may be identified using sequence data (e.g., nucleotide and/or amino acid sequences) of antigens that are related to the reference antigen.
- a reference antigen is a particular protein of a viral variant, such as a SARS-CoV 2 spike protein of a particular SARS-CoV 2 variant (e.g., an Omicron spike protein, an XBB spike protein, etc.)
- allowable mutations may be determined by first selecting a set of related lineage(s) and obtaining, for example from proprietary sequencing data and/or public data (e.g., GSAID) sequences of versions of the particular protein of for various viral variants classified as belonging to the selected set of related lineages.
- Mutations occurring within each of the related variants of the particular protein can be identified and included within a set of allowed mutations.
- mutations to include in the set of allowed mutations are filtered so as to include only those mutations observed at a sufficiently high rate or frequency, for example at or above a particular threshold frequency. For example, in certain embodiments, only those mutations observed within at least particular threshold percentage of the related variant sequences are included in set of allowed mutations.
- related variant sequences may be or comprise sequences of proteins within a similar family and/or that are known to perform a similar function to reference antigen.
- reference antigen may be a particular SARS-Cov2 protein, such as a Spike protein.
- reference sequences may be other coronavirus Spike proteins, or subsets thereof, such as Spike proteins of related SARS virus (e.g., those that are known to infect humans; e.g., SARS-Cov 1, MERS), or known to exist I originate from particular regions.
- amino acid modifications may, in certain embodiments, be selected based on and/or to utilize changes that are present in other related polypeptides.
- reference sequences may be or comprise sequences of variants of reference antigen, such as predecessors or members of other branches along an evolutionary tree. Reference sequences may be limited to variants having a particular level of prevalence and/or immune escape. In certain embodiments, reference sequences may be obtained from one or more public and/or proprietary databases, such as GSAID.
- one or more additional criteria are used to select amino acid modifications that provide a sufficient level of disruption and/or nonetheless preserve polypeptide chain stability.
- amino acid modifications that substitute an existing amino acid with certain comparable / similar amino acids are excluded, e.g., as not providing sufficient disruption.
- amino acid modifications that disrupt cysteine bridges arc disallowed and/or penalized.
- amino acid modifications that cause a charge flip and/or produce a large change in charge are disallowed and/or penalized.
- Such similarity I dis- similarity criteria and structure preservation criteria as described herein may, for example, be implemented via (e.g., encoded) rules, such as in conditional logic, lookup tables, etc.
- FIG. 2 shows an example process, whereby sequences of related antigens are accessed/obtained and analyzed to identify and generate a set of allowed mutations which can be utilized for generating amino acid modification into, and disrupting, conserved sub-regions of reference antigens.
- steps of selecting and inserting, and optionally, filtering, amino acid modifications may be repeated, e.g., iteratively, to generate, based on a single initial conserved region, multiple candidate engineered variants.
- resultant in-silico engineered variants of a reference antigen may be scored, e.g., to assess its viability, ability to infect host cells, relative similarity / dissimilarity to existing variants.
- resultant in-silico engineered variants of a reference antigen may be scored, e.g., to assess its viability, ability to infect host cells, relative similarity / dissimilarity to existing variants.
- one or more candidate disrupted polypeptide models 342a, 342b...342n are generated 340 from (e.g., by identifying 320 and disrupting 330 conserved region(s)) an initial polypeptide model 302 that represents a particular reference antigen.
- Each candidate disrupted polypeptide model corresponds to the initial polypeptide model, but with its conserved sub-region disrupted via introduction of one or more amino acid modifications.
- Candidate disrupted polypeptide models may evaluated and scored 340 in a variety of manners, for example via one or more approaches described in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the contents of each of which are incorporated herein by reference in their entirety.
- a candidate polypeptide model may be scored using a machine learning model, such as a language model.
- a language model is or has been trained to generate a predicted probability for each of one or more amino acids at each of one or more locations of an input sequence. Once trained, a language model may then be used to predict an overall likelihood, such as a log-likelihood, a conditional log likelihood, etc., of a particular variant represented via a particular input amino acid sequence.
- language models trained in this fashion may also be used to determine an embedding vector representation of an input amino acid sequence.
- Embedding vector representations of amino acid sequences are internal representations generated by a machine learning model and can be extracted as output from one or more hidden layers of a machine learning model. Extracted embedding layer representations of variants may be used themselves and/or used to generate a characteristic vector representing a particular variant.
- a polypeptide sequence is provided as input to a machine learning model which, in turn, generates an internal representation - i.e., an embedding - of the input polypeptide sequence, which may be represent an input sequence as a high dimensional matrix or vector (e.g., a numerical matrix or vector).
- machine learning models such as certain language models described in further detail herein, may receive an amino acid sequence as input and generate, an internal (e.g., embedding) representation.
- an embedding is or comprises a vector (e.g., zi), e.g., of length D, where D is a dimensionality of the vector, for each amino acid position in a sequence, such that an initial embedding corresponding directly to an internal representation output by a particular layer (e.g., a final embedding layer of a recurrent neural network, e.g., a feature map from a final transformer layer of a transformer-based model) may be a matrix of size n x D [e.g., or (n + 1) x D], where n is a length of an input amino acid sequence and D is a number of dimensions [e.g., in certain embodiments, an additional, class token may be used and appended to a network, such that a matrix is of a size (n + 1) x D], based on the particular machine learning model used.
- a vector e.g., zi
- D is a dimensionality of the vector, for each amino acid position in a sequence
- a characteristic vector may then be determined as a mean over at least a portion of sequence positions ⁇ e.g., all sequence positions, [e.g., but excluding a first (e.g., class token) position, not representing or corresponding to an amino acid in the sequence] ⁇ to, for example determine a characteristic vector z, for an amino acid sequence.
- a machine learning model may be used to generate, for each of a plurality of viral variant polypeptide sequences, a corresponding characteristic vector (e.g., based on an embedding of machine learning model). This approach can be used to associate each variant sequence with a location in a higher dimensional characteristic vector / embedding space.
- characteristic vectors and/or embeddings representing particular sequences may be compared with each other and/or with particular versions, such as a wild-type version of a protein and/or other particular variants of interest.
- a distance e.g., an LI distance, an L2 distance, etc.
- distances in embedding space may be referred to as a semantic change score, and are believed to be indicative, at least in part, of potential for a particular variant to be recognized by a host system previously exposed to one or more reference antigens.
- Machine learning models used to generate characteristic vectors may utilize and implement a variety of machine learning techniques.
- a machine learning model may be a deep learning model (e.g., an artificial neural network with one or more, e.g., a plurality of, hidden layers), such as a language or large language model (LLM).
- LLM language or large language model
- a machine learning model is or comprises one or more recurrent models, such long short-term memories (LSTMs), implemented alone or in combination, e.g., as in a bi-directional LSTM (bi-LSTM).
- a machine learning model may be or comprise one or more transformer models. Examples of machine learning models include, without limitation, evolutionary scale models (ESM), bidirectional encoder representations from transformers (BERT), and the like.
- EMM evolutionary scale models
- BERT bidirectional encoder representations from transformers
- a machine learning model may be trained, for example on protein sequence data available from public and/or private repositories, such as UniRef (see, e.g., Suzek et al. “UniRef: comprehensive and non-redundant UniProt reference clusters” 2007), the Virus Pathogen Resource (ViPR) database by the National Centre for Biotechnology Information, and the like. Training may be accomplished by providing a machine learning model with input sequences that are incomplete and/or partially masked and tasking them to predict the omitted and/or masked data. In this manner, machine learning models may be trained in an unsupervised fashion.
- UniRef see, e.g., Suzek et al. “UniRef: comprehensive and non-redundant UniProt reference clusters” 2007
- ViPR Virus Pathogen Resource
- structural modeling techniques may be used to score candidate polypeptide models.
- a viral polypeptide receptor binding score such as an ACE2 binding score
- an epitope alteration score may be determined for one or more candidate variants.
- a quantity of mutations may be evaluated and used as a score, e.g., to select candidate polypeptide models having particular ranges in terms of quantity of mutations.
- mutation co-occurrence scores may be used, to evaluate whether artificially introduced mutations in particular candidate variants are match cooccurrences observed in real-world, naturally evolved, variants.
- each candidate may be associated with and characterized by a set of performance scores 344a, 344b...344n.
- one or more candidate polypeptide models may be selected 360, e.g., on the basis of such performance scores, as a representation of an engineered antigen to be used, e.g., as an immunogenic composition.
- a set of two or more scores as described herein can be combined, for example as a (e.g., weighted) combined score.
- a set of two or more scores as described herein can be combined to identify sets of Pareto fronts and/or Pareto optimal solutions.
- a Pareto score based on a combination of two more scores, may be determined, for example as described in in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the contents of each of which arc incorporated herein by reference in their entirety.
- disrupted polypeptide models created in-silico via systems and methods described herein may be stored and/or provided for, e.g., display and/or further processing.
- a corresponding RNA sequence may be generated using the disrupted polypeptide model, the corresponding RNA sequence encoding the engineered antigen represented by the disrupted polypeptide model.
- an engineered antigen based on the disrupted polypeptide model may be manufactured and its biological activity assessed, e.g., in-vitro. In this manner, for example, a plurality of initial candidate engineered antigens may be created via in-silico design techniques described herein and subsequently manufactured and screened in-vitro.
- assays may be used to check for expression and folding.
- engineered variants designed via one or more approaches described herein may be produced and tested for viability in terms of expression and folding.
- one or more antibodies or binding agents may be used to analyze for intracellular and surface expression of an encoded antigen.
- an ACE2 binding agent may be used for example for testing engineered SARS-CoV 2 variants.
- a panel of reference (e.g., known) neutralizing antibodies may be used to evaluate an ability of an engineered construct to escape existing neutralizing antibodies.
- a reference panel of neutralizing antibodies comprises one or more antibodies that bind to distinct epitope classes (e.g., A, B, C, D, E, F).
- analysis may be extended to immune serum analysis from vaccinated/convalescent individuals.
- a pseudovirus assay is performed to check for loss of nAb titers.
- one or more (e.g., multiple) testing procedures are used and arranged in a decision tree I hierarchical fashion, for example as described in Example 8.
- all or a subset of the steps may be used, such that at each step / test, a construct is either retained, and passed along to a next step, or discarded, so as to progressively filter engineered construct designs until a subset that passes each step is obtained.
- Such steps may include any of the following:
- results of assays as described herein may be combined with in-silico design procedures, e.g., in an iterative fashion, for example to inform and/or add constraints to subsequent rounds of in-silico design.
- Engineered antigens of the present disclosure may comprise a full-length viral protein and/or a particular portion thereof, such as an RBD that has been altered, e.g., to disrupt conserved regions, according to various embodiments described herein (e.g., in section B, above, and/or various examples described below).
- Constructs comprising engineered antigens of the present disclosure may combine engineered viral proteins and/or portions thereof with one or more additional elements, such as, for example, one or more of the following: secretory signal(s), linkers, multimerization regions (e.g., fibritin domains), and membrane associated moiety(ies) (e.g., transmembrane regions).
- delivery of engineered antigens and/or constructs comprising them can be achieved by administration of such antigen(s) (e.g., of a polypeptide antigen).
- such delivery can be achieved by administration of a composition from which the antigen is subsequently generated (e.g., in and/or by the recipient).
- delivery of a polypeptide antigen is achieved through administration of a composition comprising a polynucleotide encoding the polypeptide antigen.
- such polynucleotide may be or comprise DNA; in some such embodiments, such polynucleotide may be or comprise RNA.
- composition comprising an RNA encoding an engineered antigen as described herein.
- provided pharmaceutical compositions e.g., immunogenic compositions, e.g., vaccines
- deliver antigens as described herein e.g., engineered antigens
- a nucleic acid construct e.g., in many embodiments, an RNA that encodes one or more engineered antigens as described herein and is expressed in the subject upon administration of the pharmaceutical composition (e.g., immunogenic composition, e.g., vaccine).
- the present disclosure encompasses the recognition that administration of nucleic acid, and particularly of RNA to achieve delivery (e.g., by expression) of encoded antigen can provide a variety of benefits relative to other strategies for immunizing against infections, such as a SARS-CoV-2 infection.
- RNA may be particularly useful and/or effective as an active agent in pharmaceutical compositions (e.g., immunogenic compositions, e.g., SARS-CoV-2 vaccines) for a variety of reasons including specifically that RNA can have intrinsic adjuvanticity.
- pharmaceutical compositions e.g., immunogenic compositions, e.g., SARS-CoV-2 vaccines
- RNA can have intrinsic adjuvanticity.
- ability to induce very high antibody titers to SARS-CoV-2 proteins e.g., particularly SARS-CoV-2 antigens associated with a variant of concern with high immune escape potential.
- RNA actives can also elicit significant and diverse T cell responses which, particularly when combined with strong antibody response, represents a combination of immune characteristics thought to potentially maximize the probability of protection.
- polyribonucleotides may be delivered for therapeutic applications described herein using any appropriate methods known in the art, including, e.g., delivery as naked RNAs, or delivery mediated by viral and/or non-viral vectors, polymer-based vectors, lipid compositions, nanoparticles (e.g., lipid nanoparticles, polymeric nanoparticles, lipidpolymer hybrid nanoparticles, etc.), and/or peptide-based vectors. See, e.g., Wadhwa et al.
- one or more polyribonucleotides can be formulated with lipid nanoparticles for delivery (e.g., administration).
- lipid nanoparticles can be designed to protect polyribonucleotides from extracellular RNases and/or engineered for systemic delivery of the RNA to target cells. In some embodiments, such lipid nanoparticles may be particularly useful to deliver polyribonucleotides when polyribonucleotides are intravenously or intramuscularly administered to a subject.
- systems and methods described herein may be used to design engineered antigens suitable for eliciting improved immune responses.
- approaches described herein may be used to create engineered antigens for use in vaccination of particular subjects, such as particular individuals and/or particular populations.
- technologies for antigen engineering described herein may be used to engineer antigens for delivery to subjects having been previously exposed to a particular variant (e.g., a prior naturally circulating variant) of the reference antigen.
- information pertaining to e.g., particular epitopes of the particular variant to which the subject was previously exposed may, in certain embodiments, be used to identify particular conserved epitopes to disrupt.
- Such approaches may be used to reduce memory response, and encourage naive immune respond, for subjects having been exposed previously via infection and/or prior vaccination.
- engineered antigens may be tailored (e.g., expressly created, or selected from a set of pre-existing options) for a subject whose memory B cells have been assessed.
- methods of the present disclosure include administering a composition that delivers an engineered antigen as described herein to a subject whose memory B cells have been assessed.
- approaches described herein may be used to create an engineered antigen modified memory epitope(s) target by the subject’s memory B cells (at least when non-neutralizing).
- such an engineered antigen may be administered to a subject.
- approaches for in-silico design of engineered antigens described herein may be applied toward design of immunogenic compositions.
- Such immunogenic compounds may have, for example, improved performance and/or be particularly tailored to individual subjects and/or populations, e.g., in view of particular populations of immune memory responses (e.g., depending on antigens to which various individuals and/or populations groups were first exposed).
- approaches described herein may be used to create improved immunogenic compositions.
- an infectious agent is a virus, a bacteria, or a eukaryotic cell (c.g., a plasmodium).
- an infectious agent is a bacterium.
- the bacterium is Mycobacterium.
- the bacterium is selected from Haemophilus influenzae, Chlamydophila pneumoniae, Mycoplasma pneumonia, Staphylococcus aureus, Moraxella catarrhalis, Legionella pneumophila, and Streptococcus pneumonia.
- the bacterium is Streptococcus pneumonia.
- an infectious agent is an RNA virus.
- compositions provided herein may provide a particular advantage in providing an immune response against RNA viruses, which have a relatively high mutation rate (high relative to other infectious agents).
- an infectious agent comprises a large number of strains, variants, or lineages. In some embodiments, an infectious agent has a relatively high mutation rate (e.g., relative to other infectious agents).
- an infectious agent is prone to immune escape.
- an infectious agent is one for which seasonal, variant- adapted booster shots are regularly provided.
- an infectious agent antigen is solvent exposed on the surface of the infectious agent. In some embodiments, an infectious agent antigen is a glyocoprotein. In some embodiments, an infectious agent antigen is involved in host cell recognition. In some embodiments, an infectious agent antigen is involved in host cell entry. In some embodiments, an infectious agent antigen comprises one or more B cell epitopes (e.g., one or more neutralization epitopes).
- This example describes conserved regions and hallmark mutations of XBB.1.5, and other (e.g., ACE2) regions of interest.
- the design approach in this example began by collecting 1,004 binding and neutralizing B-ccll epitopes from CoV-AbDab and IEDB. These epitopes were compared against XBB.1.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.1.5. This set of conserved epitopes was further curated via manual inspection, including removal of epitopes that were not located on an RBD surface and/or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions in the Spike protein RBD. Table 2A, below, lists these conserved epitopes, together with their respective sources.
- S SARS-CoV-2 spike
- S SARS-CoV-2 spike
- This Example describes design of an engineered antigen based on (e.g., using as a reference antigen) an XBB variant of a SARS-Cov2 Spike protein.
- Engineered antigens described in this example aim to lessen and/or avoid activation of a subject’s, such as an individual being vaccinated, memory immune response resulting from memory B and/or T cells, in order to encourage production of new neutralizing antibodies that target particular epitopes associated with an ACE2 binding interface and comprising XBB hallmark mutations.
- FIG. 4A an XBB RBD portion was used as a reference antigen for design of engineered antigens.
- the approach described in this example began with an XBB RBD sequence.
- a PDB structure including RBD positions 325 to 527 (PED ID 7EAM) was used as a polypeptide model.
- XBB hallmark mutations, as listed in Table IB, are shown in red in FIG. 4A, while non-mutated regions are shown in gray.
- 4B shows the XBB RBD model with an identification of an ACE2 interface shown in green, defined as in Lan ct al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and comprising the positions as listed in Example 1, above:
- FIGs. 4A and 4B also show an identified conserved surface in violet coloring.
- the conserved surface was identified as a continuous non-mutated surface.
- Amino acid sites identified as belonging to the conserved region in this example are listed below:
- FIG. 4C shows a colorized version of the XBB structural model. Coloration indicates which residues (amino acid sites) belong to a known neutralizing epitope that is associated with (e.g., targeted) by a neutralizing antibody.
- One-hundred and thirteen (113) known neutralizing epitopes were identified based on data in the RCSB Protein Databank (PBD) and the Immune Epitope Database and Analysis Resource (IEDB). Each residue was scored according to a number the 113 neutralizing epitopes in which it appears. The colorization in FIG.
- FIG. 4C ranges from yellow to green to blue to indicate a relative frequency of a position belonging to an epitope, with a highest count being 70, colored blue, intermediate counts colored green, and low counts colored yellow (i.e., sites that did nor, or infrequently, belonged to an epitope).
- FIG. 4C and this present example considers neutralizing epitopes in particular, but subsequent designs have also considered non-neutralizing epitopes.
- the present example aimed to generate engineered versions of XBB that would limit or avoid triggering memory immune responses and, rather, cause creation of new antibodies that were tailored to XBB hallmark mutations. Accordingly, the 113 neutralizing epitopes were compared with the XBB hallmark mutations to identify a subset of twenty-six (26) remaining epitopes that were not targeted by the XBB hallmark mutations. These conserved, non-mutated, epitopes arc listed in Table 2B, below.
- amino acid modifications were introduced at various positions. Amino acid modifications were generated by selecting from a set of allowable mutations that were known to occur in XBB, XBB sub-lineages, and Omicron. Additional criteria was used to restrict aminoacid modifications introduced to (i) require that modifications cause significant change - e.g., changes from Asp to Glu, Arg to Lys, Ile to Leu, Asn to Gin were not considered; and (ii) prohibit (e.g., exclude) modifications that would break a cysteine bond - for example, modifications to position C391 or C525 were disallowed.
- each candidate variant was evaluated against various design criteria. In particular, first, candidate designs were required to modify amino acids located within each of the 26 non-mutated epitopes, e.g., to maximize immune escape. Second, candidate designs were evaluated to ensure amino-acid modifications were sufficiently spread across the conserved surface, rather than being clustered together. Lastly, each design was scored using a combination of in-silico structural modeling and a machine learning-based language model.
- Structural modeling was used to compute an ACE2 binding score, described, for example, in PCT Publication WO 2022/235847 Al , entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022.
- FIGs. 4D and 4E show two design examples, Design 1 (FIG. 4D) and Design 2 (FIG. 4E). Both designs introduce amino acid modifications within each non-mutated epitope and also sufficiently spread modifications across the conserved surface. ACE2 binding scores, log-likelihood scores, and semantic change scores were also computed for each design. Design 1 (FIG. 4D) was identified as having a higher ACE2 binding score, while Design 2 (FIG. 4E) was identified as more immune escaping (c.g., based on a higher semantic change score).
- Cyan coloration in FIGs. 4D and 4E identifies positions where amino acid modifications were introduced in the conserved region for Designs 1 and 2, respectively.
- This Example describes another example design approach for creating an engineered antigen based on (e.g., using as a reference antigen) an XBB variant of a SARS-Cov2 Spike protein, in particular, XBB.1.5.
- engineered antigens described in this example aim to lessen and/or avoid activation of a subject’s, such as an individual being vaccinated, memory immune response, in order to novel B-cell immune responses, with XBB.1.5 as a starting point.
- FIG. 5A illustrates a SARS-CoV 2 virus and its components, including the Spike (S) protein, which is shown in greater detail, together with the ACE2 host receptor to which it binds, on the right-hand side of the figure.
- FIG. 5B shows a 3D structural representation of an XBB.1.5 RBD portion of a SARS-CoV-2 S protein, including positions 325 to 527 (PED ID 7EAM), which was used as a polypeptide model.
- XBB.1.5 hallmark mutations, listed below, are shown in red in FIG. 5B, while non-mutated regions are shown in gray.
- FIG. 5C shows a 3D structure of a full Spike protein, with XBB.1.5 hallmark mutations in red.
- the design approach in this example began by collecting 1,004 binding and neutralizing B-ccll epitopes from CoV-AbDab and IEDB. These epitopes were compared against XBB.l.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.l.5. This set of conserved epitopes was further curated via manual inspection, including removal of epitopes that were not located on an RBD surface and/or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions in the Spike protein RBD, which are listed in Table 1 above (in Example 1).
- FIG. 5D shows the XBB RBD model with an identification of an ACE2 interface shown in green, defined as in Lan et al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and comprising the positions listed below:
- FIG. 5E shows a colorized version of the XBB structural model. Coloration indicates which residues (amino acid sites) belong to a neutralizing epitopes that is associated with (e.g., targeted) by a neutralizing antibody. Each residue was scored according to a number of the 1,004 neutralizing and non-neutralizing epitopes, described above, in which it appears.
- the colorization in FIG. 5E ranges from yellow to green to blue to indicate a relative frequency of a position belonging to an epitope, with a highest count being 287, colored blue, intermediate counts colored green, and low counts colored yellow (i.e., sites that did nor, or infrequently, belonged to an epitope).
- amino acid modifications were introduced at positions across the conserved surface, which comprised 76 total positions, such that at least amino acid modification to be introduced into each conserved epitope.
- Amino acid modifications to be introduced were selected from a set of allowed mutations, using sequences of related valiants identified as belonging to XBB, BA, and their sub-lineages. Mutations present within these related lineages (i.e., XBB, BA, and their sublineages) at a frequency of greater than or equal to 1 % were included in the set of allowed mutations.
- Process 500 splits creation of candidate engineered variants via introduction of amino-acid modifications into two-steps.
- particular (amino acid) positions to modify were selected 510, to generate multiple sets of position combinations to mutate.
- each set of position combinations was generated by selecting positions until positions were selected within each of the 26 conserved epitopes. This process resulted in 200,000 sets of position combinations.
- FIG. 5G Process 500 splits creation of candidate engineered variants via introduction of amino-acid modifications into two-steps.
- particular (amino acid) positions to modify were selected 510, to generate multiple sets of position combinations to mutate.
- each set of position combinations was generated by selecting positions until positions were selected within each of the 26 conserved epitopes. This process resulted in 200,000 sets of position combinations.
- a filtering step 520 was also included in the position selection approach, whereby a position spread score was computed for each set of position combinations and the results of step 510 filtered based on the position spread score and number of mutations. Following filtering, 2,000 combinations remained, and were grouped into 20 clusters.
- FIGs. 6A-6F show 3D structures of several of the engineered antigen designs. Tables 4A-4C list, in each row, a particular design with its ID, mutations that were added, and various scores and selection criteria as described herein (identified with bold text in the first column of the tables).
- Tables 4A-4C list, in each row, a particular design with its ID, mutations that were added, and various scores and selection criteria as described herein (identified with bold text in the first column of the tables).
- Each of FIGs. 6A-6E shows XBB.1.5 hallmark mutations in red and positions where mutations were introduced in yellow, and, except for FIG. 6B, positions not considered for mutation in gray and the conserved surface in violet.
- FIGs. 6 A and 6B show design S4_l, which was selected to maximize ACE2 binding.
- FIG. 6C shows design Sl_3, which was selected to minimize a number of mutations.
- FIG. 6D shows design SI 4, which was selected based on surface distribution of mutations.
- FIG. 6E shows design S13_3, which was selected on basis of a high log-likelihood score, and
- FIG. 6F shows design SI 1, which provided a diverse set of mutations.
- Tables 4A-4C list each final engineered antigen design, indicating the mutations added to XBB.1.5, the total number of mutations, values of each of the four scores computed for the design, and selection criteria.
- Table 4B Vaccine designs with 10-12 mutations.
- Bolded text in the first column identifies designs shown in FIGs. 6A-6F.
- Bold text in the list of additional mutations (second column) highlights differences in designs mutating a same set of positions. Scores were ranked among all the designs that altered all 26 epitopes.
- Table 4C Vaccine designs with 13-15 mutations.
- Bolded text in the first column identifies designs shown in FIGs. 6A-6F.
- Bold text in the list of additional mutations (second column) highlights differences in designs mutating a same set of positions. Scores were ranked among all the designs that altered all 26 epitopes.
- FIG. 7 shows radar plots of the scores for each of the final engineered antigen designs, demonstrating a good diversity.
- a clustering algorithm based on one described in Cao et al. 2022 was used to cluster designs, according to the following steps:
- a mutation co-occurrence score was calculated by, for each pair of mutations (mi, mj), calculating, approximately the conditioned frequency that:
- the mutation cooccurrence score was calculated in this example as the averaged log frequency between all pairs of mutations.
- ACE2 binding score was computed in this example using the deep mutational scanning (DMS) data on ACE2 binding from Starr et al., 2022. This data includes DMS results for positions 331- 531 regarding the RBD:ACE2 binding of 8 variants: Alpha, Beta, Delta, Eta, Omicron BA.1, Omicron BA.2 and two versions of wild-type.
- the ACE2 binding score sums the “delta bind” (change of logio(KD)) over the mutations and variants to estimate the binding change for any RBD variant.
- Tables 5A and 5B below list the engineered antigen designs created via the approaches described in the present example.
- Table 5A lists RBD mutations according to the designs in this example.
- Table 5B provides RBD sequences for each design. As described herein, each sequence is an engineered version of a SARS-CoV 2 Spike protein RBD created using an XBB.1.5 variant RBD as a starting point.
- FIG. 5C shows metric values computed for the sequences in Tables 5 A and 5B, below.
- Table 6 below lists the (natural) XBB.1.5 RBD and the Wuhan RBD sequences. Three versions of the XBB.1.5. RBD sequence are shown in Table 6, allowing for minor variations in the (boundaries of the) particular portion of the spike protein that corresponds to the RBD region.
- Table 5B RBD Sequences for Engineered Antigen Designs.
- Table 5C RBD Mutations with Metrics for Designs in Tables 5A and B.
- FIGs. 9A-9B and 9D-9E show four manually designed RBD engineered antigens. Two constructs designed with a different spread of mutations focusing on distance from XBB.1.5 existing mutations are shown in FIGs. 9A-9B, with Table 7 A listing additional mutations added to XBB.1.5.
- FIG. 9C is a schematic of a SARS-CoV-2 trimer, adapted from Starr et al., SARS- CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021). https://doi.org/10.1038/s41586-021-03807-6. FIGs.
- 9D-9E shows two constructs designed boosting broadly conserved epitopes (class 4 and class 5 antibody sites) by disrupting epitopes identified as substantially conserved between SARS CoV 1 and SARS CoV 2 with mutations selected from SARS CoV 1 sequences.
- Table 7B lists additional mutations that were added to XBB.1.5.
- Tables 7C and 7D show various scoring metrics computed for each sequence.
- Table 7A Two manual designs focusing on spread of mutations.
- Table 7C Metrics for Designs in Table 7A.
- This example describes an embodiment of antigen engineering techniques described herein in which position spread scores that are used to encourage insertion of new amino acid modifications in a distributed fashion (e.g., as opposed to clustered) across a conserved surface account for, not only other introduced amino acid modifications, but also proximity to existing hallmark mutations (e.g., XBB hallmark mutations).
- the antigen engineering methods used in this example proceeded similarly to those in Example 3, above, but also included hallmark mutations in the position spread score (e.g., described in Example 3 above), which was used to evaluate candidate engineered antigen designs.
- this approach is believed to ensure that placing additional mutations directly adjacent to hallmark mutations is avoided.
- the embodiment described in this example maintains unique, e.g., Omicron, epitopes in their unaltered form and preferentially (e.g., only) mutates conserved epitopes.
- approaches where hallmark mutations are not expressly accounted for in this manner may present a certain likelihood of altering Omicron epitopes in a manner such that elicited antibodies bind with less affinity to real, desired, target epitopes.
- Tables 8 A and 8B below list the engineered antigen designs created via the approaches described in the present example (i.e., including XBB hallmark mutations in the position spread score).
- Table 8A lists RBD mutations according to the designs in this example.
- Table 8B provides RBD sequences for each design.
- Table 8C shows various scoring metrics computed for each sequence.
- Table 8A RBD Mutations for Engineered Antigen Designs.
- Table 8B RBD Sequences for Engineered Antigen Designs.
- This example describes certain approaches for mutation generation, scoring, and evaluation that may be used, additionally or alternatively, to various approaches described herein in certain embodiments.
- evolutionary algorithms may be used to determine mutations.
- each solution may be represented as a list of 26 mutations (in the case of SARS-CoV-2 RBD), corresponding to the 26 conserved epitopes described in examples above.
- a pool of solutions can be maintained and updated by swapping the mutations per epitope. Solutions can be scored using position spread, ACE2 binding, and machine learning-based (ML) scores, such as loglikelihood and semantic change.
- ML machine learning-based
- an epitope alteration score may be used.
- current versions of the epitope alteration score described, for example, in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the content of each of which is incorporated herein by reference in its entirety, consider mutations in any single position in an epitope sufficient to ‘evade’ the corresponding antibody.
- an, e.g., more stringent, version of an epitope alteration score may be used that considers position location and distance between a mutation and antibody CDR loops to provide increased granularity and accuracy.
- in-silico structural modelling may be used to evaluate complex kinetics.
- in-silico structural modelling may be used to evaluate complex kinetics.
- Example 7 Updated Designs Balancing Integrity Risk
- This example describes certain approaches for mutation generation, scoring, and evaluation that may be used, additionally or alternatively, to various approaches described herein in certain embodiments.
- this example provides sequences that were generated to balance risk should certain mutations be detrimental to protein integrity.
- certain in-silico sequence generation procedures may over-represent mutations at positions 384, 430, and 463 in a set of suggested designs.
- the designs presented in this example sought to balance risks (e.g., with respect to diversity of mutations across sets of designs) should certain in-silico suggested mutations be detrimental to RBD integrity.
- Table 9 A lists RBD mutations with respect to (e.g., to be added to) an XBB.1.5 RBD portion of a SARS-CoV-2 S protein (e.g., according to any one of the three XBB.1.5 RBDs provided in Table 6) according to the designs in this example.
- mutation positions identify positions in the context of (i.e., with reference to) a full length SARS-CoV-2 S protein.
- Table 9B shows scoring metrics computed for each sequence shown in Table 9A.
- Table 9A RBD Mutations for Engineered Antigen Designs.
- Table 9B RBD Mutations with Metrics for Designs in Table 9A.
- Example 8 Example In-Silico Design Testing Procedure
- This example describes an example experimental procedure for testing engineered synthetic variants designed via various approaches described herein.
- the procedure of the present example uses three backbones / constructs for evaluating engineered variants, as follows.
- a pcDNA.3.1-SARS-CoV-2-XBB.1.5-CA19 construct a mammalian expression plasmid encoding XBB.1.5 spike protein with a truncated cytoplasmic tail, harboring the respective mutated XBB.1.5 RBD.
- Use of this construct is anticipated to allow for (i) an expression check of the spike protein after transfection in HEK293T cells, (ii) production of pscudovirus, and (iii) analysis of immune escape parameters (c.g., complete escape with no detectable titer).
- a pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM construct a template for transcription of an mRNA encoding a BNT162b3-like transmembrane-anchored RBD-based vaccine antigen. Use of this construct is anticipated to allow for (i) an expression check of RBD in HEK293T cells, (ii) assessment of antibody-binding using a reference panel of different RBD epitope class binders and/or complex immune serum, and (iii) immunogenicity studies.
- a pcDNA3.4-SARS-CoV-2-XBB.1.5-RBD-his-avi construct a template for RBD protein production with a HIS tag (allowing for purification) and BAP/avi tag (allowing for assay development).
- This construct may be used for enzyme-linked immunosorbent assay (ELISA), biolayer interferometry/surface plasmon resonance (BLI/SPR ), and/or protein-protein interaction assays to evaluate, for example, whether known antibodies can bind to the created construct.
- ELISA enzyme-linked immunosorbent assay
- BI/SPR biolayer interferometry/surface plasmon resonance
- protein-protein interaction assays to evaluate, for example, whether known antibodies can bind to the created construct.
- Example procedures for testing designs may include one or more (e.g., up to all) of the below listed steps and/or steps shown in FIG. 10A, which arranges testing steps in a hierarchical fashion:
- Expression and folding check Expression and folding of various antigens (e.g., each designed antigen) will be evaluated using a flow cytometry based approach in step 1001.
- An example FACS staining protocol that can be used in connection with this approach is shown below.
- flow cytometry with hACE2- mFc as a primary binding agent can be used to analyze for intracellular and surface expression of encoded antigen.
- This approach may use a full length XBB S protein as a reference. Binding affinity of RBD proteins may also be evaluated, along with intracellular and/or extracellular surface expression. Results from this step may be compared with predictions from, e.g., machine learning algorithms used in in-silico design of engineered antigens and fed back in to (e.g., to refine) the in-silico design algorithms.
- Antibody escape - monoclonal and polyclonal Antibody escape will be evaluated using a flow-cytometry approach with a panel of selected reference antibodies binding to different epitope classes on the RBD (e.g., similar to the assay used for the expression and folding check, but with the reference antibodies used as binding agents instead of hACE2-mFc).
- a variety of reference antibodies may be used and/or selected based on the particular epitope class that they bind to, as well as, additionally or alternatively, a level of affinity to the particular SARS-CoV 2 variant (e.g., reference antigen) that is used as a starting point (i.e., which is mutated) for designing the engineered variants.
- a panel of reference antibodies may include multiple antibodies that bind to different epitope classes (e.g., A, B, C, D, E, and F).
- Antibodies that bind to the particular SARS-CoV 2 variant used as a reference antigen that is used to design engineered versions thereof may be included, along with antibodies that do not bind to that particular SARS-CoV 2 variant (but which bind to other variants).
- Cao et al. Nature, 2021 (doi.org/10.1038/s41586-021-04385-3) (in Supplementary Table 1) provides a listing of 247 neutralizing antibodies, from which a panel can be selected.
- Table 10A shows a subset of 14 antibodies from the table in Cao et al. that may be used as reference antibodies.
- FIG. 10C shows example binding assay data for several of the antibodies listed in Table 10A. Antibodies exhibiting XBB.1.5 binding are shown in Table 10B, along with IC50 values for wild-type and BA.l variants (* denotes data provided in Cao et al.).
- Table 10C shows the antibodies from Table 10A, with XBB.1.5 binding data, highlighting those antibodies that bind to XBB.1.5 in green text.
- certain classes of antibodies such as antibodies that bind to Class A-D epitopes, do not bind to BA.l S protein and, accordingly, are not expected to bind to XBB S protein.
- Certain antibodies that bind to class E and F will be checked for binding to XBB S protein. Abrogated/reduced binding of reference antibodies can be evaluated using titration EC50 values in step 1002.
- Table 10C Antibodies from Table 10A, with XBB.1.5 binding/assay data.
- panels of antibodies may include various other examples of antibody clones, for example, including (but not limited to) antibodies as listed below and/or similar antibodies: • WRAIR-2057 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants).
- WRAIR-2057 is described, for example, in further detail in Dussupt, V. et al., Low-dose in vivo protection and neutralization across SARS-CoV-2 valiants by monoclonal antibody combinations. Nat Immunol 22, 1503-1514 (2021). https://doi.org/10.1038/s41590-021-01068-z.
- COVOX-45 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants).
- COVOX-45 is described, for example, in further detail, in Dejnirattisai, W. et al., The antigenic anatomy of SARS-CoV-2 receptor binding domain, Cell, 184 (8), 2183-2200. e22, 2021, https://doi.Org/10.1016/j.cell.2021.02.032.
- S2H97 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants).
- S2H97 is described, for example, in further detail in Starr, T.N., et al. SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature 597, 97-102 (2021). https://doi.org/10.1038/s41586-021-03807-6.
- Antibody escape may, additionally or alternatively, be evaluated using immune scrum from vaccinatcd/convalcsccnt individuals. Abrogated/reduced binding of complex immune serum can be evaluated using titration EC 50 values and comparing against XBB1.5 values in step 1003.
- a third approach for assessing escape may employ a pseudovirus generation protocol and pVNT assay set-up to check for loss of nAb titers. This approach would aim to check for abrogated neutralization of respective pseudovirus in step 1004.
- a final step for example as shown in FIG. 10 A, may include a dedicated immunogenicity study, for example to evaluate whether scrum from immunized animals neutralizes XBB1.5, and/or one or more other variants.
- Vaccine compositions based on engineered SARS-CoV 2 antigens may be administered to vaccine naive animals (e.g., mice) (e.g., to evaluate immune response and its magnitude) and to vaccine experienced animals (e.g., mice), e.g., to evaluate ability to overcome immune imprinting.
- vaccine naive animals e.g., mice
- vaccine experienced animals e.g., mice
- FIG. 10B shows an example protocol, illustrating transfection and flow cytometry steps.
- pcDNA3.1-SARS-CoV-2-Swt-CA19 and/or pcDNA3.1-SARS- CoV-2-SXBB. 1 .5-CA19 may be transfected into HEK293T/17 cells and binding assays as described above can be used to evaluate hACE2 binding and (neutralizing) antibody binding, for example to carry out various steps in the decision tree shown in FIG. 10A.
- FIG. 10D also shows that XBB.l .5 Spike displays an approximately 6.7-fold higher apparent affinity towards hACE2 binding in comparison with Wild-Type Spike.
- triple-vax IM post-boost polyclonal serum pool can be tested in a dilution series ranging from 1:20 - 1:1.562.500.
- Table 11 shows several variant-of-concern (VOC) mutations, together with their numerical metrics, which may be used as references.
- VOC Variant of concern
- a first round of tests were carried out for engineered XBB.1.5 variants expressed in context of a full length SARS-CoV-2 spike (S) protein.
- S SARS-CoV-2 spike
- certain constructs described herein were incorporated into a pcDNA.3.1-SARS-CoV-2-XBB.1.5-CA19 construct harboring the respective mutated XBB.l .5 RBD.
- HEK293T/17 cells were seeded into flasks and incubated for 2 days at 37°C and 7.5% CO2. Constructs were transfected into the HEK293T/17 cells and incubated overnight at 37°C and 7.5% CO 2 .
- Flow cytometry was used to (i) evaluate expression levels via an anti-S2 fragment antibody, (ii) check for conserved ACE2 binding capacity (via hACE2 binding) and (iii) assess abrogation of RBD-targeted antibody binding, using the panel of five (5) monoclonal antibodies (mAbs) having demonstrated binding to XBB.1.5, shown in Table 10B.
- FIGs. 11A-11C binding tests were carried out for certain variant designs shown in Table 9A, expressed in the context of a full length spike (S) protein.
- S full length spike
- ACE-2 binding was evaluated using a hACE2-mFc antibody, according to the FACS protocol described in Example 8, above.
- Binding responses for each of the five mAbs listed in Table 10B were also evaluated via the protocol described in Example 8, above.
- FIGs. 1 1 A-l 1C show results for the S43 engineered XBB.1.5 variant design (FIG. 11B) and the S48 engineered XBB.1.5 variant design (FIG. 11C), along with a (parental / original) XBB.1.5 reference (FIG. 11 A).
- both the S43 and S48 variants were found to exhibit abrogated binding for all but 1-2 of the five mAbs, and the S43 variant maintained hACE-2 binding.
- FIGs. 12-14 variant designs shown in Table 9A were also expressed in the context of a BNT162b3-like (trimerized transmembrane-anchored RBD-based) vaccine antigen.
- HEK293T/17 cells were transfected with BNT162b3-XBB.1.5 and each respective variant’s RNAs.
- Flow cytometry was then used to assess ACE-2 binding via hACE2-mFc binding and binding of the five mAb antibody panel (listed in Table 10B) as before.
- Binding curves for polyclonal vaccine serum were also evaluated to assess whether impact on polyclonal serum dose-response was greater in comparison with hACE-2 dose-response for the engineered variants. Without wishing to be bound to any particular theory, it is believed that a stronger impact on polyclonal serum dose-response compared to hACE-2 dose-response when compared to the parental XBB.1.5 for a particular variant suggests successful ‘masking/mutation’ of conserved epitopes.
- FIGs. 12A-B show dose response curves for an XBB.1.5 reference (FIG. 12A), and the S43 variant design (FIG. 12B).
- design S43 expressed in the context of a trimerized TM-anchored RBD shows close to unaltered hACE-2 binding doseresponse when compared to XBB.1.5.
- binding of Class A Antibody 2, Class F Antibody 1 , and Class B Antibody 1 is abrogated, with only binding of Class F Antibody 2and Class E Antibody 2conscrvcd.
- FIGs. 13A and 13B show results repeating those of FIGs. 12A and B (FIG. 13A shown reference dose response curves and FIG. 13B showing engineered antigen design S43 dose-response curves) confirming the results shown in FIG. 12B - namely, conservation of hACE-2 binding, as well as Class F Antibody 2and Class E Antibody 2 mAb binding, but abrogation of Class A Antibody 2, Class F Antibody 1, and Class B Antibody Ibinding. Polyclonal serum binding was also compared with hACE2 binding as shown in FIGs. 13C and 13D.
- FIGs. 14A and 14B show another set of dose response curves for XBB.1.5 reference and for the S48 engineered antigen design from Table 9A expressed in the context of a trimerized TM-anchored RBD.
- S48 displays conserved hACE-2 and Class F Antibody 2binding, as did S43 (S48 harbors 5/7 mutations also found in S43); Class E Antibody 2shows minimal residual binding, whereas binding of Class A Antibody 2, Class F Antibody land Class B Antibody lis completely abrogated.
- FIGs. 14C and 14D provide additional data that suggest potential for successful disruption of conserved epitopes in S48.
- FIGs. 14C and 14D compare (i) shift in hACE2 binding curve (FIG. 14C) for the S48 variant relevant to the parental XBB.1.5 reference with (ii) the shift in binding curve for polyclonal serum (FIG. 14D).
- the shift in ACE2-binding (approximately 4-fold) is less than the shift observed in serum binding (> 10-fold) when comparing parental XBB.1.5 vs. S48, which hints towards progressive abrogation of binding antibody responses.
- Table 13 summarizes results of the screening data described in this example in a tabular format.
- the first two rows of the table show binding levels for XBB.1.5 and the ideal, desired engineered antigen.
- the columns representing binding data are arranged in two groups, corresponding to the different expression contexts analyzed - (i) full length XBB.1.5 spike (pcDNA.3. l -SARS-CoV-2-XBB.1 .5-CA19), (ii) trimerized transmembrane (TM) anchored RBD design (pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM).
- constructs S43 and S48 exhibit behavior close to that of the idealized construct.
- S43 shows conserved ACE-2 binding when expressed as a full-length S protein as well as in the trimerized TM-anchored RBD format, while abrogating binding to all but 1-2 of the mAb panel, depending on expression context.
- Construct S48 performed particularly well when expressed as trimerized TM anchored RBD, showing conserved ACE-2 binding and only binding with one of the five mAbs.
- RNA compositions encoding engineered antigens comprising SARS-CoV-2 S protein and/or portions thereof (e.g., an RBD domain) as described herein induce an immune response characterized by increased naive B cell activation and/or decreased memory B cell activation in vaccine-experienced subjects (mice in the present example).
- vaccine candidates comprising RNA encoding engineered antigens designed and screened according to the approaches described herein will be administered to mice previously exposed to full length SARS-CoV-2 S protein.
- vaccine candidates are tested on mice previously administered two doses of RNA encoding a full length SARS-CoV-2 S protein, with each engineered antigen vaccine candidate administered as a third dose (booster).
- FIG. 15A illustrates an example immunization study where engineered antigen designs S43 and S48 having sequences as described in Table 9A are evaluated in contexts of mRNA’s encoding (i) a full-length S protein, analogous to BNT162b2, but encoding the particular (e.g., S43 or S48) engineered XBB.1.5 variant and (ii) a trimerized TM anchored RBD domain.
- mice in each group will about 7-8.
- mice in each group are first administered two doses of a monovalent composition comprising RNA encoding a SARS-CoV-2 S protein of a Wuhan strain (BNT162b2).
- First and second doses of the monovalent vaccine are administered 21 days apart.
- groups of mice are divided into groups based on neutralization titers against the Wuhan strain (e.g., mice are allocated into groups so that average neutralization titers are approximately the same in each group).
- mice can be allocated into groups based on pseudo-virus neutralization titers (e.g., as shown in FIG. 15 A).
- Each candidate vaccine will then be administered as a third dose, 18 weeks, after administering a first dose of vaccine (i.e., day 126, as shown in FIG. 15A).
- a first dose of vaccine i.e., day 126, as shown in FIG. 15A.
- versions of the immunization study approach illustrated in FIG. 15A may be performed with the third, candidate vaccine, dose being administered on a shorter timeline following the first two doses.
- a third, candidate, dose may be administered quickly after a second dose, so long as sufficient time has been allowed for completion of immune reactions, e.g., generation of B-cells, in response to the initial, e.g., Wuhan, antigens.
- 28 days (4 weeks) or more may be sufficient, allowing, for example, for a first dose (of BNT162b2) to be administered on day 0, a second dose (of BNT162b2) administered on day 21, and vaccine candidates to be administered as third (e.g., booster) doses four weeks later, e.g., on day 49 (week 7) or later.
- candidates S43 and S48 may be administered as third doses, as a full spike (S) format - BNT162b2 (S43) and BNT162b2 (S48), as well as in a trimerized transmembrane-anchored RBD format - RBD-TM (S43) and RBD-TM (S48).
- S43 full spike
- S48 BNT162b2
- S48 trimerized transmembrane-anchored RBD format - RBD-TM
- a number of controls and/or reference vaccines may also be used.
- a third dose of BNT162b2 may be administered.
- full spike (S) and RBD-TM versions of the parental XBB.1.5 variant may be administered for comparison with the engineered valiants thereof.
- one test group may not be administered any third dose.
- Table 14A below, lists an example set of vaccine candidates (selected based on screening data described herein, e.g., in the previous example), several of which are also listed in FIG. 15A.
- Table 14B Several options for certain other candidate vaccines, which may be used additionally or alternatively, arc listed in Table 14B.
- Tables 14A and 14B provide description of vaccine candidates formats and references to exemplary sequences that are included in Tables 14C-14I.
- Other vaccine candidates comprising engineered antigens (e.g., variants of other VOCs) and reference compositions (vaccine doses) can be prepared and evaluated in a manner analogous to that described herein with respect to XBB.1.5.
- Table 14A Vaccine Candidates for Administering to Vaccine-Experienced Mice.
- Table 14B Additional Vaccine Candidates Options.
- Table 14C Sequences of an Exemplary RNA Construct Encoding a Soluble, Trimerized RBD (SP19-XBB.1.5_RBD- GS_Linker-Fibritin_long).
- Table 14D Sequences of an Exemplary RNA Construct Encoding a Full-Length, Prefusion- Stabilized SARS-CoV-2 S Protein (XBB.1.5_P2).
- Table 14E Sequences of an Exemplary RNA Construct Encoding a Soluble, Trimerized RBD (SP16-XBB.1.5_RBD- GS_Linker-Fibritin_long).
- Table 14F Sequences of an Exemplary RNA Construct Encoding a Soluble Trimerized SI Domain (XBB.1.5_Sl-GS_Linker- Fibritin long).
- Table 14G Sequences of an Exemplary RNA Construct Encoding a Membrane-Tethered Spike Protein Comprising a C- terminal Truncation (Spike_deltal9).
- Table 14H Sequences of an Exemplary RNA Construct Encoding a Trimerized, Membrane- Anchored RBD (SP19- XBB.1.5_RBD-GS_Linker-Fibritin_Short-GS_Linker-TM(deltal9)).
- Table 141 Sequences of an Exemplary RNA Construct Encoding a Membrane-Anchored SI Domain (XBB.1.5_S1- GS_Linker-Fibritin_Short-GS_Linker-TM(deltal9)).
- Blood samples will be collected immediately before, and 3 weeks, 5 weeks, 9 weeks, 13 weeks, 17 weeks, 18 weeks, 19 weeks, 22 weeks, 26 weeks, 30 weeks, 34 weeks, 35 weeks, 37 weeks, and 39 weeks after administering a first dose of RNA. Thirty-nine (39) weeks after administering a first dose of RNA, mice will be sacrificed, and final blood, lymph node, and spleen samples collected for analysis.
- Blood samples will be screened for titers of antibodies that bind and neutralize various SARS-CoV-2 strains and variants (e.g., using ELISA and pseudovirus assays that described herein). Each spleen sample can be analyzed individually. For lymph node samples, samples from two mice can be combined for analysis.
- B cells are isolated from spleens and lymph nodes will be phenotyped to determine B cell type (e.g., naive, memory, or plasma) and binding specificity (e.g., specificity to the S proteins of various SARS-CoV-2 strains and variants). Phenotyping can be performed using the FACS-based and depletion assays analogous to those depicted in FIG. 16A and 16B and described in further detail ineuer and Muik et al., Science Immunol. , 7(75) eabq2427 (2022) (doi/10.1126/sciimmunol.abq2427), the contents of which is incorporated by reference herein in its entirety. Pseudovirus neutralization titers will also be collected for each of the sera samples. BCR repertoire analysis will also be performed for B cells isolated from spleen samples.
- BCR repertoire analysis will also be performed for B cells isolated from spleen samples.
- samples from all mice can be grouped together to generate enough sample to perform the methods. Each sample can be screened for negative, Wuhan-specific, XBB.1.5 exclusive binders, and cross reaction between Wuhan and XBB.1.5 binding.
- the protocol described in the present example can be used to characterize the binding specificity of B cells in subjects administered a booster vaccine that delivers an antigen of a variant of concern - here, an XBB. 1.5 S protein or an immunogenic portion thereof.
- this experimental protocol can be used to determine the relative number of B cells that are specific to XBB.1.5 and/or what portions of the XBB.1.5 S protein those B cells recognize.
- Vaccine candidates that are less susceptible to immune imprinting can be characterized by one or more of: (i) an increased proportion of B cells that are specific to a variant of concern (namely, XBB.1.5 or an immunogenic portion thereof) delivered by the vaccine candidate, (ii) increased neutralization titers against a variant of concern encoded by the vaccine candidate, and/or (iii) an increased number of B cell receptors that recognize epitopes that are unique to an antigen encoded by the vaccine candidate (i.e., increased B cell breadth).
- Additional analysis techniques include spleen sample analysis, lymph node analysis, and blood sample characterization as shown in FIG. 17.
- the spleen sample analysis may involve a preparation of single cell suspension that results in >5xl0 7 leucocytes/spleen after isolation and red blood cells (RBC) depletion (with 80-90% cell viability). Out of these isolated cells, approximately 1.5xl0 7 leucocytes/spleen may be used for FACS phenotyping for immunogenicity testing.
- This analysis may involve, for example, naive, memory, and plasma cell specific for the different S proteins.
- spleen isolated cells approximately 1.5x107 leucocytes/spleen may be used for tagging and pooling with subsequent magnetic- activated cell sorting (MACS).
- Subsequent steps may involve FACS sorting and staining for memory and plasma cell population.
- Another subsequent step may include BCR repertoire analysis.
- Remaining leucocytes extracted from spleen may be used for enzyme-linked immunosorbent spot (ELISpot) assays and freezing of remaining samples or freezing and ELISpot.
- Lymph node analysis may involve a preparation of single cell suspension that results in approximately IxlO 6 leucocytes per inguinal (iLN) and pelvic (pLN) lymph nodes.
- This analysis may involve FACS phenotyping for immunogenicity testing.
- This analysis may involve, for example, naive, memory, and plasma cell specific for the different S proteins.
- blood sample characterization may involve pseudotyped virus neutralization tests (pVNTs) ELISA.
- vaccine candidates may be tested in vaccine naive mice.
- tests in vaccine naive mice may be carried out on shorter time scales and can be used to evaluate and/or confirm that vaccine candidates based on the various engineered antigen compounds can elicit meaningful immunogenic response to XBB.1.5 on their own.
- FIG. 15B shows an example protocol where vaccine candidates and references as described herein are administered to vaccine naive mice in a two-dose format, with the first dose administered on day zero and the second dose administered three weeks later, on day 21.
- Blood samples can be collected immediately before, and 2 weeks, 3 weeks, 4 weeks, and 7 weeks after administering a first dose of RNA. Seven (7) weeks after administering a first dose of RNA, mice will be sacrificed, and final blood, lymph node, and spleen samples collected for analysis. Neutralizing antibody generation can be tested via pVNT assays using XBB.1.5 pseudo-virus, while binding antibodies can be tested via ELISA, using XBB.1.5 RBD as a target.
- construct S48 corresponds to a SARS-CoV-2 S protein RBD with (i) XBB.1.5 hallmark mutations together with (ii) an additional set of introduced mutations that were engineered to disrupt memory triggering conserved regions of a baseline, reference, XBB.1.5 reference antigen. As shown in Table 9A, this additional set of introduced mutations in construct S48 included the following mutations:
- a subset was determined to be beneficial for expression of constructs comprising the S48 RBD design, while another subset was selected for further evaluation, e.g., as potentially sub-optimal and/or causing a reduction in expression.
- Mutations determined to be beneficial include, without limitation, K356T, P384S, T43O1, F464Y, and H519N.
- Mutations selected for further evaluation included L335F and L390R.
- FIG. 18 illustrates this additional set of tailored constructs.
- Construct design labels are shown along the vertical, such that each row corresponds to a particular construct design and a listing of possible mutations is shown along the horizontal.
- XBB.1.5 hallmark mutations are identified via a black asterisk (“*”) symbol, the mutations determined to be beneficial identified via a green star, the two mutations selected for further evaluation identified via a red downward pointing triangle, and new mutations identified via purple diamonds. Shading (dark green) identifies mutations present in a particular construct. As illustrated in FIG. 18, all constructs included the XBB.1.5 hallmark mutations as well as the five mutations determined to be beneficial.
- the BNT162b3 construct, shown in FIG. 21A is 1397 base pair mRNA encoding a membrane- anchored RBD and Fibritin domain (F) together with a viral signal peptide.
- the construct includes a secretory signal (“sec”), an RBD domain (“RBD”), a Fibritin domain (“F”) and a transmembrane anchor (“TM”).
- the RBD encoding domain was modified to encode the XBB.1.5 hallmark mutations and each construct design’s particular set of additional mutations.
- FIGs. 20A-C shows the binding affinity for XBB.1.5, S48, and seven construct designs shown in FIG. 18 to various SARS-CoV-2 monoclonal antibodies.
- FIG. 20A shows results for five antibodies for which the seven construct designs showed a full regain of binding
- FIG. 20B shows results for five antibodies for which the seven construct designs showed a partial regain of binding
- FIG. 20C shows results for two antibodies for which the seven constructs exhibited little to no binding. The results appear consistent with the polyclonal serum data, with S48 still showing most significant decrease in binding with the studied monoclonal antibodies.
- Table 15A RBD Mutations for Engineered Antigen Designs.
- Table 15B RBD Mutations for Engineered Antigen Designs.
- the cloud computing environment 2100 may include one or more resource providers 2102a, 2102b, 2102c (collectively, 2102). Each resource provider 2102 may include computing resources.
- computing resources may include any hardware and/or software used to process data.
- computing resources may include hardware and/or software capable of executing algorithms, computer programs, and/or computer applications.
- exemplary computing resources may include application servers and/or databases with storage and retrieval capabilities.
- Each resource provider 2102 may be connected to any other resource provider 2102 in the cloud computing environment 2100.
- the resource providers 2102 may be connected over a computer network 2108.
- Each resource provider 2102 may be connected to one or more computing device 2104a, 2104b, 2104c (collectively, 2104), over the computer network 2108.
- the cloud computing environment 2100 may include a resource manager 2106.
- the resource manager 2106 may be connected to the resource providers 2102 and the computing devices 2104 over the computer network 2108.
- the resource manager 2106 may facilitate the provision of computing resources by one or more resource providers 2102 to one or more computing devices 2104.
- the resource manager 2106 may receive a request for a computing resource from a particular computing device 2104.
- the resource manager 2106 may identify one or more resource providers 2102 capable of providing the computing resource requested by the computing device 2104.
- the resource manager 2106 may select a resource provider 2102 to provide the computing resource.
- the resource manager 2106 may facilitate a connection between the resource provider 2102 and a particular computing device 2104.
- the resource manager 2106 may establish a connection between a particular resource provider 2102 and a particular computing device 2104. In some implementations, the resource manager 2106 may redirect a particular computing device 2104 to a particular resource provider 2102 with the requested computing resource.
- FIG. 22 shows an example of a computing device 2200 and a mobile computing device 2250 that can be used to implement the techniques described in this disclosure.
- the computing device 2200 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers.
- the mobile computing device 2250 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices.
- the components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.
- the computing device 2200 includes a processor 2202, a memory 2204, a storage device 2206, a high-speed interface 2208 connecting to the memory 2204 and multiple high-speed expansion ports 2210, and a low-speed interface 2212 connecting to a low-speed expansion port 2214 and the storage device 2206.
- Each of the processor 2202, the memory 2204, the storage device 2206, the high-speed interface 2208, the high-speed expansion ports 2210, and the low-speed interface 2212 are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate.
- the processor 2202 can process instructions for execution within the computing device 2200, including instructions stored in the memory 2204 or on the storage device 2206 to display graphical information for a GUI on an external input/output device, such as a display 2216 coupled to the high-speed interface 2208.
- an external input/output device such as a display 2216 coupled to the high-speed interface 2208.
- multiple processors and/or multiple buses may be used, as appropriate, along with multiple memories and types of memory.
- multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
- a processor any number of processors (one or more) of any number of computing devices (one or more).
- a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (one or more) of any number of computing devices (one or more) (e.g., in a distributed computing system).
- the memory 2204 stores information within the computing device 2200.
- the memory 2204 is a volatile memory unit or units.
- the memory 2204 is a non-volatile memory unit or units.
- the memory 2204 may also be another form of computer-readable medium, such as a magnetic or optical disk.
- the storage device 2206 is capable of providing mass storage for the computing device 2200.
- the storage device 2206 may be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations.
- Instructions can be stored in an information carrier.
- the instructions when executed by one or more processing devices (for example, processor 2202), perform one or more methods, such as those described above.
- the instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 2204, the storage device 2206, or memory on the processor 2202).
- the high-speed interface 2208 manages bandwidth-intensive operations for the computing device 2200, while the low-speed interface 2212 manages lower bandwidthintensive operations. Such allocation of functions is an example only.
- the high-speed interface 2208 is coupled to the memory 2204, the display 2216 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 2210, which may accept various expansion cards (not shown).
- the low-speed interface 2212 is coupled to the storage device 2206 and the low-speed expansion port 2214.
- the low-speed expansion port 2214 which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input/output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
- input/output devices such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
- the computing device 2200 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 2220, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 2222. It may also be implemented as part of a rack server system 2224. Alternatively, components from the computing device 2200 may be combined with other components in a mobile device (not shown), such as a mobile computing device 2250. Each of such devices may contain one or more of the computing device 2200 and the mobile computing device 2250, and an entire system may be made up of multiple computing devices communicating with each other.
- the mobile computing device 2250 includes a processor 2252, a memory 2264, an input/output device such as a display 2254, a communication interface 2266, and a transceiver 2268, among other components.
- the mobile computing device 2250 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage.
- a storage device such as a micro-drive or other device, to provide additional storage.
- Each of the processor 2252, the memory 2264, the display 2254, the communication interface 2266, and the transceiver 2268, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.
- the processor 2252 can execute instructions within the mobile computing device 2250, including instructions stored in the memory 2264.
- the processor 2252 may be implemented as a chipset of chips that include separate and multiple analog and digital processors.
- the processor 2252 may provide, for example, for coordination of the other components of the mobile computing device 2250, such as control of user interfaces, applications run by the mobile computing device 2250, and wireless communication by the mobile computing device 2250.
- the processor 2252 may communicate with a user through a control interface 558 and a display interface 2256 coupled to the display 2254.
- the display 2254 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology.
- the display interface 2256 may comprise appropriate circuitry for driving the display 2254 to present graphical and other information to a user.
- the control interface 2258 may receive commands from a user and convert them for submission to the processor 2252.
- an external interface 2262 may provide communication with the processor 2252, so as to enable near area communication of the mobile computing device 2250 with other devices.
- the external interface 2262 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.
- the memory 2264 stores information within the mobile computing device 2250.
- the memory 2264 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units.
- An expansion memory 2274 may also be provided and connected to the mobile computing device 2250 through an expansion interface 2272, which may include, for example, a SIMM (Single In Line Memory Module) card interface.
- SIMM Single In Line Memory Module
- the expansion memory 2274 may provide extra storage space for the mobile computing device 2250, or may also store applications or other information for the mobile computing device 2250.
- the expansion memory 2274 may include instructions to carry out or supplement the processes described above, and may include secure information also.
- the expansion memory 2274 may be provided as a security module for the mobile computing device 2250, and may be programmed with instructions that permit secure use of the mobile computing device 2250.
- secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non- hackable manner.
- the memory may include, for example, flash memory and/or NVRAM memory (non-volatile random access memory), as discussed below.
- instructions are stored in an information carrier.
- the instructions when executed by one or more processing devices (for example, processor 2252), perform one or more methods, such as those described above.
- the instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 2264, the expansion memory 2274, or memory on the processor 2252).
- the instructions can be received in a propagated signal, for example, over the transceiver 2268 or the external interface 2262.
- the mobile computing device 2250 may communicate wirelessly through the communication interface 2266, which may include digital signal processing circuitry where necessary.
- the communication interface 2266 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others.
- GSM voice calls Global System for Mobile communications
- SMS Short Message Service
- EMS Enhanced Messaging Service
- MMS messaging Multimedia Messaging Service
- CDMA code division multiple access
- TDMA time division multiple access
- PDC Personal Digital Cellular
- WCDMA Wideband Code Division Multiple Access
- CDMA2000 Code Division Multiple Access
- GPRS General Packet Radio Service
- a GPS (Global Positioning System) receiver module 2270 may provide additional navigation- and location-related wireless data to the mobile computing device 2250, which may be used as appropriate by applications running on the mobile computing device 2250.
- the mobile computing device 2250 may also communicate audibly using an audio codec 2260, which may receive spoken information from a user and convert it to usable digital information.
- the audio codec 2260 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 2250.
- Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 2250.
- the mobile computing device 2250 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 2280. It may also be implemented as part of a smart-phone 2282, personal digital assistant, or other similar mobile device.
- Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and/or combinations thereof.
- ASICs application specific integrated circuits
- These various implementations can include implementation in one or more computer programs that are executable and/or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
- machine -readable medium and computer-readable medium refer to any computer program product, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine- readable medium that receives machine instructions as a machine-readable signal.
- machine-readable signal refers to any signal used to provide machine instructions and/or data to a programmable processor.
- the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer.
- a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor
- a keyboard and a pointing device e.g., a mouse or a trackball
- Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
- the systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components.
- the components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
- LAN local area network
- WAN wide area network
- the Internet the global information network
- the computing system can include clients and servers.
- a client and server are generally remote from each other and typically interact through a communication network.
- the relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
- modules described herein can be separated, combined or incorporated into single or combined modules. Modules depicted in the figures are not intended to limit the systems described herein to the software architectures shown therein.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Virology (AREA)
- General Health & Medical Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Molecular Biology (AREA)
- Medicinal Chemistry (AREA)
- Veterinary Medicine (AREA)
- Public Health (AREA)
- Pharmacology & Pharmacy (AREA)
- Animal Behavior & Ethology (AREA)
- Microbiology (AREA)
- Biophysics (AREA)
- Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
- Immunology (AREA)
- General Chemical & Material Sciences (AREA)
- Mycology (AREA)
- Epidemiology (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Oncology (AREA)
- Analytical Chemistry (AREA)
- Communicable Diseases (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Peptides Or Proteins (AREA)
- Medicines Containing Antibodies Or Antigens For Use As Internal Diagnostic Agents (AREA)
Abstract
Presented herein are technologies directed to in-silico design of custom, engineered, antigens. In particular, in certain embodiments, methods and system of the present disclosure provide for engineering antigens to reduce their activation of a memory immune response, such as B cell and/or T cell based response, when introduced into a subject. Designing antigens in this manner, may, for example, improve their performance as immunogenic compositions for purposes of vaccination. For example, without wishing to be bound to any particular theory, it is believed that, among other things, reducing the extent to which a memory immune response is triggered can lead to improved production of new antibodies that are selectively tailored by the subject's immune system to neutralize particular (e.g., arisen) epitopes of a reference antigen.
Description
SYSTEMS AND METHODS FOR ENGINEERING SYNTHETIC ANTIGENS TO PROMOTE TAILORED IMMUNE RESPONSES
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority and benefit of U.S. provisional application numbers 63/448,217 (filed February 24, 2023), 63/448,215 (filed February 24, 2023), 63/448,987 (filed February 28, 2023), 63/449,031 (filed February 28, 2023), 63/449,936 (filed March 3, 2023), 63/452,989 (filed March 17, 2023), 63/452,987 (filed March 17, 2023), 63/514,242 (filed July 18, 2023), and 63/514221 (filed July 18, 2023), the contents of each of which are incorporated by reference herein in their entirety.
BACKGROUND
[0002] Vaccination can play a pivotal role in managing and ensuring public health. When vaccinated with sufficiently effective immunogenic compositions against a particular infectious agent, individuals experience reduced risk of infection and/or reduced severity of disease in the event of an in infection. Accordingly, development of highly effective vaccination techniques that can keep pace with ever evolving valiants of circulating pathogens and newly emergent diseases is a critical challenge.
SUMMARY
[0003] Presented herein are technologies directed to in-silico design of custom, engineered, antigens. In particular, in certain embodiments, methods and system of the present disclosure provide for engineering antigens to reduce their activation of a memory immune response, such as B cell and/or T cell based response, when introduced into a subject. Designing antigens in this manner, may, for example, improve their performance as immunogenic compositions for purposes of vaccination. For example, without wishing to be bound to any particular theory, it is believed that, among other things, reducing the extent to which a memory immune response is triggered can lead to improved production of new antibodies that are
selectively tailored by the subject’s immune system to neutralize particular (e.g., arisen) epitopes of a reference antigen.
[0004] For example, in certain embodiments, systems and methods of the present disclosure identify, within computer representations of a reference antigen, conserved regions that are similar to regions of other, for example previously circulating, variants of the reference antigen and thus likely to trigger a memory response. Approaches described herein then disrupt these conserved region(s), for example by introducing amino acid modifications within I across them. In this manner, an engineered antigen can be generated that retains certain portions of the reference antigen, such particular target epitopes, but replaces a conserved region with disrupted versions thereof. Without wishing to be bound to any particular theory, it is believed that engineered antigens with disrupted conserved region(s) when manufactured and introduced into a subject, arc less likely to trigger a memory immune response, e.g., from memory B or T cells, and, instead, encourage naive responses and, accordingly, production of new neutralizing antibodies that are expressly tailored to the retained target epitopes. Such engineered antigens may, accordingly, offer improved efficacy when used as immunogenic compositions, particularly for viral infectious agents that are prone to mutation.
[0005] In one aspect, the present disclosure provides methods for in-silico design of an engineered antigen [e.g., for (e.g., characterized in that) eliciting an immune response directed to one or more target epitopes of a reference antigen of an infectious agent while reducing the engineered antigen’s activation of a (e.g., B cell and/or T cell) memory immune response (e.g., relative to) to the reference antigen], the method comprising: (a) receiving and/or accessing, by a processor of a computing device, a polypeptide model representing (e.g., as a sequence of and/or 3D structural model of) a reference antigen of an infectious agent; (b) identifying, by the processor, within the polypeptide model, one or more memory-triggering conserved region(s) representing conserved portions of the reference antigen that are determined likely to trigger a memory immune response [e.g., portions of the reference antigen that are determined to (i) correspond to known epitopes and/or potential epitopes and/or (ii) do not comprise hallmark mutations of the reference antigen]; (c) generating, by the processor, one or more amino-acid modifications within at least a portion of the one or more (memory -triggering) conserved region(s) [e.g., the one or more amino-acid modifications being one or more point modifications,
insertions, and/or deletions], thereby creating a disrupted polypeptide model representing the engineered antigen [c.g., representing an artificially engineered version of the reference antigen with mutations introduced to disrupt its conserved portions (e.g., at least a portion of the one or more memory-triggering conserved regions, e.g., that are/were determined likely to trigger memory immune response)]; and (d) storing and/or providing, by the processor, the disrupted polypeptide model for display and/or further processing.
[0006] In some embodiments, a reference antigen is or comprises at least a portion of a naturally occurring variant of a viral protein [e.g., wherein the infectious agent is a variant of a particular virus (e.g., an influenza virus, a coronavirus, a respiratory syncytial virus, a filovirus) and wherein the reference antigen is or comprises at least a portion of a protein thereof],
[0007] In some embodiments, a reference antigen (e.g., and/or viral protein) is or comprises at least a portion of a SARS-Cov2 Spike polypeptide [e.g., a Receptor Binding Domain (RBD); e.g., an N-Terminal region; e.g., substantially all of a (e.g., an entire) Spike protein] [e.g., a portion selected to focus on a minimal relevant vaccine antigen, e.g., to facilitate removal of as many conserved epitopes as possible without e.g., resorting to introducing point mutations (e.g., thereby limiting number of epitopes in which point mutations are to be introduced)] .
[0008] In some embodiments, a reference antigen is or comprises at least a portion of a particular SARS-CoV-2 variant Spike polypeptide.
[0009] In some embodiments, a particular SARS-CoV-2 variant is a member of an Omicron and/or XBB lineage classification (e.g., according to a WHO, Pango, Nextstrain, etc. classification) (e.g., wherein the particular SARS-CoV-2 variant is XBB.1.5; e.g., wherein the particular SARS-CoV-2 variant is JN.l).
[0010] In some embodiments, an infectious agent is or comprises (e.g., a particular variant of) an RNA virus (e.g., a virus that encodes its genetic information with RNA) and the reference antigen is or comprises at least a portion of a protein thereof.
[0011] In some embodiments, a reference antigen is or comprises a bacterial protein (e.g., wherein the infectious agent is a bacteria and the reference antigen is or comprises at least a portion of a particular protein thereof).
[0012] In some embodiments, a reference antigen is or comprises an antigen (e.g., a surface antigen; e.g., a protein) of a parasite (e.g., wherein the infectious agent is a parasite and the reference antigen is or comprises at least a portion of a particular protein thereof) (e.g., wherein the infectious agent is a malaria parasite and the reference antigen is an antigen thereof).
[0013] In some embodiments, one or more memory -triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar’ to (i) one or more (e.g., pre-existing) variants thereof and/or (ii) an initial/wild-type strain (e.g., a first-observed strain) {e.g., wherein the one or more memory-triggering conserved region(s) represent portion(s) of the reference antigen having sufficient sequence similarity [e.g., at least 80% (including, e.g., at least 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher] identical; e.g., identical] to corresponding portions of (i) one or more (e.g., pre-existing) variants thereof and/or (ii) an initial/wild-type strain (e.g., a first-observed strain)}.
[0014] In some embodiments, a reference antigen is a particular target SARS-CoV-2 variant (e.g., XBB.1.5; e.g., JN.l) S polypeptide or portion thereof (e.g., wherein the reference antigen is an RBD of the target SARS-CoV-2S protein) and wherein the one or more memorytriggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar to corresponding portion(s) of (i) one or more (e.g., pre-existing) other SARS-CoV-2 variant polypeptides and/or (ii) Wuhan SARS-CoV-2 polypeptide [e.g., wherein the one or more memory-triggering conserved region(s) represent unmutated portion(s) of the reference antigen that are common to the reference antigen and corresponding portions of (i) the one or more (e.g., pre-existing) SARS-CoV-2 variant polypeptide) s) and/or (ii) Wuhan SARS-CoV-2 polypeptide].
[0015] In some embodiments, one or more memory -triggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and unmutated on the reference antigen (e.g., known epitopes without any hallmark mutations) [e.g., wherein the reference antigen is a particular sub-region (e.g., an RBD) of a target SARS-CoV-2 variant S protein (e.g., XBB.1.5; e.g., JN.l) and wherein the set of conserved epitope regions represent known epitopes that are present and un-muted on the particular sub-region of the target SARS-CoV-2 variant S protein relative corresponding sub-regions of SARS-CoV-2 S proteins of to one or more pre-existing variants and/or Wuhan strain].
[0016] In some embodiments, step (b) comprises: obtaining, by the processor, data corresponding to a set of known epitopes and identifying, within the reference antigen, each of one or more particular known epitopes of the set; obtaining, by the processor, an identification of a set of hallmark mutations of the reference antigen; and identifying, by the processor, as the set of conserved epitope regions, those particular known epitopes that correspond to portions of the reference antigen without any hallmark mutations.
[0017] In some embodiments, a set of known epitopes comprise one or more of the epitopes listed in Table 2A.
[0018] In some embodiments, a set of known epitopes comprise one or more of the epitopes listed in Table 2B.
[0019] In some embodiments, one or more memory -triggering conserved region(s) are or comprise a conserved surface representing a (e.g., continuous; e.g., contiguous) un-mutated (e.g., lacking any hallmark mutations) surface of the reference antigen [e.g., wherein the reference antigen is or comprises an XBB.1.5 variant of an RBD of a SARS-Cov-2 S protein and the conserved surface is or comprises at least a portion (e.g., up to all) of the following positions: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y38O, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, 1468, S469, T470, E471, 1472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527].
[0020] In some embodiments, provided methods comprise identifying, and/or accessing an identification of (e.g., positions of), a set of one or more hallmark mutations of the reference antigen [e.g., wherein the reference antigen is a protein of a particular viral variant classification and the set of hallmark mutations comprises those mutations of the reference antigen that are present at or above a particular threshold rate (e.g., appearing in sequences identified as belonging to the viral variant classification at or above the threshold rate (e.g., a fraction, percentage, etc.))].
[0021] In some embodiments, a reference antigen is or comprises a SARS-Cov2 XBB.1.5 Spike protein {e.g., a SARS-CoV-2 S protein with XBB.1.5 hallmark mutations [e.g.,
comprising at least a portion (e.g., all) of the following mutations: T191, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0022] In some embodiments, step (c) comprises selecting, by the processor the one or more amino acid modifications from a set of allowed mutations [e.g., selecting a location and/or particular modification (e.g., substitution, deletion, insertion) from the set of allowed mutations].
[0023] In some embodiments, a set of allowed mutations is or comprises (e.g., a list, table, etc. representing) a plurality of mutations observed as occurring (e.g., present at or above a particular rate) within a set of related antigens.
[0024] In some embodiments, a reference antigen is a SARS-CoV-2 (e.g., Spike) protein of a particular variant and the set of related antigens comprises corresponding proteins of other, related, variants (e.g., within a particular set of lineages and/or sub-lineages including the reference antigen).
[0025] In some embodiments, a reference antigen is a member of an Omicron lineage and the set of related antigens comprises observed variants belonging to the Omicron lineage.
[0026] In some embodiments, a set of allowed mutations comprises at least a portion (e.g., a subset; e.g., all) of mutations listed in Table 3 (e.g., excluding those removed, as marked with strikethrough in Table 3) (e.g., and, optionally, one or more additional mutations; e.g., and not any other mutations).
[0027] In some embodiments, a set of allowed mutations comprises at least a portion (e.g., a subset) of mutations listed in Table 3, excluding one or both of L335F and L390R.
[0028] In some embodiments, a reference antigen is a SARS-CoV2 (e.g., Spike) protein and the set of related antigens comprises corresponding proteins of other (e.g., SARS-CoV 1, MERS) coronavirus.
[0029] In some embodiments, one or more memory -triggering conserved regions are or comprise a set of conserved epitope regions and step (c) comprises introducing at least one
amino acid modification within each of at least a portion (e.g., all; e.g., a particular subset) of the conserved epitope regions.
[0030] In some embodiments, one or more memory -triggering conserved regions are or comprise a conserved surface and step (c) comprises generating the one or more amino acid modifications at positions distributed (e.g., approximately evenly, as opposed to clustered) throughout / across the conserved surface [e.g., distributed in a substantially equidistant manner over a 3D representation of the conserved surface (e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
[0031] In some embodiments, provided methods comprise performing steps (b) and (c) repeatedly to generate a plurality of candidate polypeptide models, each representing a candidate engineered variant (e.g., each candidate polypeptide model comprising a distinct combination of amino acid modifications and representing a unique artificially engineered version of the reference antigen).
[0032] In some embodiments, provided methods comprise determining, by the processor, values of one or more performance scores for each of the candidate polypeptide models and selecting a subset of the candidate polypeptide models based at least in part on the determined performance score values.
[0033] In some embodiments, one or more performance scores comprise one or both of: (a) an immune escape score indicative of a likelihood and/or relative capability of a particular candidate engineered variant to be recognized and neutralized by antibodies, and (b) a fitness score indicative of a likelihood and/or viability of a particular candidate engineered variant.
[0034] In some embodiments, determining one or both of (a) an immune escape score and (b) a fitness score comprises using a machine learning model [e.g., a language model that receives, as input, a representation of an amino acid sequence of the particular candidate engineered variant (e.g., wherein the input does not comprise a 3D structural representation of the reference)] [e.g., to compute a likelihood score representing a predicted likelihood of the particular candidate variant occurring; e.g., to compute a semantic change score indicative of a distance between an embedding representation of (i) the particular candidate variant (e.g.,
generated from one or more hidden layers of the machine learning model) and (ii) one or more reference variants (c.g.., a WT variant; c.g., a variant of the reference antigen with which a subject has previously been infected, e.g., via vaccination and/or natural exposure; e.g., the reference antigen)].
[0035] In some embodiments, determining one or both of (a) an immune escape score and (b) a fitness score comprises using a 3D structural model of at least a portion of the particular candidate variant [e.g., computing a viral polypeptide receptor binding score (e.g., an ACE2 binding score); e.g., computing an epitope alteration score],
[0036] In some embodiments, one or more performance scores comprise(s) a position spread score that measures an extent to which amino acid modifications are evenly distributed across a surface (e.g., the conserved surface) of the candidate engineered variant.
[0037] In some embodiments, one or more performance scores comprise(s) a mutation co-occurrence score (e.g., that measures a degree to which one or more amino acid modifications of a particular candidate variant are aligned with a natural co-occurrence rates).
[0038] In some embodiments, provided methods comprise identifying, by the processor, within the polypeptide model, one or more target regions representing portions (e.g., target epitopes) of the reference antigen to be retained and excluding the one or more target regions from the one or more memory-triggering conserved region(s) (e.g., wherein the reference antigen is a SARS-CoV-2 S protein and/or portion thereof and the one or more target regions are or comprise an ACE2 binding interface, e.g., thereby preserving ACE2 binding interface regions).
[0039] In some embodiments, provided methods comprise causing, by the processor, rendering of the disrupted polypeptide model for graphical display.
[0040] In some embodiments, provided methods comprise generating, from the disrupted polypeptide model, a corresponding RNA sequence.
[0041] In some embodiments, provided methods comprise producing (e.g., as a vaccine) a composition comprising a polypeptide based on (e.g., having a substantially same amino acid sequence as represented by) the disrupted polypeptide model.
[0042] In some embodiments, provided methods comprise assessing the biological activity of the engineered antigen in vitro.
[0043] In some embodiments, biological activity of the engineered antigen is characterized in that: the engineered antigen is expressed and folded properly [e.g., based on a binding assay, such as a flow-cytometry-based binding assay (e.g., with hACE2)]; and/or the engineered antigen does not bind to antibodies that bind to the reference antigen [e.g., shows reduced/abrogated binding (in comparison with the reference antigen) for a panel of one or more neutralizing antibodies (e.g., antibodies that bind to distinct epitope classes; e.g., one or more (e.g., each of) classes A, B, C, D, E, and F)]; and/or pseudoviruses loaded with the engineered antigen are able to enter cells; and/or the engineered antigen is immunogenic, and/or the engineered antigen reduces the engineered antigen’s activation of the B cell memory immune response to the reference antigen.
[0044] In some embodiments, provided methods comprise producing (e.g., as a vaccine) a composition comprising a nucleic acid encoding the amino acid sequence represented by the disrupted polypeptide model.
[0045] In some aspects, the present disclosure provides vaccine compositions comprising a polypeptide and/or nucleic acid of one or more aspects or embodiments described herein (e.g., in paragraphs above).
[0046] In some aspects, the present disclosure provides methods of vaccination comprising administering to a subject or a population of subjects provided vaccines according to one or more aspects or embodiments described herein (e.g., in paragraphs above).
[0047] In some aspects, the present disclosure provides systems comprising a processor of a computing device and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method of one or more aspects or embodiments described herein (e.g., in paragraphs above).
[0048] In some aspects, the present disclosure provides methods of manufacturing an immunogenic composition comprising: comparing sequences of viral proteins from different variants of an infectious disease agent (e.g., influenza; e.g., SARS-CoV-2) to identify (residual) conserved sites in an antigen of interest (e.g., via one or more systems and/or methods, including any of those recited in the preceding claims); replacing at least one or more of the residual conserved sites with a sequence that is characterized by [features/characteristics of such amino acid substitutions, e.g., providing new immunogenic epitopes from different variants] to generate
a new sequence (e.g., via one or more systems and/or methods, including any of those recited in the preceding claims); and producing a vaccine that delivers at least a portion of the new sequence including at least one of the residual conserved sites replaced.
[0049] In some aspects, the present disclosure provides RNA comprising a nucleotide sequence that encodes an engineered antigen [e.g., for (e.g., characterized in that) eliciting an immune response directed to one or more target epitopes of a reference antigen of an infectious agent while reducing the engineered antigen’s activation of a (e.g., B cell and/or T cell) memory immune response to the reference antigen], wherein the engineered antigen corresponds to (e.g., has a sequence of) a particular reference antigen having been (e.g., artificially) altered to introduce one or more amino acid modifications within at least a portion of one or more memory-triggering conserved regions having been identified as portions of the reference antigen that are determined likely to trigger a memory immune response [e.g., portions that are determined to (i) correspond to known epitopes and/or potential epitopes and/or (ii) do not comprise hallmark mutations of the reference antigen] [e.g., such that the engineered antigen is an artificially engineered version of the reference antigen with mutations introduced to disrupt its conserved portions (e.g., that are/were determined likely to trigger memory immune response)].
[0050] In some embodiments, one or more memory -triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar (e.g., having sufficient sequence similarity; e.g., common) to one or more (e.g., pre-existing) variants thereof.
[0051] In some embodiments, one or more memory -triggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and unmutated on the reference antigen (e.g., known epitopes without any hallmark mutations).
[0052] In some embodiments, a set of known epitopes comprise one or more of the epitopes listed in Tables 2A and/or 2B.
[0053] In some embodiments, one or more memory -triggering conserved region(s) are or comprise a conserved surface representing a (e.g., continuous; e.g., contiguous) un-mutated (e.g., lacking any hallmark mutations) surface of the reference antigen [e.g., wherein the target polypeptide is or comprises an XBB.1.5 variant of an RBD of a SARS-CoV-2 Spike (S) protein and the conserved surface is or comprises the following positions: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372,
F377, K378, Y38O, G381 , S383, P384, T385, K386, N388, D389, L390, C391 , F392, T393, N394, Y396, P412, G413, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, 1468, S469, T470, E471, 1472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527],
[0054] In some embodiments, a reference antigen is or comprises at least a portion of (e.g., an RBD of) an XBB.1.5 variant of a SARS-Cov2 Spike protein [e.g., and the hallmark mutations comprise at least a portion (e.g., all) of the following mutations: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K],
[0055] In some embodiments, one or more amino acid modifications are selected from a set of allowed mutations [e.g., selecting a location and/or particular modification (e.g., substitution, deletion, insertion) from the set of allowed mutations].
[0056] In some embodiments, a set of allowed mutations is or comprises a plurality of mutations observed as occurring (e.g., present at or above a particular rate) within a set of related polypeptides.
[0057] In some embodiments, a set of related polypeptides comprises corresponding polypeptides of other, related, SARS-CoV 2 variants (e.g., within a particular set of lineages and/or sub-lineages including the particular SARS-CoV 2 variant).
[0058] In some embodiments, a target (SARS-CoV 2 variant) polypeptide is a member of an Omicron lineage and the set of related polypeptides comprises corresponding polypeptides (e.g., Spike protein sequences) of observed variants belonging to the Omicron lineage.
[0059] In some embodiments, one or more memory -triggering conserved regions are or comprise a set of conserved epitope regions and the engineered antigen has at least one amino acid modification within each of at least a portion (e.g., all; e.g., a particular subset) of the conserved epitope regions.
[0060] In some embodiments, one or more memory-triggering conserved regions are or comprise a conserved surface and the one or more amino acid modifications occur at positions distributed (e.g., approximately evenly, as opposed to clustered) throughout / across the conserved surface [e.g., distributed in a substantially equidistant manner over a 3D representation of the conserved surface (e.g., wherein 3D linear distances and/or geodesic distances (traversing the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and/or distributed according to a particular pre-defined statistical distribution (e.g., normal distribution)].
[0061] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, second or third column) in Table 5A [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in the third column of Table 5A)][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in the second column of Table 5A].
[0062] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the sequences listed in (e.g., each row of) Table 5B [e.g., wherein the engineered antigen is or comprises a SARS- CoV-2 S protein RBD having any one of the sequences listed in Table 5B (e.g., any one of SEQ ID NOs: 7 through 34)].
[0063] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine
compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, first or second column) in Table 7A and/or Table 7B [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 7A and/or Table 7B (e.g., identified as LI, L2, L3, and L4)].
[0064] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (e.g., in each row, second or third column) in Table 8A [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in third column of Table 8A][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 8A].
[0065] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the sequences listed in (e.g., each row of) Table 8B [e.g., wherein the engineered antigen is or comprises a SARS- CoV-2 S protein RBD having any one of the sequences listed in Table 8B (e.g., any one of SEQ ID NOs: 39 through 64 and 100)].
[0066] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation
combinations listed (e.g., in each row, second column) in Table 9A [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 9A].
[0067] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S43 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 S protein RBD with additional mutations according to S43 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.l .5 variant S protein RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with additional mutations N360D P384S L390R T430I F464Y H519N].
[0068] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: N360D, P384S, L390R, T430I, F464Y, H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0069] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise the mutation combinations identified as S48 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 RBD with additional mutations according to S48 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the XBB.1.5 hallmark
mutations identified in Table IB and) with additional mutations L335F K356T P384S L390R T430I F464Y H519N].
[0070] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: L335F, K356T, P384S, L390R, T4301, F464Y, and H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0071] In some embodiments, [e.g., wherein the reference antigen is a particular target valiant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigen of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion of a SARS-CoV-2 S protein (e.g., a RBD) with mutations P384S L390R T430I F464Y H519N.
[0072] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion of a SARS-CoV-2 S protein (e.g., a RBD) with one or more of (e.g., a subset of; e.g., all of) mutations K356T P384S L390R T430I F464Y H519N.
[0073] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is or comprises an S protein RBD) and] provided methods, systems, vaccine compositions, method of manufacturing, or RNAs do not include one or both of mutations L335F and L390R.
[0074] In some embodiments, [e.g., wherein the reference antigen is a particular target variant of a SARS-CoV-2 S protein and/or portion thereof (e.g., wherein the reference antigen is
or comprises an S protein RBD) and], provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise one or more of the mutation combinations listed (c.g., in each row of) Table 15A and/or Table 15B [e.g., wherein the engineered antigen comprises an RBD of a SARS-CoV-2 S protein (e.g., according to any one of SEQ ID NOs: 4 or 38) (e.g., residues 327 to 528 of SEQ ID NO: 1) with one or more of the mutation combinations listed in Table 15A][e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant RBD (e.g., according to any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table IB and) with one or more of the additional mutation combinations listed in second column of Table 15B],
[0075] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 23 in Table 15 A and/or Table 15B.
[0076] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: I332V L335F K356T P384S T430I L452Q F464Y H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0077] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: I332V L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452Q N460K F464Y S477N T478K E484A F486P F490S Q498R N501 Y Y505H E516Q H519N.
[0078] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 22 in Table 15A and/or Table 15B.
[0079] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: L335F K.356T P384S T430I L452R F464Y H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.l .5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0080] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H E516Q H519N T523S.
[0081] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 09 in Table 15A and/or Table 15B.
[0082] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T N360S P384S N388K T430I N450D F464Y H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.l.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F,
S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0083] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S N388K D405N R408S K417N T4301 N440K V445P G446S N450D N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N T523S.
[0084] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S129 in Table 15 A and/or Table 15B.
[0085] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T L335F P384S D389G T430I N450D F464Y H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0086] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
[0087] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 56 in Table 15 A and/or Table 15B.
[0088] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T P384S L390R T430I N450D F464Y I472V H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0089] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S L390R D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y I472V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
[0090] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as SI 12 in Table 15A and/or Table 15B.
[0091] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs are or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T P384S D389G T430I N450D F464Y I468V H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0092] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (c.g., up to all) of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G D405N R408S K417N T430I N440K V445P G446S N450D N460K F464Y I468V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
[0093] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise mutation combinations identified as S125 in Table 15A and/or Table 15B.
[0094] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs arc or comprise at least a portion [e.g., a particular domain (e.g., an RBD)] of a SARS-CoV-2 S protein (e.g., according to SEQ ID NO: 1) with at least a portion of (e.g., all) of the following mutations: K356T L335F P384S T430I F464Y I468V H519N {e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., those within an RBD of the SARS-CoV-2 S protein) [e.g., wherein the one or more XBB.1.5 hallmark mutations as follows T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L3681, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K] }.
[0095] In some embodiments, engineered antigens of provided methods, systems, vaccine compositions, method of manufacturing, or RNAs comprise at least a portion (e.g., up to all) of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S N460K F464Y I468V S477N T478K E484A F486P F490S Q498R N501 Y Y505H E516Q H519N T523S.
[0096] In some aspects, the present disclosure provides methods of manufacturing an RNA comprising a nucleotide sequence that encodes an engineered antigen corresponding to an engineered version of a reference antigen, the method comprising producing an RNA whose nucleotide sequence, when compared with (e.g., a nucleotide sequence of) the reference antigen shows difference(s) (e.g., comprises nucleotides encoding for one or more mutations) relative to the reference antigen in one or more memory-triggering conserved regions [e.g., one or more
conserved epitopes and/or one or more conserved surface(s)] that are common to the reference antigen and (i) one or more pre-existing variants of the reference antigen and/or (ii) a wild-type strain of the reference antigen.
[0097] In some embodiments, a reference antigen is or comprises a particular target variant SARS-CoV-2 S protein RBD [e.g., an XBB.1.5 RBD (e.g., having an amino acid sequence according to any one of SEQ ID NOs: 35, 36, and 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 with the XBB.1.5 hallmark mutations identified in Table IB)].
[0098] In some embodiments, RNA manufactured according to provided methods (e.g., as described above) is or comprises the RNA of any one the aspects and embodiments described herein (e.g., wherein the differences in comparison with the reference antigen correspond/comprise nucleotide differences encoding for any of the mutation combinations of various aspects and embodiments described herein, e.g., in paragraphs above).
[0099] In some embodiments, provided methods of manufacturing comprise producing RNA via in vitro transcription (IVT).
[0100] Features of embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention
BRIEF DESCRIPTION OF THE DRAWING
[0101] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0102] FIG. 1 is a block-flow diagram and schematic illustrating an example process for design of an engineered antigen, according to an illustrative embodiment.
[0103] FIG. 2 is a block-flow diagram of an example process for inserting amino acid modifications, according to an illustrative embodiment.
[0104] FIG. 3 is a block-flow diagram of an example process for generating and scoring multiple candidate antigen designs, according to an illustrative embodiment.
[0105] FIG. 4A is a set of views of a 3D model of a receptor-binding domain (RBD) of a SARS-Cov2 Spike (.S') protein of an XBB variant, colorized to show XBB hallmark mutations and an identified conserved region, according to an illustrative embodiment.
[0106] FIG. 4B is another set of views of the XBB RBD .S' protein model shown in FIG. 6A, colorized to show XBB hallmark mutations, an identified conserved region, and an ACE2 binding interface, according to an illustrative embodiment.
[0107] FIG. 4C is another set of views of the XBB RBD .S' protein model shown in FIGs. 6 A and 6B, colorized to show a relative frequency of epitope involvement at various amino acid sites within the protein, according to an illustrative embodiment.
[0108] FIG. 4D is a set of views of a 3D model of a first engineered antigen design, generated using the XBB RBD 5 protein as a reference antigen.
[0109] FIG. 4E is a set of views of a 3D model of a second engineered antigen design, generated using the XBB RBD ,S' protein as a reference antigen.
[0110] FIG. 5A is a schematic of SARS CoV 2 and its relevant proteins, adapted from Jain el al., Vaccines 2020, 8(4), 649 (www.mdpi.com/2076-393X/8/4/649.).
[0111] FIG. 5B is a set of views of a 3D model of a receptor-binding domain (RBD) of a SARS-Cov2 Spike (.S') protein of an XBB. 1 .5 variant, colorized to show XBB hallmark mutations and an identified conserved region, according to an illustrative embodiment.
[0112] FIG. 5C is a set of views of a 3D model of an XBB.1.5 spike protein (full spike protein), according to an illustrative embodiment.
[0113] FIG. 5D is another set of views of the XBB RBD 5 protein model shown in FIG. 7B, colorized to show XBB hallmark mutations, an identified conserved region, and an ACE2 binding interface, according to an illustrative embodiment.
[0114] FIG. 5E is another set of views of the XBB RBD 5 protein model shown in FIGs. 5B and 5D, colorized to show a relative frequency of epitope involvement at various amino acid sites within the protein, according to an illustrative embodiment.
[0115] FIG. 5F is another set of views of the XBB RBD S protein model shown in FIGs. 5B and 5D, colorized to show a relative frequency of epitope involvement at various amino acid sites within a conserved surface of the protein, according to an illustrative embodiment.
[0116] FIG. 5G is a block flow diagram of an example process for generating candidate variant designs, according to an illustrative embodiment.
[0117] FIGs. 6A shows a set of views of a 3D model of candidate engineered antigen design, generated using an XBB.1.5 RBD 5 protein as a reference antigen, according to an illustrative embodiment. Colorization/shading in FIG. 6A identifies XBB.1.5 hallmark mutations, an identified conserved region, and amino acid sites mutated based on various in- silico antigen design approaches described herein.
[0118] FIG. 6B shows a set of views of a 3D model of a SARS-CoV-2 spike protein with XBB.1.5 hallmark mutations and additional mutations corresponding to those shown in FIG. 6A highlighted.
[0119] FIG. 6C shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
[0120] FIG. 6D shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
[0121] FIG. 6E shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 5 protein RBD as a reference antigen, according to an illustrative embodiment.
[0122] FIG. 6F shows a set of views of a 3D model of another candidate engineered antigen design based on (e.g., using) an XBB.1.5 .S' protein RBD as a reference antigen, according to an illustrative embodiment.
[0123] FIG. 7 shows radar plots of certain performance scores for candidate engineered antigen designs, according to an illustrative embodiment.
[0124] FIG. 8 shows a plot of predicted versus experimental ACE2 binding change, according to an illustrative embodiment.
[0125] FIG. 9A shows a view of a 3D model for an engineered RBD construct, according to an illustrative embodiment .
[0126] FIG. 9B shows a view of a 3D model for an engineered RBD construct, according to an illustrative embodiment.
[0127] FIG. 9C is a schematic of a SARS-CoV-2 trimer, adapted from Starr et al., SARS- CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021 ). https://doi.org/! 0.1038/s41586-021 -03807-6.
[0128] FIG. 9D shows a view of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes on an XBB.1.5 backbone, according to an illustrative embodiment.
[0129] FIG. 9E shows a view of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes on an XBB.1.5 backbone, according to an illustrative embodiment.
[0130] FIG. 10A is a diagram of an example procedure for testing engineered antigens, according to an illustrative embodiment.
[0131] FIG. 10B is a block-flow diagram of an example protocol for generating and evaluating engineered constructs as described herein, according to an illustrative embodiment.
[0132] FIG. 10C is a set of graphs evaluating antibody binding to WT and XBB.1.5 S protein for different epitope classes, according to an illustrative embodiment.
[0133] FIG. 10D is a graph of a dilution series for a hACE2-mFc binding agent.
[0134] FIG. 10E is a graph of a dilution series of human BNT162b23 (triple-vax) polyclonal serum.
[0135] FIG. 11 A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0136] FIG. 11B shows binding dose response curves for certain engineered SARS-CoV
2 antigens described herein.
[0137] FIG. 11C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0138] FIG. 12A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0139] FIG. 12B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0140] FIG. 13A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0141] FIG. 13B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0142] FIG. 13C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0143] FIG. 13D shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0144] FIG. 14A shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0145] FIG. 14B shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0146] FIG. 14C shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0147] FIG. 14D shows binding dose response curves for certain engineered SARS-CoV 2 antigens described herein.
[0148] FIG. 15A is a schematic illustrating example groups and a dosage / sampling schedule for evaluating immune imprinting and performance of candidate vaccines in vaccine- cxpcricnccd mice. Yellow-filled cells indicate days on which sera sample will be collected, gray-filled cells indicate days on which vaccines will be administered, and green-filled cells indicate days on which mice are sacrificed and final samples collected, according to an illustrative embodiment.
[0149] FIG. 15B is a schematic illustrating example groups and a dosage I sampling schedule for evaluating performance of candidate vaccines in vaccine naive mice. Yellow-filled cells indicate days on which sera sample will be collected, gray-filled cells indicate days on which vaccines will be administered, and green-filled cells indicate days on which mice are sacrificed and final samples collected, according to an illustrative embodiment.
[0150] FIG. 16A is a schematic illustrating FACS-based assay, e.g., using approaches described in Quandt and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety, according to an illustrative embodiment.
[0151] FIG. 16B is a schematic illustrating depletion assay, e.g., using approaches described in Quandt and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety, according to an illustrative embodiment.
[0152] FIG. 17 is a schematic illustrating approaches for collecting and analyzing spleen sample, lymph nodes as well as analysis of blood samples, according to certain embodiments.
[0153] FIG. 18 is a diagram illustrating mutations included in various engineered antigen construct designs, according to an illustrative embodiment.
[0154] FIG. 19A is a schematic illustrating an example construct comprising a membrane anchored RBD+fibritin domain with viral signal peptide, according to an illustrative embodiment.
[0155] FIG. 19B is a graph showing ACE2 binding assay results for a set of engineered construct designs, according to an illustrative embodiment.
[0156] FIG. 19C is a graph showing serum (from vaccinated patients) binding assay results for a set of engineered construct designs, according to an illustrative embodiment.
[0157] FIG. 20A is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
[0158] FIG. 20B is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
[0159] FIG. 20C is a set of graphs showing binding assay results against a set of monoclonal antibodies, according to an illustrative embodiment.
[0160] FIG. 21 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.
[0161] FIG. 22 is a diagram of an example computing device and an example mobile computing device used in certain embodiments.
[0162] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and/or structurally similar elements.
CERTAIN DEFINITIONS
[0163] About or Approximately: The term “about” or “approximately”, when used herein in reference to a value, refers to a value that is similar to the referenced value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” or “approximately” in that context. For example, in some embodiments, the term “about” or “approximately” may encompass a range of values that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.
[0164] Administration'. As used herein, the term “administration” typically refers to the administration of a composition to a subject or system. Those of ordinary skill in the art will be aware of a variety of routes that may, in appropriate circumstances, be utilized for administration to a subject, for example a human. For example, in some embodiments, administration may be ocular, oral, parenteral, topical, etc. In some particular embodiments, administration may be bronchial (e.g., by bronchial instillation), buccal, dermal (which may be or comprise, for example, one or more of topical to the dermis, intradermal, interdermal, transdermal, etc.), enteral, intra-arterial, intradermal, intragastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intraventricular, within a specific organ (e. g. intrahepatic), mucosal, nasal, oral, rectal, subcutaneous, sublingual, topical, tracheal (e.g., by intratracheal instillation), vaginal, vitreal, etc. In some embodiments, administration may involve
dosing that is intermittent (e.g., a plurality of doses separated in time) and/or periodic (e.g., individual doses separated by a common period of time) dosing. In some embodiments, administration may involve continuous dosing (e.g., perfusion) for at least a selected period of time.
[0165] Adult: As used herein, the term “adult” refers to a human eighteen years of age or older. In some embodiments, a human adult has a weight within the range of about 90 pounds to about 250 pounds.
[0166] Affinity. As is known in the art, “affinity” is a measure of the tightness with which two or more binding partners associate with one another. Those skilled in the art are aware of a variety of assays that can be used to assess affinity, and will furthermore be aware of appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity is assessed over a plurality of concentrations (e.g., of one binding partner at a time). In some embodiments, affinity is assessed in the presence of one or more potential competitor entities (e.g., that might be present in a relevant - e.g., physiological - setting). In some embodiments, affinity is assessed relative to a reference (e.g., that has a known affinity above a particular threshold [a “positive control” reference] or that has a known affinity below a particular threshold [ a “negative control” reference”]. In some embodiments, affinity may be assessed relative to a contemporaneous reference; in some embodiments, affinity may be assessed relative to a historical reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
[0167] Agent'. In general, the term “agent”, as used herein, is used to refer to an entity (e.g., for example, a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or complex, combination, mixture or system [e.g., cell, tissue, organism] thereof), or phenomenon (e.g., heat, electric current or field, magnetic force or field, etc.). In appropriate circumstances, as will be clear from context to those skilled in the art, the term may be utilized to refer to an entity that is or comprises a cell or organism, or a fraction, extract, or component thereof. Alternatively or additionally, as context will make clear, the term may be used to refer to a natural product in that it is found in and/or is obtained from nature. In some instances, again as will be clear from context, the term may be used to refer to one or more entities that is man-made in that it is designed, engineered, and/or produced through action of the hand of man and/or is
not found in nature. In some embodiments, an agent may be utilized in isolated or pure form; in some embodiments, an agent may be utilized in crude form. In some embodiments, potential agents may be provided as collections or libraries, for example that may be screened to identify or characterize active agents within them. In some cases, the term “agent” may refer to a compound or entity that is or comprises a polymer; in some cases, the term may refer to a compound or entity that comprises one or more polymeric moieties. In some embodiments, the term “agent” may refer to a compound or entity that is not a polymer and/or is substantially free of any polymer and/or of one or more particular polymeric moieties. In some embodiments, the term may refer to a compound or entity that lacks or is substantially free of any polymeric moiety.
[0168] Amelioration-. The term “amelioration,” as used herein, refers to the prevention, reduction or palliation of a state, or improvement of the state of a subject. Amelioration includes, but does not require complete recovery or complete prevention of a disease, disorder or condition (e.g., radiation injury).
[0169] Amino acid: The term “amino acid,” its broadest sense, as used herein, the term “amino acid” refers to a compound and/or substance that can be, is, or has been incorporated into a polypeptide chain, e.g., through formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure H2N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally-occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. “Standard amino acid” refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides.
“Nonstandard amino acid” refers to any amino acid, other than the standard amino acids, regardless of whether it is prepared synthetically or obtained from a natural source. In some embodiments, an amino acid, including a carboxy-and/or amino-terminal amino acid in a polypeptide, can contain a structural modification as compared with the general structure above. For example, in some embodiments, an amino acid may be modified by methylation, amidation, acetylation, pegylation, glycosylation, phosphorylation, and/or substitution (e.g., of the amino group, the carboxylic acid group, one or more protons, and/or the hydroxyl group) as compared with the general structure. In some embodiments, such modification may, for example, alter the
circulating half-life of a polypeptide containing the modified amino acid as compared with one containing an otherwise identical unmodified amino acid. In some embodiments, such modification does not significantly alter a relevant activity of a polypeptide containing the modified amino acid, as compared with one containing an otherwise identical unmodified amino acid. As will be clear from context, in some embodiments, the term “amino acid” may be used to refer to a free amino acid; in some embodiments it may be used to refer to an amino acid residue of a polypeptide.
[0170] Animal'. As used herein, the term “animal” refers to any member of the animal kingdom. In some embodiments, "animal" refers to humans, of either sex and at any stage of development. In some embodiments, "animal" refers to non-human animals, at any stage of development. In certain embodiments, the non-human animal is a mammal (e.g., a rodent, a mouse, a rat, a rabbit, a monkey, a dog, a cat, a sheep, cattle, a primate, and/or a pig). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and/or worms. In some embodiments, an animal may be a transgenic animal, genetically engineered animal, and/or a clone.
[0171] Antibody. As used herein, the term “antibody” refers to a polypeptide that includes canonical immunoglobulin sequence elements sufficient to confer specific binding to a particular target antigen. As is known in the ail, intact antibodies as produced in nature are approximately 150 kD tetrameric agents comprised of two identical heavy chain polypeptides (about 50 kD each) and two identical light chain polypeptides (about 25 kD each) that associate with each other into what is commonly referred to as a “Y-shaped” structure. Each heavy chain is comprised of at least four domains (each about 110 amino acids long)- an amino-terminal variable (VH) domain (located at the tips of the Y structure), followed by three constant domains: CHI, CH2, and the carboxy -terminal CH3 (located at the base of the Y’s stem). A short region, known as the “switch”, connects the heavy chain variable and constant regions. The “hinge” connects CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect the two heavy chain polypeptides to one another in an intact antibody. Each light chain is comprised of two domains - an amino-terminal variable (VL) domain, followed by a carboxy -terminal constant (CL) domain, separated from one another by another “switch”. Intact antibody tetramers are comprised of two heavy chain-light chain dimers in
which the heavy and light chains are linked to one another by a single disulfide bond; two other disulfide bonds connect the heavy chain hinge regions to one another, so that the dimers arc connected to one another and the tetramer is formed. Naturally-produced antibodies are also glycosylated, typically on the CH2 domain. Each domain in a natural antibody has a structure characterized by an “immunoglobulin fold” formed from two beta sheets (e.g., 3-, 4-, or 5- stranded sheets) packed against each other in a compressed antiparallel beta barrel. Each variable domain contains three hypervariable loops known as “complement determining regions” (CDR1, CDR2, and CDR3) and four somewhat invariant “framework” regions (FR1, FR2, FR3, and FR4). When natural antibodies fold, the FR regions form the beta sheets that provide the structural framework for the domains, and the CDR loop regions from both the heavy and light chains are brought together in three-dimensional space so that they create a single hypervariable antigen binding site located at the tip of the Y structure. The Fc region of naturally-occurring antibodies binds to elements of the complement system, and also to receptors on effector cells, including for example effector cells that mediate cytotoxicity. As is known in the art, affinity and/or other binding attributes of Fc regions for Fc receptors can be modulated through glycosylation or other modification. In some embodiments, antibodies produced and/or utilized in accordance with the present disclosure include glycosylated Fc domains, including Fc domains with modified or engineered such glycosylation. For purposes of the present disclosure, in certain embodiments, any polypeptide or complex of polypeptides that includes sufficient immunoglobulin domain sequences as found in natural antibodies can be referred to and/or used as an “antibody”, whether such polypeptide is naturally produced (e.g., generated by an organism reacting to an antigen), or produced by recombinant engineering, chemical synthesis, or other artificial system or methodology. In some embodiments, an antibody is polyclonal; in some embodiments, an antibody is monoclonal. In some embodiments, an antibody has constant region sequences that are characteristic of mouse, rabbit, primate, or human antibodies. In some embodiments, antibody sequence elements are humanized, primatized, chimeric, etc., as is known in the ail. Moreover, the term “antibody” as used herein, can refer in appropriate embodiments (unless otherwise stated or clear from context) to any of the art-known or developed constructs or formats for utilizing antibody structural and functional features in alternative presentation.
[0172] For example, in some embodiments, an antibody utilized in accordance with the present disclosure is in a format selected from, but not limited to, intact IgA, IgG, IgE or IgM antibodies; bi- or multi- specific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab’ fragments, F(ab’)2 fragments, Fd’ fragments, Fd fragments, and isolated CDRs or sets thereof; single chain Fvs; polypeptide-Fc fusions; single domain antibodies, alternative scaffolds or antibody mimetics (e.g., anticalins, FN3 monobodies, DARPins, Affibodies, Affilins, Affimers, Affitins, Alphabodies, Avimers, Fynomers, Im7, VLR, VNAR, Trimab, CrossMab, Trident); nanobodies, binanobodies, F(ab’)2, Fab’, di-sdFv, single domain antibodies, trifunctional antibodies, diabodies, and minibodies etc. In some embodiments, relevant formats may be or include: Adnectins®; Affibodies®; Affilins®; Anticalins®;
Avimers®; BiTE®s; cameloid antibodies; Centyrins®; ankyrin repeat proteins or DARPINs®; dual-affinity re targeting (DART) agents; Fynomers®; shark single domain antibodies such as IgNAR; immune mobilizing monoclonal T cell receptors against cancer (ImmTACs); KALBITOR®s; MicroProteins; Nanobodies® minibodies; masked antibodies (e.g., Probodies®); Small Modular ImmunoPharmaceuticals (“SMIPsTM”); single chain or Tandem diabodies (TandAb®); TCR-like antibodies;, Trans-bodies®; TrimerX®; VHHs. In some embodiments, an antibody may lack a covalent modification (e.g., attachment of a glycan) that it would have if produced naturally. In some embodiments, an antibody may contain a covalent modification (e.g., attachment of a glycan, a payload [e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.], or other pendant group [e.g., poly-ethylene glycol, etc.])
[0173] Antigen: The term “antigen”, as used herein, refers to an agent that elicits an immune response; and/or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or to an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including production of antigen- specific antibodies); in some embodiments, an antigen elicits a cellular response (e.g., involving T-cells whose receptors specifically interact with the antigen). In some embodiments, an antigen binds to an antibody and may or may not induce a particular physiological response in an organism. In general, an antigen may be or include any chemical entity such as, for example, a small molecule, a nucleic acid, a polypeptide, a carbohydrate, a lipid, a polymer (in some embodiments other than a biologic polymer [e.g., other than a nucleic acid or amino acid polymer) etc. In some embodiments, an antigen is or comprises a polypeptide. In some embodiments, an antigen is or comprises a glycan. Those of
ordinary skill in the art will appreciate that, in general, an antigen may he provided in isolated or pure form, or alternatively may be provided in crude form (c.g., together with other materials, for example in an extract such as a cellular extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in a crude form. In some embodiments, an antigen is a recombinant antigen.
[0174] Arisen epitope: As used herein, the term “arisen epitope(s)” is used to refer an epitope, e.g., a portion of an antigen, which a subject’s immune system has not encountered previously. For example, as circulating viral pathogens mutate, mutations in various epitopes found on viral proteins can result arisen epitopes that are versions of epitopes present on proteins of previously circulating variants, with one or more mutations introduced.
[0175] Composition '. Those skilled in the art will appreciate that the term “composition” may be used to refer to a discrete physical entity that comprises one or more specified components. In general, unless otherwise specified, a composition may be of any form - e.g., gas, gel, liquid, solid, etc.
[0176] Comprising: A composition or method described herein as “comprising” one or more named elements or steps is open-ended, meaning that the named elements or steps are essential, but other elements or steps may be added within the scope of the composition or method. To avoid prolixity, it is also understood that any composition or method described as "comprising" (or which "comprises") one or more named elements or steps also describes the corresponding, more limited composition or method "consisting essentially of (or which "consists essentially of) the same named elements or steps, meaning that the composition or method includes the named essential elements or steps and may also include additional elements or steps that do not materially affect the basic and novel characteristic(s) of the composition or method. It is also understood that any composition or method described herein as "comprising" or "consisting essentially of one or more named elements or steps also describes the corresponding, more limited, and closed-ended composition or method "consisting of (or "consists of) the named elements or steps to the exclusion of any other unnamed element or step. In any composition or method disclosed herein, known or disclosed equivalents of any named essential element or step may be substituted for that element or step.
[0177] Determine'. In some embodiments, the methodologies described herein include a step of “determining”. Those of ordinary skill in the art, reading the present specification, will appreciate that such “determining” can utilize or be accomplished through use of any of a variety of techniques available to those skilled in the art, including for example specific techniques explicitly referred to herein. In some embodiments, determining involves manipulation of a physical sample. In some embodiments, determining involves consideration and/or manipulation of data or information, for example utilizing a computer or other processing unit adapted to perform a relevant analysis. In some embodiments, determining involves receiving relevant information and/or materials from a source. In some embodiments, determining involves comparing one or more features of a sample or entity to a comparable reference.
[0178] Engineered Antigen: As used herein, the term “engineered antigen,” is used to refer to an antigen that is or is intended to be artificially created and intentionally introduced to a subject, e.g., in order to generate an immune response (e.g., via vaccination), as opposed to a naturally occurring antigen, evolved through natural processes. For example, as described in further detail herein, in certain embodiments, engineered antigens are or comprise polypeptides. Engineered polypeptide antigens may be designed to match, be similar to, or based on other, reference antigens, which may themselves be engineered antigens or may be naturally occurring antigens. For example, in certain embodiments, an engineered antigen is created by beginning with a structure of a reference antigen and introducing one or more amino acid modifications therein, e.g., to achieve a desired behavior / result upon planned introduction to a subject. In certain embodiments, engineered antigens may be designed in-silico - i.e., via computer- implemented systems and methods, using polypeptide models that represent various physical antigens and other computer representations. In certain embodiments, engineered antigens may be encoded by ribonucleic acid (RNA), which, in turn, may be used to manufacture an engineered antigen (e.g., in-vitro), or which may be directly administered to a subject, e.g., as an RNA vaccine.
[0179] Epitope: As used herein, the term “epitope,” is used to refer to used herein, includes any moiety that is specifically recognized by an immunoglobulin (e.g., antibody or receptor) binding component. In some embodiments, an epitope is comprised of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups
are surface-exposed when the antigen adopts a relevant three-dimensional conformation. In some embodiments, such chemical atoms or groups arc physically near to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms are groups are physically separated from one another when the antigen adopts an alternative conformation (e.g., is linearized).
[0180] Epitope Alteration Score: As used interchangeably herein, the terms “epitope alteration score” and “epitope score” both refer to a measure of alteration to a viral polypeptide at epitope positions. In some embodiments, such alteration can be characterized by the impact of mutation(s) in one or more epitopes of a viral variant on recognition by antibodies (e.g., neutralizing antibodies). For example, in some embodiments, such alteration can be characterized by determining the number of antibodies potentially escaped. In some embodiments, antibodies for characterization have been isolated from patients who have been vaccinated against a disease or who have previously been infected with a disease (e.g., SARS- CoV-2). In some embodiments, antibodies for characterization have previously been shown to bind a reference sequence. In some embodiments, an epitope alteration score can be determined by comparison of mutations in a variant candidate to one or more regions of a reference sequence that have previously been shown to bind antibodies (e.g., through structural data). In some embodiments, an epitope alteration score can be determined by enumerating the number of unique epitopes involving altered positions, as measured across one or more known antibody- viral polypeptide complex structures (e.g., all known antibody -viral polypeptide complex structures).
[0181] In some embodiments, an epitope alteration score is a measure of how many distinct epitopes are evaded by a variant candidate as compared to a reference sequence (e.g., as compared to a wild type sequence). In some embodiments, an epitope alteration score is computed based on known binding sites of antibodies, e.g., as reported in Protein Data Bank. In some embodiments, an epitope alteration score can change over time with identification of new epitope positions and/or discoveries of epitope-binding antibodies. In some embodiments, an epitope alteration score can be used to characterize degree of alteration of a SARS-CoV-2 Spike polypeptide at epitope positions, for example, in some embodiments by counting the number or
percentage of antibodies potentially escaped. In various embodiments described herein, an epitope alteration score can be normalized such that it ranks between 0 and 100%.
[0182] Growth score: As used interchangeably herein, the term “growth,” “growth metric,” or “growth score” refers to a measure of the rate at which a given variant is growing in a subject population (e.g., at a given time). In some embodiments, a growth score refers to lineagelevel growth. For example, in some embodiments, a growth score of a given variant can be determined by referencing growth of a parent species or a known variant of substantially the same lineage, or a known variant having a similar sequence (e.g., a sequence that is at least 90% identical to the given variant). In some embodiments, growth of a given variant is a function of the change in the number of subjects within a subject population who are reported as being infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a defined period of time). In some embodiments, growth of a given variant is a function of the change in the proportion of a subject population infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a defined period of time). In some embodiments, a growth score of a given variant can be an empirically determined by considering sequences associated with a given variant (e.g., in some embodiments including sequences associated with a lineage) that have been observed within a defined period and computing its proportion among all observed sequences at a given time relative to a reference level (e.g., its proportion determined over a defined period of time). For example, in some embodiments, for each lineage, its proportion of sequences among all observed sequences is calculated for an extended period of time (e.g., an eight- week window) and for the most recent time window (e.g., for the last 24 hours, last 48 hours, last 72 hours, last 4 days, last 5 days, last 6 days, or last week), denoted by rextended and riast, respectively. The growth of the lineage is defined by their ratio rextended / rlast, measuring the change of the proportion. In various embodiments described herein, a growth score can be normalized such that it ranks between 0 and 100%.
[0183] Human: In some embodiments, a human is an embryo, a fetus, an infant, a child, a teenager, an adult, or a senior citizen.
[0184] Infectivity Score or “Fitness Prior Score” as used interchangeably herein, the term “infectivity score” or “fitness prior score" is a measure of a viral variant’s evolutionary fitness,
and is a function of the efficiency with which a virus replicates and/or the efficiency with which a virus infects host cells. In some embodiments, calculation of a fitness prior score comprises determining one or more of a log-likelihood score, a viral polypeptide receptor binding score, and/or a growth score. In some embodiments, a fitness prior score is determined by referencing each of a log-likelihood score, a viral polypeptide receptor binding score, and a growth score.
[0185] “Improve ” "increase” . “inhibit” or “reduce”: As used herein, the terms “improve”, “increase”, “inhibit’, “reduce”, or grammatical equivalents thereof, indicate values that are relative to a baseline or other reference measurement. In some embodiments, an appropriate reference measurement may be or comprise a measurement in a particular system (e.g., in a single individual) under otherwise comparable conditions absent presence of (e.g., prior to and/or after) a particular agent or treatment, or in presence of an appropriate comparable reference agent. In some embodiments, an appropriate reference measurement may be or comprise a measurement in comparable system known or expected to respond in a particular way, in presence of the relevant agent or treatment.
[0186] Log-likelihood: As used herein, the term “log-likelihood” refers to a measure of the existence probability of a variant polypeptide sequence, which has been determined using natural language processing algorithms. In some embodiments, log-likelihood can be determined using a transformer model. In some embodiments, log-likelihood can be determined without a reference sequence. In some embodiments, log-likelihood is a transformer-derived log-likelihood without reference. The higher the log-likelihood of a variant, the more probable the variant is to occur from a language model perspective. In various embodiments described herein, loglikelihood can be normalized such that it ranks between 0 and 100%. In some embodiments, a log-likelihood measures how log-likelihood of a variant polypeptide sequence compares to the entire population of known variants. In some embodiments, a log-likelihood measures how loglikelihood of a variant polypeptide sequence compares to other variants with similar mutational loads (“conditional log-likelihood”). Such conditional log-likelihood is particularly useful for assessing variants with high mutation counts (e.g., at least 30 or more, including, e.g., at least 40, at least 50, at least 60, at least 70, or more mutation counts).
[0187] Machine learning module, machine learning model: As used herein, the terms “machine learning module” and “machine learning model” arc used interchangeably and refer to
a computer implemented process (e.g., a software function) that implements one or more particular machine learning algorithms, such as an artificial neural networks (ANN), random forest, decision trees, support vector machines, and the like, in order to determine, for a given input, one or more output values. In certain embodiments, machine learning models are deep learning models or deep neural networks - ANNs that comprise, in addition to an input layer and an output layer, one or more hidden layers (e.g., in between). Examples of deep learning models include, without limitation, recurrent neural networks (RNNs) (e.g., long short-term memory networks (LSTMs), bi-directional LSTMs (biLSTMs)), attention-based networks, such as transformer models, and convolutional neural networks (CNNs). In some embodiments, machine learning modules implementing machine learning techniques are trained in a supervised manner, for example using curated and/or manually annotated datasets. In certain embodiments, machine learning models may be trained in an unsupervised manner, using unlabeled data. In certain embodiments, a machine learning model may be trained via a reinforcement approach, for example wherein a reward / penalty system is used to train a machine learning model to learn strategies for accomplishing specified tasks. Training a machine learning model may be used to determine various parameters of a model, such as weights associated with layers in neural networks. In some embodiments, once a machine learning module is trained, e.g., to accomplish a specific task such as predicting types of hidden amino acids within of polypeptide sequences based on their context, values of determined parameters are fixed and the (e.g., unchanging, static) machine learning module is used to process new data (e.g., different from the training data), such as a new amino acid sequence, referred to as an inference. In some embodiments, machine learning modules may receive feedback, e.g., based on user review of accuracy, and such feedback may be used as additional training data, for example to dynamically update the machine learning module. In some embodiments, a trained machine learning module is a classification algorithm with adjustable and/or fixed (e.g., locked) parameters, e.g., a random forest classifier. In some embodiments, two or more machine learning modules may be combined and implemented as a single module and/or a single software application. In some embodiments, two or more machine learning modules may also be implemented separately, e.g., as separate software applications. A machine learning module may be software and/or hardware. For example, a machine learning module may be implemented entirely as software, or certain functions of an ANN module may be carried out via specialized hardware (e.g. , via an
application specific integrated circuit (ASIC), field programmable gate arrays (FPGAs), and the like).
[0188] Model: As used herein, the term “model” is used to identify a computer representation of a particular physical object or quantity, e.g., that is accessed, displayed by, used as input to, generated as output of, etc., computer-implemented methods and systems and/or one or more steps and/or modules or functions thereof. For example, as described in further detail herein, various computer implemented systems and methods may operate on, process, and generate polypeptide models that represent physical polypeptides, such as particular proteins or portions thereof. Computer representations of polypeptides, e.g., polypeptide models, may be implemented in a variety of formats, such as a string of characters (e.g., letters, each representing a particular amino acid type) representing an amino acid sequence (e.g., a FASTA file), or a 3D structural model that includes information about a (e.g., relative) 3D location of amino acids and/or atoms thereof, such as a Protein Data Bank (PDB) file format which may be used to describe a 3D structure of a particular protein and includes, among other things, atomic coordinates of atoms of particular protein.
[0189] Nucleic acid'. As used herein, the term “nucleic acid” in its broadest sense, refers to any compound and/or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, a nucleic acid is a compound and/or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester linkage. As will be clear from context, in some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and/or nucleoside); in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, a "nucleic acid" is or comprises RNA; in some embodiments, a "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, a nucleic acid is, comprises, or consists of one or more "peptide nucleic acids", which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, are considered within the scope of the present disclosure. Alternatively or additionally, in some embodiments, a nucleic
acid has one or more phosphorothioate and/or 5'-N-phosphoramidite linkages rather than phosphodicstcr bonds. In some embodiments, a nucleic acid is, comprises, or consists of one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxy guanosine, and deoxy cytidine). In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2- aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyl adenosine, 5- methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5- bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5 -propynyl-cytidine, C5 -methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8- oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, a nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) as compared with those in natural nucleic acids. In some embodiments, a nucleic acid has a nucleotide sequence that encodes a functional gene product such as an RNA or protein. In some embodiments, a nucleic acid includes one or more introns. In some embodiments, nucleic acids are prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), reproduction in a recombinant cell or system, and chemical synthesis. In some embodiments, a nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 1 10, 120, 130, 140, 150, 160, 170 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues long. In some embodiments, a nucleic acid is partly or wholly single stranded; in some embodiments, a nucleic acid is partly or wholly double stranded. In some embodiments a nucleic acid has a nucleotide sequence comprising at least one element that encodes, or is the complement of a sequence that encodes, a polypeptide. In some embodiments, a nucleic acid has enzymatic activity.
[0190] Pareto Score: As used herein, the term “Pareto score” refers to a measure of a variant’s performance with respect to / as evaluated via one or more scoring metrics, such as various scoring metrics and/or combinations thereof, described herein. These may include, but are not necessarily limited to, scores that evaluate fitness and ability to escape an immune response. In some embodiments, a Pareto score comprises a combination of an immune escape score (e.g., as described herein) and a fitness prior score (e.g., as described herein). In some
embodiments, a Pareto score captures the relative evolutionary advantage of a given strain. In some embodiments, such a Pareto score can be determined as described in the Examples. In some embodiments, a Pareto score is an optimality score, which, for example in some embodiments ranks a variant relative to other sequences, e.g., ones that are observed in a population. A high Pareto score at a given time for a specific lineage indicates that fewer variants have higher scores for fitness prior and immune escape at that time. As a Pareto score in some embodiments is a ranking system, and fitness prior and immune escape scores incorporated therein can change as new data are acquired, the Pareto score for a given variant can change over time. As used herein, in some embodiments, Pareto optimality is defined over a set of lineages. In some embodiments, lineages are Pareto optimal within a set if there are no lineages in the set with higher immune escape and higher fitness prior scores. In some embodiments, a Pareto score is a measure of the degree of Pareto optimality. For example, in some embodiments, lineages with the highest Pareto score are Pareto optimal; and lineages with the second-best Pareto score would be Pareto optimal, if the Pareto optimal lineages were removed from the set, and so on.
[0191] Patient'. As used herein, the term “patient” refers to any organism to which a provided composition is or may be administered, e.g., for experimental, diagnostic, prophylactic, cosmetic, and/or therapeutic purposes. Typical patients include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and/or humans). In some embodiments, a patient is a human. In some embodiments, a patient is suffering from or susceptible to one or more disorders or conditions. In some embodiments, a patient displays one or more symptoms of a disorder or condition. In some embodiments, a patient has been diagnosed with one or more disorders or conditions. In some embodiments, the disorder or condition is or includes a viral infection (e.g., a SARS-CoV-2 infection). In some embodiments, the patient is receiving or has received certain therapy to diagnose and/or to treat a disease, disorder, or condition.
[0192] Peptide'. The term “peptide” as used herein refers to a polypeptide that is typically relatively short, for example having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
[0193] Pharmaceutical composition '. As used herein, the term “pharmaceutical composition” refers to an active agent, formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, active agent is present in a unit dose amount that is appropriate for administration in a therapeutic regimen that shows a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population. In some embodiments, pharmaceutical compositions may be specially formulated for administration in solid or liquid form, including those adapted for the following: oral administration, for example, drenches (aqueous or non-aqueous solutions or suspensions), tablets, e.g., those targeted for buccal, sublingual, and systemic absorption, boluses, powders, granules, pastes for application to the tongue; parenteral administration, for example, by subcutaneous, intramuscular, intravenous or epidural injection as, for example, a sterile solution or suspension, or sustained-release formulation; topical application, for example, as a cream, ointment, or a controlled-release patch or spray applied to the skin, lungs, or oral cavity; intravaginally or intrarectally, for example, as a pessary, cream, or foam; sublingually; ocularly; transdermally; or nasally, pulmonary, and to other mucosal surfaces.
[0194] Pharmaceutically acceptable'. As used herein, the term "pharmaceutically acceptable" applied to the carrier, diluent, or excipient used to formulate a composition as disclosed herein means that the carrier, diluent, or excipient must be compatible with the other ingredients of the composition and not deleterious to the recipient thereof.
[0195] Pharmaceutically acceptable carrier : As used herein, the term “pharmaceutically acceptable carrier” means a pharmaceutically-acceptable material, composition or vehicle, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, involved in carrying or transporting the subject compound from one organ, or portion of the body, to another organ, or portion of the body. Each carrier must be “acceptable” in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials which can serve as pharmaceutically-acceptable carriers include: sugars, such as lactose, glucose and sucrose; starches, such as corn starch and potato starch; cellulose, and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, com oil and soybean
oil; glycols, such as propylene glycol; polyols, such as glycerin, sorhitol, mannitol and polyethylene glycol; esters, such as ethyl oleate and ethyl laurate; agar; buffering agents, such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer’s solution; ethyl alcohol; pH buffered solutions; polyesters, polycarbonates and/or poly anhydrides; and other non-toxic compatible substances employed in pharmaceutical formulations.
[0196] Pharmaceutical grade'. The term “pharmaceutical grade” as used herein refers to standards for chemical and biological drug substances, drug products, dosage forms, compounded preparations, excipients, medical devices, and dietary supplements, established by a recognized national or regional pharmacopeia (e.g., The United States Pharmacopeia and The Formulary (USP-NF)).
[0197] Polypeptide'. As used herein refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that occurs in nature. In some embodiments, a polypeptide has an amino acid sequence that does not occur in nature. In some embodiments, a polypeptide has an amino acid sequence that is engineered in that it is designed and/or produced through action of the hand of man. In some embodiments, a polypeptide may comprise or consist of natural amino acids, non-natural amino acids, or both. In some embodiments, a polypeptide may comprise or consist of only natural amino acids or only nonnatural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L- amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more pendant groups or other modifications, e.g., modifying or attached to one or more amino acid side chains, at the polypeptide’s N-terminus, at the polypeptide’s C-terminus, or any combination thereof. In some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, pegylation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic, and/or may comprise a cyclic portion. In some embodiments, a polypeptide is not cyclic and/or does not comprise any cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or comprise a stapled polypeptide. In some embodiments, the term “polypeptide” may be appended to a name of a
reference polypeptide, activity, or structure; in such instances it is used herein to refer to polypeptides that share the relevant activity or structure and thus can be considered to be members of the same class or family of polypeptides. For each such class, the present specification provides and/or those skilled in the art will be aware of exemplary polypeptides within the class whose amino acid sequences and/or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides for the polypeptide class or family. In some embodiments, a member of a polypeptide class or family shows significant sequence homology or identity with, shares a common sequence motif (e.g., a characteristic sequence element) with, and/or shares a common activity (in some embodiments at a comparable level or within a designated range) with a reference polypeptide of the class; in some embodiments with all polypeptides within the class). For example, in some embodiments, a member polypeptide shows an overall degree of sequence homology or identity with a reference polypeptide that is at least about 30-40%, and is often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more and/or includes at least one region (e.g., a conserved region that may in some embodiments be or comprise a characteristic sequence element) that shows very high sequence identity, often greater than 90% or even 95%, 96%, 97%, 98%, or 99%. Such a conserved region usually encompasses at least 3-4 and often up to 20 or more amino acids; in some embodiments, a conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more contiguous amino acids. In some embodiments, a relevant polypeptide may comprise or consist of a fragment of a parent polypeptide. In some embodiments, a useful polypeptide as may comprise or consist of a plurality of fragments, each of which is found in the same parent polypeptide in a different spatial arrangement relative to one another than is found in the polypeptide of interest (e.g., fragments that are directly linked in the parent may be spatially separated in the polypeptide of interest or vice versa, and/or fragments may be present in a different order in the polypeptide of interest than in the parent), so that the polypeptide of interest is a derivative of its parent polypeptide.
[0198] Prevent or prevention : as used herein when used in connection with the occurrence of a disease, disorder, and/or condition, refers to reducing the risk of developing the disease, disorder and/or condition and/or to delaying onset of one or more characteristics or
symptoms of the disease, disorder or condition. Prevention may be considered complete when onset of a disease, disorder or condition has been delayed for a predefined period of time.
[0199] Ribonucleotide'. As used herein, the term “ribonucleotide” encompasses unmodified ribonucleotides and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G), and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications including, but not limited to, for example, (a) end modifications, e.g., 5' end modifications ( e.g. , phosphorylation, dephosphorylation, conjugation, inverted linkages, etc.), 3' end modifications (e.g., conjugation, inverted linkages, etc.), (b) base modifications, e.g. , replacement with modified bases, stabilizing bases, destabilizing bases, or bases that base pair with an expanded repertoire of partners, or conjugated bases, (c) sugar modifications (e.g., at the 2' position or 4' position) or replacement of the sugar, and (d) intemucleoside linkage modifications, including modification or replacement of the phosphodiester linkages. The term “ribonucleotide” also encompasses ribonucleotide triphosphates including modified and non-modified ribonucleotide triphosphates.
[0200] Ribonucleic acid (RNA): As used herein, the term “RNA” refers to a polymer of ribonucleotides. In some embodiments, an RNA is single stranded. In some embodiments, an RNA is double stranded. In some embodiments, an RNA comprises both single and double stranded portions. In some embodiments, an RNA can comprise a backbone structure as described in the definition of “Nucleic acid I Polynucleotide” above. An RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.), or a messenger RNA (mRNA). In some embodiments where an RNA is an mRNA. In some embodiments where an RNA is an mRNA, a RNA typically comprises at its 3’ end a poly(A) region. In some embodiments where an RNA is an mRNA, an RNA typically comprises at its 5’ end an art-recognized cap structure, e.g., for recognizing and attachment of an mRNA to a ribosome to initiate translation. In some embodiments, an RNA is a synthetic RNA. Synthetic RNAs include RNAs that are synthesized in vitro (e.g., by enzymatic synthesis methods and/or by chemical synthesis methods). In some embodiments, an RNA is a single-stranded RNA. In some embodiments, a single- stranded RNA may comprise self-complementary elements and/or may establish a secondary and/or tertiary structure. One of ordinary skill in the art will understand that when a single-stranded RNA is
referred to as “encoding,” it can mean that it comprises a nucleic acid sequence that itself encodes or that it comprises a complement of the nucleic acid sequence that encodes. In some embodiments, a single- stranded RNA can be a self-amplifying RNA (also known as selfreplicating RNA).
[0201] Semantic Change'. As used herein, the term “semantic change” refers to a measure of a functional change of a viral polypeptide of a variant (e.g., in some embodiments a viral polypeptide that interacts with a host cell receptor and/or is otherwise involved in host cell entry) with respect to at least one or a plurality of (e.g., at least two, at least three, at least four, or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species and/or known variants, e.g., of the same lineage) from the language model perspective. In some embodiments, a semantic change is a measure of a functional change of a viral polypeptide of a variant (e.g., in some embodiments a viral polypeptide that interacts with a host cell receptor and/or otherwise involved in host cell entry) with respect to a plurality of (e.g., at least two or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species and/or known variants, e.g., of the same lineage) from the language model perspective. In some embodiments, a relevant language model can comprise Transformer-derived embedding differences (e.g., as described herein) with respect to at least one or a plurality of (e.g., at least two, at least three, at least four, or more) reference viral polypeptide(s) (e.g., in some embodiments reference viral polypeptides of wild type species or known variants, e.g., of the same lineage). In some embodiments, a semantic change score can be computed using LI norm. In some embodiments, a sematic change score can be computed using L2 norm (also known as Euclidean norm). In some embodiments, semantic change describes how different a variant is with regard to an underlying statistical model (e.g., in some embodiments a large machine learning model fine-tuned on viral protein sequences observed until a given time point). In some embodiments, semantic change score depends on sequences observed, and thus the semantic change score may change over time, as an underlying model is trained on new variant sequences and/or reference sequences. In some embodiments, a semantic change score is determined for a variant Spike polypeptide from SARS-Co-V-2 as described herein. In various embodiments described herein, a semantic change score can be normalized such that it ranks between 0 and 100%.
[0202] Subject: As used herein, the term “subject” refers an organism, typically a mammal (e.g., a human, in some embodiments including prenatal human forms). In some embodiments, a subject is suffering from a relevant disease, disorder or condition. In some embodiments, a subject is susceptible to a disease, disorder, or condition. In some embodiments, a subject displays one or more symptoms or characteristics of a disease, disorder or condition. In some embodiments, a subject does not display any symptom or characteristic of a disease, disorder, or condition. In some embodiments, a subject is someone with one or more features characteristic of susceptibility to or risk of a disease, disorder, or condition. In some embodiments, a subject is a patient. In some embodiments, a subject is an individual to whom diagnosis and/or therapy is and/or has been administered.
[0203] Substantially. As used herein, the term “substantially” refers to the qualitative condition of exhibiting total or near-total extent or degree of a characteristic or property of interest. One of ordinary skill in the biological arts will understand that biological and chemical phenomena rarely, if ever, go to completion and/or proceed to completeness or achieve or avoid an absolute result. The term “substantially” is therefore used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena.
[0204] Variant: As used herein, in the context of molecules, e.g., nucleic acids, proteins, or small molecules, the term “variant” refers to a molecule that shows significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., in the presence or absence or in the level of one or more chemical moieties as compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. In general, whether a particular molecule is properly considered to be a “variant” of a reference molecule is based on its degree of structural identity with the reference molecule. As will be appreciated by those skilled in the art, any biological or chemical reference molecule has certain characteristic structural elements. A variant, by definition, is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and/or one or more differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalently components of the polypeptide or nucleic acid (e.g.,
that are attached to the polypeptide or nucleic acid backbone). In some embodiments, a variant polypeptide or nucleic acid shows an overall sequence identity with a reference polypeptide or nucleic acid that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of the reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid shows a reduced level of one or more biological activities as compared to the reference polypeptide or nucleic acid. In some embodiments, a polypeptide or nucleic acid of interest is considered to be a “variant” of a reference polypeptide or nucleic acid if it has an amino acid or nucleotide sequence that is identical to that of the reference but for a small number of sequence alterations at particular positions. Typically, fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of the residues in a variant are substituted, inserted, or deleted, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residues as compared to a reference. Often, a variant polypeptide or nucleic acid comprises a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) number of substituted, inserted, or deleted, functional residues (i.e., residues that participate in a particular biological activity) relative to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises not more than about 5, about 4, about 3, about 2, or about 1 addition or deletion, and, in some embodiments, comprises no additions or deletions, as compared to the reference. In some embodiments, a variant polypeptide or nucleic acid comprises fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and commonly fewer than about 5, about 4, about 3, or about 2 additions or deletions as compared to the reference. In some embodiments, a reference polypeptide or nucleic acid is one found in nature.
[0205] Vaccination'. As used herein, the term “vaccination” refers to the administration of a composition intended to generate an immune response, for example to a disease (c.g., to a viral epitope). In some embodiments, vaccination can be administered before, during, and/or after development of a disease. In some embodiments, vaccination includes multiple administrations, appropriately spaced in time, of a vaccinating composition.
[0206] Viral polypeptide receptor binding score: As used herein, the term “viral polypeptide receptor binding score” refers to a measure of binding affinity between a viral polypeptide that plays a role in host recognition and/or host cell entry, and a corresponding host protein with which the viral polypeptide interacts to recognize and/or enter a host cell. In some embodiments, a viral polypeptide receptor binding score is determined in silico. In some embodiments, a viral polypeptide receptor binding score can be determined using a conformational sampling algorithm. In some embodiments, a viral polypeptide receptor binding score can be determined using structures that have been optimized using a probabilistic optimization algorithm (for example, in some embodiments a variant of simulated annealing, aiming to overcome local energy barriers and follow a kinetically accessible path toward an attainable deep energy minimum with respect to a knowledge-based, protein-oriented potential). In some embodiments, a viral polypeptide receptor binding score can be calculated using the change in solvent accessible surface area (SAS A) of a viral polypeptide in a complexed state (e.g., a bound state) and a non-complexed state (e.g., a non-bound state). In some embodiments, a viral polypeptide receptor binding score can be determined by calculating the change in energy of the complexed (e.g., bound) and non-complexed (e.g., non-bound) structures of a viral polypeptide and its cognate host receptor. In some embodiments, change in binding energy can be estimated by differences in Gibbs free energy between bound and unbound states. In various embodiments described herein, a viral polypeptide receptor binding score can be normalized such that it ranks between 0 and 100%. In some embodiments, a viral polypeptide receptor binding score can be calculated in silico, e.g., by calculating the change in Gibbs Free Energy, or the change in solvent accessible surface area in the bound and unbound states. In some embodiments, a viral polypeptide receptor binding score can be calculated using in vitro binding data (e.g., using a dissociation constant, KD, or an association rate, kOn). In some embodiments, such in vitro binding data can be determined methods known in the art, including, e.g., but not limited to biolayer interferometry (BLI) and/or surface plasmon resonance (SPR).
[0207] ACE2 Binding Score: As used herein, the term “ACE2 binding score” is a viral polypeptide receptor binding score (as described herein), wherein the viral polypeptide receptor is angiotensin-converting enzyme 2 (ACE2). An “ACE2 binding score” is a measure of binding affinity between an S protein of a coronavirus (e.g., SARS-CoV-2) or an immunogenic fragment of the S protein (e.g., the RBD domain) and the ACE2 protein. In some embodiments, an ACE2 binding score can be calculated in silico, e.g., by calculating the change in Gibbs Free Energy, or the change in solvent accessible surface area in the bound and unbound states. In some embodiments, an ACE2 binding score can be calculated using in vitro binding data (e.g., using a dissociation constant, KD, or an association rate, kon). In some embodiments, such in vitro binding data can be determined methods known in the art, including, e.g., but not limited to biolayer interferometry (BLI) and/or surface plasmon resonance (SPR).
[0208] Wild-type: As used herein, the term “wild-type” has its art-understood meaning that refers to an entity having a structure and/or activity as found in nature in a “normal” (as contrasted with mutant, diseased, altered, etc.) state or context. Those of ordinary skill in the art will appreciate that wild-type genes and polypeptides often exist in multiple different forms (e.g., alleles). In some embodiments, in the context of SARS-CoV-2, “wild-type” refers to the Wuhan variant.
DETAILED DESCRIPTION
[0209] It is contemplated that systems, architectures, devices, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and/or modification of the systems, architectures, devices, methods, and processes described herein may be performed, as contemplated by this description.
[0210] Throughout the description, where articles, devices, systems, and architectures are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are articles, devices, systems, and architectures of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and
methods according to the present invention that consist essentially of, or consist of, the recited processing steps.
[0211] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0212] The mention herein of any publication, for example, in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not meant as a description of prior art with respect to any claim.
[0213] Documents arc incorporated herein by reference as noted. Where there is any discrepancy in the meaning of a particular term, the meaning provided in the Definition section above is controlling.
[0214] Headers are provided for the convenience of the reader - the presence and/or placement of a header is not intended to limit the scope of the subject matter described herein.
[0215] Among other things, the present disclosure provides systems and methods for design of engineered antigens with tailored immunological features. In certain embodiments, engineered antigens created via the technologies described herein are computer engineered versions of reference antigens, designed to interact with a subject’s immune system in particular, e.g., desired, ways.
[0216] For example, in certain embodiments, a reference antigen may be a naturally occurring protein, such as a variant of a particular viral protein, or portion thereof. Antigen engineering techniques described in further detail herein may be used to design a custom, tailored version of such reference antigens to encourage production of new antibodies and reduce likelihood of triggering and/or extent of a memory immune response that may result in generation of antibodies from memory B-cells that result prior exposure to a variant of reference antigen. New antibodies produced in this manner may, accordingly, be tailored to, e.g., particular versions of arisen epitopes (e.g., mutations) that are present in a reference antigens, in contrast to antibodies generated from a memory response, which may effectively target similar
epitopes on prior variants, but be evaded by new arisen epitopes on a particular reference antigen.
[0217] Without wishing to be bound to any particular theory, we propose that it may be desirable to vaccinate with antigen polypeptides designed to encourage immune responses, and particularly antibody responses (e.g., neutralizing antibody responses), to arisen epitopes in valiant polypeptides. Among other things, the present disclosure provides an insight that it may be particularly desirable, especially for circulating infectious diseases (e.g., for which variants can be expected to arise), to encourage immune responses, specifically including antibody responses (e.g., neutralizing responses) to arisen epitopes.
[0218] Among other things, the present disclosure provides technologies for engineering antigens with reduced risk of triggering memory immune response(s) as described herein by identifying and mutating (e.g., introducing amino acid modifications into) portions of (e.g., sequences, particular sets of amino acid sites that are determined to be members of an identified surface region) reference antigens that are determined to be likely to include one or more shared epitope(s) and/or (e.g., contiguous) conserved surfaces. As described in further detail herein, particular approaches described herein include recognition and establishment of various criteria pertaining to particular distributions of amino acid modifications, sources and rules for generating amino acid modifications, and approaches for evaluating prospective performance of engineered antigen designs that aim to sufficiently disrupt memory triggering portions of an input reference antigen while at a same time preserve features such as stability, 3D structure (e.g., folding), and the like.
[0219] The present disclosure exemplifies certain aspects of provided technologies via approaches for designing engineered versions of a recently evolved XBB SARS-Cov2 Spike protein variant. While certain examples provided herein are described with reference to particular proteins and viral variants, one of skill in the art, having read the present disclosure, will appreciate that approaches described herein may be applied and adopted, e.g., to other pathogens (e.g., virus types), proteins, sub-regions, and the like, to encourage production of desired immune response.
A. Natural Evolution of Infectious Agents
[0220] Immune imprinting is a phenomenon whereby initial exposure to a particular antigen can limit (e.g., subsequent) development of immune responses against epitopes that are unique to new variants of the antigen. In particular, when exposed to a new, previously unencountered infectious agent, such as a virus, immune systems respond, among other things, by generating antibodies that bind to and neutralize portions of antigen(s) of the agent, in a highly specific fashion. Subsequently, the immune system retains a ‘memory’ of the antigen(s), along with the ability to produce the particular antibodies that target it, in the form of memory B and T cells.
[0221] On the one hand, following an initial exposure to a particular agent, this immune memory allows the body to rapidly recognize and defend against it when it is subsequently encountered. On the other hand, processes such as natural mutation and evolution can give rise to variants of the agent that are similar enough to the originally encountered strain to be recognized and trigger a memory response, prompting production of antibodies that were generated to defend against the original strain, rather than being expressly tailored to the new variant. If the new variant includes sufficient mutations in key regions (e.g., implicated in host cell infection, viral replication, etc.) targeted by these antibodies, the efficacy of this memory response can be reduced. Accordingly, immune imprinting can be particularly concerning for pathogens having a high concentration of mutations at neutralization sensitive epitopes.
[0222] In the context of viral infection and vaccination, this immune imprinting phenomenon can lead to increased rates or reinfection by mutated variants and limit efficacy of vaccination in individuals after their initial exposure to an earlier strain (e.g., whether due to natural infection or an earlier vaccine dose). Immune imprinting can, according, be especially problematic for vaccination against viruses that have higher mutation rates, such as RNA viruses. These include, without limitation, influenza, coronavirus (e.g., severe acute respiratory syndrome-related coronavirus), human immunodeficiency virus (HIV), Respiratory syncytial vims (RSV), and the like.
[0223] For example, in the context of the recent SARS-CoV 2 pandemic, mutation of circulating virus has given rise to tens of thousands of viral variants, several of which - such as Omicron and recently emergent XBB (e.g., XBB.1.5) - are characterized by their immune escape
potential. In particular, these variants include several mutations that allow them to evade existing (e.g., memory) immune responses that individuals have developed as a result of prior exposure - either through vaccination and/or natural infection - to previous strains, such as the original Wild-Type (WT) Wuhan variant.
[0224] Among other things, the present disclosure appreciates that, for example, XBB includes mutations across several epitopes. Many of these mutations as shared with other, earlier variants, while a subset are unique to XBB. When these shared, pre-existing, epitopes trigger memory immune responses, antibodies that are tailored to earlier variants to which an individual was exposed are produced in lieu of new antibodies that are expressly designed to neutralize XBB.
[0225] Accordingly, without wishing to be bound to any particular theory, systems and methods described herein include approaches for designing engineered antigens that reduce an extent to which a memory response is triggered by shared epitopes of a new antigen and encourage the immune system to generate novel responses to particular target epitopes that are unique to the new antigen.
B, Engineered Antigens
[0226] Turning to FIG. 1, among other things, the present disclosure provides systems and methods for in-silico design of engineered antigens that are designed to encourage an immune response to particular epitopes of a reference antigen of an infectious agent.
Approaches for engineering antigens arc described herein often with reference to virus and their proteins, but may also be utilized to design engineered antigens based on reference antigens from I associated with other types of infectious agents, such as bacteria, parasites (e.g., malaria), etc. i. Reference Antigens & Computer Representations Thereof
[0227] A reference antigen may be a protein (e.g., of or produced by) the infectious agent and/or a portion thereof. For example, in certain embodiments, a reference antigen is a particular viral protein and/or portion thereof, such as a surface protein. In certain embodiment, a reference antigen is a protein or portion of a protein that has been determined to be utilized by the virus to
infect host cells, for example implicated in binding to particular host cell proteins and/or facilitating fusion with a host cell membrane. In certain embodiments, a reference antigen is or comprises a particular portion - e.g., a sub-unit, domain, etc. of a viral protein.
[0228] For example, in certain embodiments, an infectious agent is or comprises a coronavirus, such as SARS-CoV 2 and a reference antigen thereof is a SARS-CoV 2 Spike protein, or a portion of the SARS-CoV 2 Spike protein. For example, a reference antigen may be an entire Spike protein, or may be a particular portion, such as an /V-tcrminal region (e.g., an N- Terminal Domain (NTD)) or a Receptor Binding Domain (RBD). In certain embodiments, a particular portion is selected to focus on a minimal relevant vaccine antigen, e.g., to facilitate removal of as many conserved epitopes as possible without, e.g., resorting to introducing point mutations (e.g., thereby limiting number of epitopes in which point mutations are to be introduced).
[0229] As described herein, in certain embodiments, antigen engineering technologies of the present disclosure aim to design engineered versions of a reference antigen that, when introduced (e.g., administered) to a subject, encourage their immune system to generate new antibodies that are expressly tailored to particular epitopes of the reference antigen. In certain embodiments, this involves reducing a likelihood and/or an extent to which an engineered version of a reference antigen will trigger a memory immune response, thereby ameliorating certain obstacles immune imprinting phenomena can, as described herein, cause in regards to e.g., vaccination.
[0230] Accordingly, in certain embodiments, systems and methods described herein identify those portions of a particular reference antigen that remain are likely to trigger memory immune response, and generate, in-silico, one or more engineered variants in which these memory triggering sub-region(s) are disrupted.
[0231] FIG. 1 shows an example process 100 for generating engineered antigens according to certain embodiments. In certain embodiments, example process 100 begins via accessing, generating, or otherwise obtaining, a computer representation of at least a portion of a reference antigen of interest, referred to herein as a polypeptide model 102.
[0232] A variety of formats (e.g., data structures) may be used for computer representation of reference antigens such as particular proteins and/or portions thereof. For
example, a polypeptide model 102 be or comprise an amino acid sequence, such as ordered string of letters, each representing a particular amino acid (c.g., FASTA a file). In certain embodiments, a polypeptide model may be or comprise a structural model, such as a 3D model that includes an indication of a location of each amino acid (or atom thereof) of a reference antigen in 3D space. For example, protein data bank (PDB) format files may be used and/or accessed to obtain 3D structural information about particular proteins, for example based on derived crystallographic structures.
[0233] In certain embodiments, a reference antigen may be a protein of a particular target viral valiant, such as a particular target SARS-CoV-2 variant's spike (S) protein or a portion thereof. As used herein, a full-length SARS-CoV-2 S protein comprising a “Wild-Type” or
“Wuhan” sequence has a sequence corresponding to that of the first detected SARS-CoV-2 strain, consisting of 1273 amino acids and having an amino acid sequence according to SEQ ID
NO: 1:
[0234] Unless otherwise indicated, position numberings in a SARS-CoV-2 S protein and/or portion thereof given herein are in relation to the amino acid sequence of SEQ ID NO: 1. One of skill in the art reading the present disclosure will understand and be able to determine corresponding positions in a SARS-CoV-2 S protein variant sequence from locations of positions provided relative to the amino acid sequence of SEQ ID NO: 1 (i.e. , a person of skill in the art provided positions relative to SEQ ID NO: 1, or another variant, will be able to determine corresponding positions in the S protein sequence of another SARS-CoV-2 variant or a fragment thereof). Where a portion of a SARS-CoV-2 S protein is described as having certain mutations, unless otherwise indicated, it should be understood that position numberings of the mutations are given, and identify locations, in relation to the amino acid sequence of SEQ ID NO: 1.
[0235] In specific embodiments, a spike (S) protein described herein can be modified in such a way that the prototypical prefusion conformation is stabilized. Certain mutations that stabilize a prefusion confirmation are known in the art, e.g., as disclosed in WO 2021243122 A2 and Hsieh, Ching-Lin, et al. (“Structure-based design of prefusion- stabilized SARS-CoV-2 spikes,” Science 369.6510 (2020): 1501-1505), the contents of each which are incorporated by reference herein in their entirety. In some embodiments, a SARS-CoV-2 S protein may be stabilized by introducing one or more proline mutations. In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to residues 986 and/or 987 of SEQ ”D NO: 1. In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at one or more positions corresponding to residues 817, 892, 899, and 942 of SEQ ID NO: 1 . In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to each of residues 817, 892, 899, and 942 of SEQ ID NO: 1. In some embodiments, a SARS-CoV-2 S protein comprises a proline substitution at positions corresponding to each of residues 817, 892, 899, 942, 986, and 987 of SEQ ID NO: 1.
[0236] In some embodiments, stabilization of the prototypical prefusion conformation of a SARS-CoV-2 S protein may be obtained by introducing two consecutive proline substitutions at residues 986 and 987. Specifically, spike (S) protein stabilized protein variants are obtained in
a way that the amino acid residue at position 986 is exchanged to proline and the amino acid residue at position 987 is also exchanged to proline. In one embodiment, a SARS-CoV-2 S protein variant wherein the prototypical prefusion conformation is stabilized comprises the amino acid sequence shown in SEQ ID NO: 2:
[0237] Those skilled in the art are aware of various SARS-COV-2 Spike variants, and/or resources that document them. For example, the following strains, their SARS-CoV-2 S protein amino acid sequences and, in particular, modifications thereof compared to wildtype SARS-
CoV-2 S protein amino acid sequence, e.g., as compared to SEQ ID NO: 1, are useful herein.
[0238] B.1.1.7 (“Variant of Concern 202012/01” (VOC-202012/01))
[0239] B.l .1 .7 (“alpha variant”) is a SARS-CoV-2 variant that was first detected in October 2020 in the United Kingdom from a sample taken the previous month, and quickly began to spread by mid-December. It is correlated with a significant increase in the rate of COVID-19 infection; this increase is thought to be at least partly due to a change of N501 Y inside the spike glycoprotein’s receptor-binding domain, which is needed for binding to ACE2 in human cells. B.l.1.7 is defined by 23 mutations: 13 non-synonymous mutations, 4 deletions, and 6 synonymous mutations (i.e., there are 17 mutations that change proteins and six that do not). Spike protein changes in B.l.1.7 include deletion 69-70, deletion 144, N501Y, A570D, D614G, P681H, T716I, S982A, and D1118H.
[0240] B.1.351 (501.V2)
[0241] B.1.351 lineage ( “Beta variant”), colloquially known as South African COVID-
19 variant, has increased transmissibility relative to the original Wuhan strain. The B.1.351 variant is defined by multiple spike protein changes including: L18F, D80A, D215G, deletion 242-244, R246I, K417N, E484K, N501Y, D614G and A701V. There are three mutations of particular interest in the spike region of the B.l.351 genome: K417N, E484K, N501Y.
[0242] B.l.1.298 (Cluster 5)
[0243] B .1.1.298 was discovered in North Jutland, Denmark, and is believed to have been spread from minks to humans via mink farms. Several different mutations in the spike protein of the virus have been confirmed. The specific mutations include deletion 69-70, Y453F, D614G, I692V, M1229I, and optionally S1147L.
[0244] P.l (B.l.1.248)
[0245] Lineage B.l.1.248 (the “gamma variant”), known as the Brazil(ian) variant, is one of the variants of SARS-CoV-2 which has been named P.l lineage. P.l has a number of S- protein modifications (L18F, T20N, P26S, D138Y, R190S, K417T, E484K, N501Y, D614G, H655Y, T1027I, VI 176F) and is similar in certain key RBD positions (K417, E484, N501) to variant B.1.351 from South Africa.
[0246] B.1.427/B.1.429 (CAL.20C)
[0247] Lineage B.1.427/B.1.429 (the “epsilon variant”), also known as CAL.20C, is defined by the following modifications in the S-protein: S 131, W152C, L452R, and D614G, of
which the L452R modification is of particular concern. CDC has listed B.1 ,427/B.1 .429 as a “variant of concern”.
[0248] B.1.525
[0249] B.1.525 ( “eta variant”) carries the same E484K modification as found in the P.l, and B.1.351 variants, and also carries the same AH69/AV70 deletion as found in B.1.1.7, and B.1.1.298. It also carries the modifications D614G, Q677H and F888L.
[0250] B.1.526
[0251] B.1.526 ( “iota variant”) was detected as an emerging lineage of viral isolates in the New York region that shares mutations with previously reported variants. The most common sets of spike mutations in this lineage are L5F, T95I, D253G, E484K, D614G, and A701V.
[0252] B.1.1.529
[0253] B .1.529 (“Omicron variant”) was first detected in South Africa in November 2021. Omicron multiplies around 70 times faster than Delta variants, and quickly became the dominant strain of SARS-CoV-2 worldwide. Since its initial detection, a number of Omicron sublineages have arisen. Listed below are the current Omicron variants of concern, along with certain characteristic mutations associated with the S protein of each. The S protein of BA.4 and BA.5 have the same set of characteristic mutations, which is why the below table has a single row for “BA.4 or BA.5”, and why the present disclosure refers to a “BA.4/5” S protein in some embodiments. Similarly, the S proteins of the BA.4.6 and BF.7 Omicron variants have the same set of characteristic mutations, which is why the below table has a single row for “BA.4.6 or BF.7”).
Table 1A: Certain Omicron Variants of Concern and Their Characteristic Mutations.
[0254] In some embodiments, SARS-CoV-2 S proteins described herein comprise one or more mutations (including, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more) characteristic of a certain Omicron variant (e.g., one or more mutations of an Omicron variant listed in Table 1A, e.g., each of the mutations associated with a given XBB variant in the above Table 1A).
[0255] As noted elsewhere in the present disclosure, in some embodiments a particular’ immunogenic portion (e.g., sub-region) of a full-length coronavirus S protein (e.g., SARS-CoV-2 S protein) may be used as a reference antigen for creation of engineered antigens as described herein.
[0256] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS- CoV-2) S protein lacks certain features that are in the full-length polypeptide (e.g., features that have been shown or predicted to interfere with induction of a naive immune response). For example, in some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks regions that have (i) a low number or density of B cell neutralization epitopes and/or (ii) a high number or density of B cell epitopes not associated with neutralization. For example, in some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a full S2 domain. In some embodiments, a coronavirus (e.g., SARS-CoV-2) S protein lacking a full S2 domain lacks regions of S2 that have (i) a low number or density of B cell epitopes associated with neutralization or (ii) a high number of B cell epitopes not associated with neutralization, but retains other portions of S2. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a full S2 domain, but comprises certain sequences that can improve immunogenicity and/or stability of an immunogenic portion (e.g., in some embodiments, an immunogenic portion lacks a full S2 domain but retains a TM sequence).
[0257] A person of skill in the art reading the present disclosure will be able to identify B cell epitopes in a coronavirus (e.g., SARS-CoV-2) S protein and determine which epitopes arc or are not associated with neutralization. For example, a person of skill in the ail will be aware of numerous studies that have identified such regions using antibody binding studies (e.g., studies characterizing antibodies produced in subjects infected with or vaccinated against SARS-CoV- 2).
[0258] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS- CoV-2) S protein comprises certain regions that have been determined to have a high number or density of neutralization epitopes and optionally a high mutation rate. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises an N-terminal domain (NTD) of the S protein. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) protein comprises a receptor binding domain (RED) of the S protein. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises an SI domain of the S protein.
[0259] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-
CoV-2) S protein comprises an RED and an NTD and omits other features of the SI domain.
[0260] Coronavirus (e.g., SARS-CoV-2) S proteins are well characterized, and a person of skill in the art will be able to determine which portions of an S protein sequence correspond to immunogenic portions discussed herein (e.g., which portions of an S protein sequence correspond to the NTD, the RED, the SI, and the S2 domains). In some embodiments, an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises residues 327 to 528 of SEQ ID NO: 1 or a corresponding region.
[0261] In some embodiments, an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises the amino acid sequence:
or a corresponding region.
[0262] In some embodiments, an RED of a coronavirus (e.g., SARS-CoV-2) S protein comprises the amino acid sequence:
or a corresponding region.
[0263] In some embodiments, an SI domain of a coronavirus (e.g., SARS-CoV-2) S protein comprises amino acids 1 to 678 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant. In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises amino acids 1 to 683 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant. In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises amino acids 1 to 685 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS-CoV-2 variant.
[0264] In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises the amino acid sequence:
[0265] In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises the amino acid sequence:
65
[0266] In some embodiments, an S2 domain of a SARS-CoV-2 S protein comprises amino acids 679 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant. In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises amino acids 684 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant. In some embodiments, an SI domain of a SARS-CoV-2 S protein comprises amino acids 686 to 1273 of SEQ ID NO: 1, or a corresponding region in an S protein of a SARS- CoV-2 variant.
[0267] In some embodiments, compositions described herein deliver an immunogenic portion of an S protein of a SARS-CoV-2 variant. In some embodiments, the variant is a variant of concern (e.g., a variant that has been predicted to and/or has been shown to spread rapidly in a relevant jurisdiction, e.g., as identified by certain public health agencies, e.g., the Center for Disease Control and Prevention (CDC), Public Health England and the CO VID- 19 Genomics UK Consortium for the UK, the Canadian COVID Genomics Network (CanCOGeN), and/or the
World Health Organization (WHO)). In some embodiments, a variant has been predicted to have a highly likelihood of becoming a variant of concern (e.g., using sequence-based algorithms that predict the ability of a variant to escape previously developed immune responses and/or measure the “fitness” of a given variant, such as described, e.g., in WO2022/235847 and WO2022/235853, the contents of each of which are incorporated by reference herein in their entirety).
[0268] In some embodiments, an RBD comprises mutations associated with a variant described herein. A person of skill in the art will be able to identify which portions of a given variant correspond to immunogenic portions described herein.
[0269] In some embodiments, a polypeptide comprises two or more SARS-CoV-2 subdomains (e.g., two or more SI domains or RBDs). In some embodiments, a polypeptide comprises two or more receptor binding domains linked in tandem, e.g., as described in Dai, Lianpan, et al. “A universal design of betacoronavirus vaccines against COVID- 19, MERS, and SARS,” Cell 182.3 (2020): 722-733, and Han, Yuxuan, et al. “mRNA vaccines expressing homo-prototype/Omicron and hetero -chimeric RBD-dimers against SARS-CoV-2,” Cell Research 32.1 1 (2022): 1022-1025, the contents of each of which are incorporated by reference herein in their entirety. In some embodiments, the two or more subdomains are from the same SARS-CoV-2 variant (e.g., a variant described herein). In some embodiments, at least two of the two or more subdomains are from different SARS-CoV-2 variants (e.g., from different variants of concern, different Omicron variants, an Omicron variant and a non-Omicron variant, or a Wuhan strain and an Omicron variant). ii. Hallmark Mutations
[0270] In certain embodiments, approaches described herein leverage information describing unique features of particular reference antigens, such as identifications of hallmark mutations. Hallmark mutations are those mutations that serve to distinguish particular reference antigens (e.g., proteins) from other, similar, antigens. In certain embodiments, for example, a reference antigen is a particular protein or portion thereof of a particular viral variant, and hallmark mutations are those mutations that serve to distinguish the particular viral variant, as well as closely related variants, from other, for example variants. For example,
[0271] For example, in the context of SARS-CoV 2, hallmark mutations may be identified based on one or more classification schemes, such as World Health Organization (WHO) classification(s), GISAID, Pango Lineage, Nextstrain clade, etc.
[0272] For example, in certain embodiments, a reference antigen may be a SARS CoV 2 Spike protein of a particular XBB variant. Hallmark mutations may then be identified as those
mutations that appear in greater than a particular hallmark threshold percentage of all sequences (c.g., within a particular database) classified as members of an XBB lineage, according to the WHO designations. In certain embodiments, a hallmark threshold is fifty-percent or greater. In certain embodiments, a hallmark threshold is sixty percent or greater. In certain embodiments, a hallmark threshold is seventy percent or greater. In certain embodiments, a hallmark threshold is seventy-five percent or greater.
[0273] Exemplary XBB.1.5 hallmark mutations (e.g., mutations relative to a Wuhan strain, e.g., identified via SEQ ID NO: 1 and/or SEQ ID NO. 2) identified as those mutations occurring in greater than 50% of sequences for variants classified as XBB.1.5 by the WHO are as shown in Table IB, below:
Table IB. XBB.1.5 hallmark mutations.
iii. Memory-Triggering Conserved Region(s)
[0274] In certain embodiments, one or more memory-triggering conserved region(s) 122 of polypeptide model 102 are identified 120. Identified (memory triggering) conserved region(s) 122 represent those particular portions of a reference antigen (that is represented by polypeptide model 102) that are determined to be likely to trigger a memory immune response. A memorytriggering conserved region may represent these particular portions of the reference antigen in a variety of manners. For example, a memory -triggering conserved region may be or comprise a list of amino-acid positions, or, additionally or alternatively, be or comprise an identified 3D region on a surface of a 3D polypeptide model.
[0275] Particular portions of a reference antigen may be determined likely to trigger a memory immune response via several approaches, based on various criteria.
[0276] In certain embodiments, memory triggering conserved regions may be identified and determined to be implicated in triggering of memory immune rcsponsc(s) in a variety of manners, using, for example and without limitation, a-priori known or determined data, such as lists of epitopes, structural models, in combination with experimental techniques (e.g., binding assays), and computational methods, individually or in combination.
Conserved Epitopes
[0277] In certain embodiments, memory-triggering portions of a reference antigen determined by identifying a set of conserved epitopes, within the reference antigen, that are likely to trigger a memory immune response. Sets of shared epitopes of a reference antigen may be identified by comparing an initial set of known epitopes with mutations present in the reference antigen. For example, in certain embodiments, a reference antigen is a version of a particular protein, e.g., a particular viral variant’s version of the protein. In certain embodiments, a set of known epitopes comprises epitopes that satisfy particular criteria, such as having been identified as targets of antibody binding (e.g., binding epitopes) and/or neutralizing antibodies (e.g., neutralizing epitopes).
[0278] Known epitopes may be, for example, previously characterized epitopes of the particular virus of which the reference antigen is a variant, such as epitopes to which antibodies have been determined to bind. In certain embodiments, a known epitope is a site to which a neutralizing epitope binds (e.g., a neutralizing B-cell epitope). A set of known epitopes may include binding and/or neutralizing epitopes. Data identifying a set of known epitopes may be or comprise, for each known epitope of the set, a list of amino acid positions. Such data may be obtained from sources such as results of targeted experiments, proprietary datasets, and public information such as literature and publicly available datasets.
[0279] For example, data on epitopes of SARS-CoV 2 Spike protein epitopes can be found via CoV-AbDab, IEDB, etc.
[0280] A known epitope data set may then be screened analyzed to determine which of the known epitopes are mutated in the reference antigen and which are un-mutated and, accordingly, conserved. For example, data identifying a set of known epitopes can be compared
against a list of hallmark mutations of a particular reference antigen. Epitopes of the set of known epitopes that comprise / overlap with one or more hallmark mutations of reference antigen may be identified as mutated epitopes, and other epitopes of the set, e.g., which do not comprise a hallmark mutation, identified as conserved epitopes In certain embodiments, a set of known epitopes may be filtered using additional criteria, such as identifying and removing epitopes that are subsets of other known epitopes or known epitopes that are or are not located within particular regions such as, e.g., within the context of SARS-CoV2, an RBD and/or N- terminal region.
Conserved Surfaces
[0281] In certain embodiments, identifying a memory-triggering conserved region comprises identifying a conserved surface of a polypeptide model representing a particular reference antigen. A conserved surface represents portions of a reference antigen that are available and/or have potential to bind with antibodies created from a memory immune response (e.g., memory B-cells), but, for example, may not necessarily have been previously identified as known epitopes. For example, a conserved surface may represent those portions (e.g., amino acid sites) of the particular reference antigen that are (i) sufficiently surface accessible and (ii) determined to be sufficiently similar to other (e.g., related) antigens and/or unaffected by mutations present in the particular reference antigen such that they could be a binding target of antibodies associated with a memory immune response.
[0282] For example, in certain embodiments, amino acid positions that are determined to be located at a surface (e.g., sufficiently surface accessible) may be identified (e.g., using polypeptide model 102). In certain embodiments, a set of surface amino acids may be compared with hallmark mutations of the reference antigen to, for example, filter out (e.g., remove from the set) amino acid positions that are sites of hallmark mutations and/or within a particular distance [e.g., in 3D space, ‘as the crow flies;’ or traversed over a 3D surface (e.g., a geodesic distance)] of one or more hallmark mutations. The set of remaining amino acid sites can then be used to define a conserved surface. In certain embodiments, a conserved surface may be further refined via evaluation of a spatial distribution of amino acid sites, e.g., to exclude (from the conserved surface) portions that are determined too small for an antibody to bind to. For example, isolated
patches below a particular threshold size (e.g., surface area) may be removed from the conserve surface.
Memory-Triggering Conserved Regions
[0283] In certain embodiments, one or more memory triggering conserved regions are or comprise a set of conserved epitopes. In certain embodiments, one or more memory triggering sub-regions are or comprise a conserved surface. In certain embodiments, both a set of conserved epitopes and a conserved surface are included/identified as one or more memory triggering sub-regions. iv. Disrupting Conserved Region(s)
[0284] Turning again to FIG. 1, following identification of one or more memory triggering conserved region(s) of a reference antigen, systems and methods of the present disclosure aim to disrupt these regions, e.g., in order to lessen or avoid (e.g., reduce a likelihood of and/or extent of) triggering a memory immune response. In particular, in certain embodiments, approaches described herein introduce (e.g., in-silico) amino acid modifications, such as substitutions, insertions, deletions, etc., into at least a portion of one or more identified memory triggering conserved regions of a reference antigen.
Distribution Criteria
[0285] In certain embodiments, approaches for introducing amino acid modifications into conserved regions aim to distribute amino acid modifications throughout conserved regions. For example, in certain embodiments, one or more amino acid modifications are introduced into at least a portion of one or more conserved epitope regions. In certain embodiments, one or more amino acid modifications are introduced into each of one or more conserved epitope region(s), e.g., to disrupt each identified conserved epitope. This may be accomplished by, for example, by separately considering each of at least a portion (e.g., all) of the conserved epitope regions and introducing one or more amino acid modifications into each, e.g., in a step-wise fashion, or,
additionally or alternatively, by repeatedly introducing amino acid modifications at random / according to various statistical selection criteria, until each conserved epitope region is modified via at least one amino acid modification, or other stop criteria is met.
[0286] In certain embodiments, one or more amino acid modifications are introduced into (e.g., across) a conserved surface. In certain embodiments, a plurality of amino acid modifications are generated. In certain embodiments, a quantity and location of amino acid modifications are introduced to create a particular desired spatial distribution of mutations. Design criteria may include, without limitation, a density across a 3D surface representing the conserved region, a minimum and/or maximal separation between amino acid modifications, spatially across a 3D conserved surface. In certain embodiments, desirable distribution of mutations across a conserved surface may be achieved by computing a position spread score that measures an extent to which locations of amino acid modifications are, e.g., evenly, distributed across a conserved surface and, e.g., evaluating designs using, at least in part, the computed position spread score. In certain embodiments, criteria, such as density, minimum/maximum separation, position spread score, etc. is elevated using introduced amino acid modifications (e.g., to ensure introduced amino acid modifications are introduced in a manner to distribute them across a conserved region). In certain embodiments, criteria, such as density, minimum/maximum separation, position spread score, etc. is elevated using introduced amino acid modifications and existing, e.g., hallmark, mutations of a particular antigen such as a viral protein variant (e.g., to ensure introduced amino acid modifications are introduced in a manner to distribute them across a conserved region and also to avoid mutating existing unique epitopes, e.g., which are desired to be retained, e.g., in order to promote production of antibodies directed to (e.g., having an affinity to) these regions).
Generating Amino Acid Modifications
[0287] In certain embodiments, amino acid modifications introduced into memorytriggering sub-regions are selected from a set of allowed mutations. In certain embodiments, a set of allowed mutations comprises, for each of at least a portion of amino acid positions within a particular sequence, a set of one or more mutations having been identified as allowable.
Allowable mutations may be identified using sequence data (e.g., nucleotide and/or amino acid
sequences) of antigens that are related to the reference antigen. For example, wherein a reference antigen is a particular protein of a viral variant, such as a SARS-CoV 2 spike protein of a particular SARS-CoV 2 variant (e.g., an Omicron spike protein, an XBB spike protein, etc.), allowable mutations may be determined by first selecting a set of related lineage(s) and obtaining, for example from proprietary sequencing data and/or public data (e.g., GSAID) sequences of versions of the particular protein of for various viral variants classified as belonging to the selected set of related lineages. Mutations occurring within each of the related variants of the particular protein can be identified and included within a set of allowed mutations. In certain embodiments, mutations to include in the set of allowed mutations are filtered so as to include only those mutations observed at a sufficiently high rate or frequency, for example at or above a particular threshold frequency. For example, in certain embodiments, only those mutations observed within at least particular threshold percentage of the related variant sequences are included in set of allowed mutations.
[0288] In certain embodiments, related variant sequences may be or comprise sequences of proteins within a similar family and/or that are known to perform a similar function to reference antigen. For example, reference antigen may be a particular SARS-Cov2 protein, such as a Spike protein. In certain embodiments, reference sequences may be other coronavirus Spike proteins, or subsets thereof, such as Spike proteins of related SARS virus (e.g., those that are known to infect humans; e.g., SARS-Cov 1, MERS), or known to exist I originate from particular regions. In this manner, amino acid modifications may, in certain embodiments, be selected based on and/or to utilize changes that are present in other related polypeptides.
[0289] For example, in certain embodiments, reference sequences may be or comprise sequences of variants of reference antigen, such as predecessors or members of other branches along an evolutionary tree. Reference sequences may be limited to variants having a particular level of prevalence and/or immune escape. In certain embodiments, reference sequences may be obtained from one or more public and/or proprietary databases, such as GSAID.
[0290] In certain embodiments, one or more additional criteria are used to select amino acid modifications that provide a sufficient level of disruption and/or nonetheless preserve polypeptide chain stability. For example, in certain embodiments, amino acid modifications that substitute an existing amino acid with certain comparable / similar amino acids are excluded,
e.g., as not providing sufficient disruption. For example, in certain embodiments, amino acid modifications that disrupt cysteine bridges arc disallowed and/or penalized. In certain embodiments, amino acid modifications that cause a charge flip and/or produce a large change in charge are disallowed and/or penalized. Such similarity I dis- similarity criteria and structure preservation criteria as described herein may, for example, be implemented via (e.g., encoded) rules, such as in conditional logic, lookup tables, etc.
[0291] FIG. 2 shows an example process, whereby sequences of related antigens are accessed/obtained and analyzed to identify and generate a set of allowed mutations which can be utilized for generating amino acid modification into, and disrupting, conserved sub-regions of reference antigens. In certain embodiments, steps of selecting and inserting, and optionally, filtering, amino acid modifications may be repeated, e.g., iteratively, to generate, based on a single initial conserved region, multiple candidate engineered variants.
Scoring and Selecting Candidate Variants
[0292] In certain embodiments, as conserved sub-region(s) are disrupted via introduction of amino acid modifications, resultant in-silico engineered variants of a reference antigen may be scored, e.g., to assess its viability, ability to infect host cells, relative similarity / dissimilarity to existing variants. For example, in certain embodiments, as illustrated in FIG. 3, one or more candidate disrupted polypeptide models 342a, 342b...342n are generated 340 from (e.g., by identifying 320 and disrupting 330 conserved region(s)) an initial polypeptide model 302 that represents a particular reference antigen. Each candidate disrupted polypeptide model corresponds to the initial polypeptide model, but with its conserved sub-region disrupted via introduction of one or more amino acid modifications.
[0293] Candidate disrupted polypeptide models may evaluated and scored 340 in a variety of manners, for example via one or more approaches described in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the contents of each of which are incorporated herein by reference in their entirety.
[0294] For example, in certain embodiments a candidate polypeptide model may be scored using a machine learning model, such as a language model. In certain embodiments, a language model is or has been trained to generate a predicted probability for each of one or more amino acids at each of one or more locations of an input sequence. Once trained, a language model may then be used to predict an overall likelihood, such as a log-likelihood, a conditional log likelihood, etc., of a particular variant represented via a particular input amino acid sequence.
[0295] In certain embodiments, language models trained in this fashion may also be used to determine an embedding vector representation of an input amino acid sequence. Embedding vector representations of amino acid sequences are internal representations generated by a machine learning model and can be extracted as output from one or more hidden layers of a machine learning model. Extracted embedding layer representations of variants may be used themselves and/or used to generate a characteristic vector representing a particular variant.
[0296] For example, in certain embodiments, a polypeptide sequence is provided as input to a machine learning model which, in turn, generates an internal representation - i.e., an embedding - of the input polypeptide sequence, which may be represent an input sequence as a high dimensional matrix or vector (e.g., a numerical matrix or vector). For example, in certain embodiments, machine learning models, such as certain language models described in further detail herein, may receive an amino acid sequence as input and generate, an internal (e.g., embedding) representation. In certain embodiments, an embedding is or comprises a vector (e.g., zi), e.g., of length D, where D is a dimensionality of the vector, for each amino acid position in a sequence, such that an initial embedding corresponding directly to an internal representation output by a particular layer (e.g., a final embedding layer of a recurrent neural network, e.g., a feature map from a final transformer layer of a transformer-based model) may be a matrix of size n x D [e.g., or (n + 1) x D], where n is a length of an input amino acid sequence and D is a number of dimensions [e.g., in certain embodiments, an additional, class token may be used and appended to a network, such that a matrix is of a size (n + 1) x D], based on the particular machine learning model used. In certain embodiments, a characteristic vector may then be determined as a mean over at least a portion of sequence positions {e.g., all sequence positions, [e.g., but excluding a first (e.g., class token) position, not representing or corresponding to an amino acid in the sequence] } to, for example determine a characteristic
vector z, for an amino acid sequence. Further details of embedding vectors and approaches for determining characteristic vectors based thereon are provided in in PCT publications WO 2022/235847 and WO/2022/235853, the contents of each of which are incorporated by reference herein in their entirety.
[0297] In this manner, in certain embodiments, a machine learning model may be used to generate, for each of a plurality of viral variant polypeptide sequences, a corresponding characteristic vector (e.g., based on an embedding of machine learning model). This approach can be used to associate each variant sequence with a location in a higher dimensional characteristic vector / embedding space.
[0298] In certain embodiments, characteristic vectors and/or embeddings representing particular sequences (e.g., viral variants) may be compared with each other and/or with particular versions, such as a wild-type version of a protein and/or other particular variants of interest. In certain embodiments, a distance (e.g., an LI distance, an L2 distance, etc.) in embedding space may be computed, and used to evaluate a similarity and/or dissimilarity between a particular candidate variant and one or more other variants for comparison. In certain embodiments, distances in embedding space may be referred to as a semantic change score, and are believed to be indicative, at least in part, of potential for a particular variant to be recognized by a host system previously exposed to one or more reference antigens.
[0299] Machine learning models used to generate characteristic vectors may utilize and implement a variety of machine learning techniques. For example, a machine learning model may be a deep learning model (e.g., an artificial neural network with one or more, e.g., a plurality of, hidden layers), such as a language or large language model (LLM). In certain embodiments, a machine learning model is or comprises one or more recurrent models, such long short-term memories (LSTMs), implemented alone or in combination, e.g., as in a bi-directional LSTM (bi-LSTM). In certain embodiments, a machine learning model may be or comprise one or more transformer models. Examples of machine learning models include, without limitation, evolutionary scale models (ESM), bidirectional encoder representations from transformers (BERT), and the like.
[0300] A machine learning model may be trained, for example on protein sequence data available from public and/or private repositories, such as UniRef (see, e.g., Suzek et al. “UniRef:
comprehensive and non-redundant UniProt reference clusters” 2007), the Virus Pathogen Resource (ViPR) database by the National Centre for Biotechnology Information, and the like. Training may be accomplished by providing a machine learning model with input sequences that are incomplete and/or partially masked and tasking them to predict the omitted and/or masked data. In this manner, machine learning models may be trained in an unsupervised fashion. Additional details of machine learning models and approaches for training them, such as particular techniques for training recurrent and/or transformer models, may be found, for example, in PCT publications WO 2022/235847 and WO/2022/235853, the contents of each of which are incorporated by reference herein in their entirety.
[0301] In certain embodiments, structural modeling techniques may be used to score candidate polypeptide models. For example, in certain embodiments, a viral polypeptide receptor binding score, such as an ACE2 binding score, may be determined for one or more candidate synthetic variants. In certain embodiments, an epitope alteration score may be determined for one or more candidate variants.
[0302] In certain embodiments, a quantity of mutations may be evaluated and used as a score, e.g., to select candidate polypeptide models having particular ranges in terms of quantity of mutations. In certain embodiments, mutation co-occurrence scores may be used, to evaluate whether artificially introduced mutations in particular candidate variants are match cooccurrences observed in real-world, naturally evolved, variants.
[0303] In this manner, each candidate may be associated with and characterized by a set of performance scores 344a, 344b...344n.
[0304] In certain embodiments, one or more candidate polypeptide models may be selected 360, e.g., on the basis of such performance scores, as a representation of an engineered antigen to be used, e.g., as an immunogenic composition. In certain embodiments, a set of two or more scores as described herein can be combined, for example as a (e.g., weighted) combined score. In certain embodiments, a set of two or more scores as described herein can be combined to identify sets of Pareto fronts and/or Pareto optimal solutions. In certain embodiments, a Pareto score, based on a combination of two more scores, may be determined, for example as described in in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and
PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the contents of each of which arc incorporated herein by reference in their entirety.
V. Processing Engineered Antigens
[0305] In certain embodiments, disrupted polypeptide models created in-silico via systems and methods described herein may be stored and/or provided for, e.g., display and/or further processing. For example, in certain embodiments, a corresponding RNA sequence may be generated using the disrupted polypeptide model, the corresponding RNA sequence encoding the engineered antigen represented by the disrupted polypeptide model. In certain embodiments, an engineered antigen based on the disrupted polypeptide model may be manufactured and its biological activity assessed, e.g., in-vitro. In this manner, for example, a plurality of initial candidate engineered antigens may be created via in-silico design techniques described herein and subsequently manufactured and screened in-vitro.
[0306] For example, in certain embodiments, assays may be used to check for expression and folding. For example, engineered variants designed via one or more approaches described herein may be produced and tested for viability in terms of expression and folding. In certain embodiments, one or more antibodies or binding agents may be used to analyze for intracellular and surface expression of an encoded antigen. In certain embodiments, for example for testing engineered SARS-CoV 2 variants, an ACE2 binding agent may be used. In certain embodiments, a panel of reference (e.g., known) neutralizing antibodies may be used to evaluate an ability of an engineered construct to escape existing neutralizing antibodies. In certain embodiments, a reference panel of neutralizing antibodies comprises one or more antibodies that bind to distinct epitope classes (e.g., A, B, C, D, E, F). In certain embodiments, analysis may be extended to immune serum analysis from vaccinated/convalescent individuals. In certain embodiments, a pseudovirus assay is performed to check for loss of nAb titers.
[0307] In certain embodiments, one or more (e.g., multiple) testing procedures are used and arranged in a decision tree I hierarchical fashion, for example as described in Example 8. For example, in certain embodiments, all or a subset of the steps may be used, such that at each step / test, a construct is either retained, and passed along to a next step, or discarded, so as to
progressively filter engineered construct designs until a subset that passes each step is obtained.
Such steps may include any of the following:
1. Expression and folding check
2. Abrogated/reduced binding of a panel of reference antibodies;
3. Abrogated/reduced binding of complex immune serum;
4. Abrogated neutralization of respective pseudovirus;
5. Dedicated immunogenicity studies.
An example approach used for engineered SARS-CoV-2 antigen designs is described in further detail in Example 8.
[0308] In certain embodiments, results of assays as described herein may be combined with in-silico design procedures, e.g., in an iterative fashion, for example to inform and/or add constraints to subsequent rounds of in-silico design.
C. Delivery of Engineered Antigens
[0309] Those skilled in the art will appreciate that effective vaccination can be achieved through delivery of engineered antigens to a subject so that the antigen is exposed to B cells of the subject.
[0310] Engineered antigens of the present disclosure may comprise a full-length viral protein and/or a particular portion thereof, such as an RBD that has been altered, e.g., to disrupt conserved regions, according to various embodiments described herein (e.g., in section B, above, and/or various examples described below). Constructs comprising engineered antigens of the present disclosure may combine engineered viral proteins and/or portions thereof with one or more additional elements, such as, for example, one or more of the following: secretory signal(s), linkers, multimerization regions (e.g., fibritin domains), and membrane associated moiety(ies) (e.g., transmembrane regions).
[0311] In some embodiments, delivery of engineered antigens and/or constructs comprising them can be achieved by administration of such antigen(s) (e.g., of a polypeptide antigen). In some embodiments, such delivery can be achieved by administration of a
composition from which the antigen is subsequently generated (e.g., in and/or by the recipient). In some embodiments, delivery of a polypeptide antigen is achieved through administration of a composition comprising a polynucleotide encoding the polypeptide antigen. In some such embodiments, such polynucleotide may be or comprise DNA; in some such embodiments, such polynucleotide may be or comprise RNA. Those skilled in the art, furthermore, will be aware that non-natural residues, linkages, and/or other elements, are often utilized in therapeutic nucleic acids administered to subjects. Various example polypeptide and/or nucleic acid sequences suitable for providing engineered antigens of the present disclosure are described in Examples provided herein, for example Examples 10 and 11.
[0312] Of particular interest in certain embodiments of the present disclosure is administration of a composition comprising an RNA encoding an engineered antigen as described herein.
[0313] In many embodiments, provided pharmaceutical compositions (e.g., immunogenic compositions, e.g., vaccines) deliver antigens as described herein (e.g., engineered antigens) by delivering a nucleic acid construct, e.g., in many embodiments, an RNA that encodes one or more engineered antigens as described herein and is expressed in the subject upon administration of the pharmaceutical composition (e.g., immunogenic composition, e.g., vaccine).
[0314] Among other things, the present disclosure encompasses the recognition that administration of nucleic acid, and particularly of RNA to achieve delivery (e.g., by expression) of encoded antigen can provide a variety of benefits relative to other strategies for immunizing against infections, such as a SARS-CoV-2 infection.
[0315] Among other things, the present disclosure provides an insight that RNA may be particularly useful and/or effective as an active agent in pharmaceutical compositions (e.g., immunogenic compositions, e.g., SARS-CoV-2 vaccines) for a variety of reasons including specifically that RNA can have intrinsic adjuvanticity. As noted herein, ability to induce very high antibody titers to SARS-CoV-2 proteins, e.g., particularly SARS-CoV-2 antigens associated with a variant of concern with high immune escape potential.
[0316] Still further, experience with SARS-CoV-2 vaccines has demonstrated that RNA actives, can also elicit significant and diverse T cell responses which, particularly when
combined with strong antibody response, represents a combination of immune characteristics thought to potentially maximize the probability of protection.
[0317] Provided polyribonucleotides may be delivered for therapeutic applications described herein using any appropriate methods known in the art, including, e.g., delivery as naked RNAs, or delivery mediated by viral and/or non-viral vectors, polymer-based vectors, lipid compositions, nanoparticles (e.g., lipid nanoparticles, polymeric nanoparticles, lipidpolymer hybrid nanoparticles, etc.), and/or peptide-based vectors. See, e.g., Wadhwa et al. “Opportunities and Challenges in the Delivery of mRNA-Based Vaccines” Pharmaceutics (2020) 102 (27 pages), the content of which is incorporated herein by reference, for information on various approaches that may be useful for delivery polyribonucleotides described herein.
[0318] In some embodiments, one or more polyribonucleotides can be formulated with lipid nanoparticles for delivery (e.g., administration).
[0319] In some embodiments, lipid nanoparticles can be designed to protect polyribonucleotides from extracellular RNases and/or engineered for systemic delivery of the RNA to target cells. In some embodiments, such lipid nanoparticles may be particularly useful to deliver polyribonucleotides when polyribonucleotides are intravenously or intramuscularly administered to a subject.
[0320] Additional detail regarding approaches suitable for use in delivery of engineered antigens described herein may be found in in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 the content of which is incorporated herein by reference in its entirety.
D. Sub jects and Indications
[0321] In certain embodiments, systems and methods described herein may be used to design engineered antigens suitable for eliciting improved immune responses. In certain embodiments, approaches described herein may be used to create engineered antigens for use in vaccination of particular subjects, such as particular individuals and/or particular populations. For example, in certain embodiments, technologies for antigen engineering described herein may
be used to engineer antigens for delivery to subjects having been previously exposed to a particular variant (e.g., a prior naturally circulating variant) of the reference antigen. For example, information pertaining to e.g., particular epitopes of the particular variant to which the subject was previously exposed may, in certain embodiments, be used to identify particular conserved epitopes to disrupt. Such approaches may be used to reduce memory response, and encourage naive immune respond, for subjects having been exposed previously via infection and/or prior vaccination.
[0322] In certain embodiments, engineered antigens may be tailored (e.g., expressly created, or selected from a set of pre-existing options) for a subject whose memory B cells have been assessed. In certain embodiments, methods of the present disclosure include administering a composition that delivers an engineered antigen as described herein to a subject whose memory B cells have been assessed. For example, in certain embodiments, approaches described herein may be used to create an engineered antigen modified memory epitope(s) target by the subject’s memory B cells (at least when non-neutralizing). In certain embodiments, such an engineered antigen may be administered to a subject.
E. Certain Example Applications
[0323] Among other things, in certain embodiments, approaches for in-silico design of engineered antigens described herein may be applied toward design of immunogenic compositions. Such immunogenic compounds may have, for example, improved performance and/or be particularly tailored to individual subjects and/or populations, e.g., in view of particular populations of immune memory responses (e.g., depending on antigens to which various individuals and/or populations groups were first exposed).
[0324] Accordingly, in certain embodiments, approaches described herein may be used to create improved immunogenic compositions.
[0325] Technologies described herein, including various approaches described with particularity in the Examples with respect to SARS-CoV 2 antigen engineering, may be applied to a variety of types of antigens and pathogens.
[0326] In some embodiments, an infectious agent is a virus, a bacteria, or a eukaryotic cell (c.g., a plasmodium).
[0327] In some embodiments, an infectious agent is a respiratory virus. In some embodiments, an infectious agent is an RNA virus. In some embodiments, an infectious agent is a coronavirus (e.g., MERS, SARS, or SARS-CoV-2). In some embodiments, an infectious agent is HIV. In some embodiments, an infectious agent is HSV (e.g., HSV-1 or HSV-2). In some embodiments, an infectious agent is RSV. In some embodiments, an infectious agent is a norovirus. In some embodiments, an infectious agent is an influenza virus. In some embodiments, an infectious agent is P. falciparum. In some embodiments, an infectious agent is an Orthopoxvirus (e.g., Monkeypox).
[0328] In some embodiments, an infectious agent is a bacterium. In some embodiments, the bacterium is Mycobacterium. In some embodiments, the bacterium is selected from Haemophilus influenzae, Chlamydophila pneumoniae, Mycoplasma pneumonia, Staphylococcus aureus, Moraxella catarrhalis, Legionella pneumophila, and Streptococcus pneumonia. In some embodiments, the bacterium is Streptococcus pneumonia.
[0329] In some embodiments, an infectious agent is an RNA virus. Compositions provided herein may provide a particular advantage in providing an immune response against RNA viruses, which have a relatively high mutation rate (high relative to other infectious agents).
[0330] In some embodiments, an infectious agent comprises a large number of strains, variants, or lineages. In some embodiments, an infectious agent has a relatively high mutation rate (e.g., relative to other infectious agents).
[0331] In some embodiments, an infectious agent is prone to immune escape.
[0332] In some embodiments, an infectious agent is one for which seasonal, variant- adapted booster shots are regularly provided.
In some embodiments, an infectious agent antigen is solvent exposed on the surface of the infectious agent. In some embodiments, an infectious agent antigen is a glyocoprotein. In some embodiments, an infectious agent antigen is involved in host cell recognition. In some embodiments, an infectious agent antigen is involved in host cell entry. In some embodiments,
an infectious agent antigen comprises one or more B cell epitopes (e.g., one or more neutralization epitopes).
F. Examples i. Example 1: XBB.1.5 Conserved Epitopes and Hallmark Mutations
[0333] This example describes conserved regions and hallmark mutations of XBB.1.5, and other (e.g., ACE2) regions of interest.
[0334] In order to encourage novel B-cell responses to the XBB.l .5 variant, the design approach in this example began by collecting 1,004 binding and neutralizing B-ccll epitopes from CoV-AbDab and IEDB. These epitopes were compared against XBB.1.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.1.5. This set of conserved epitopes was further curated via manual inspection, including removal of epitopes that were not located on an RBD surface and/or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions in the Spike protein RBD. Table 2A, below, lists these conserved epitopes, together with their respective sources.
Table 2A. Twenty-six Non-Mutated Binding/Neutralizing Epitopes.
[0335] Certain approached described herein (e.g., identification of target regions) utilize identification of an ACE2 interface of a SARS-CoV-2 spike (S) protein, which may, in certain embodiments, be defined as in Lan et al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and comprise the positions listed below:
• ACE2 interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505 ii. Example 2: Engineering a Tailored XBB-Based Antigen
[0336] This Example describes design of an engineered antigen based on (e.g., using as a reference antigen) an XBB variant of a SARS-Cov2 Spike protein. Engineered antigens described in this example aim to lessen and/or avoid activation of a subject’s, such as an individual being vaccinated, memory immune response resulting from memory B and/or T cells, in order to encourage production of new neutralizing antibodies that target particular epitopes associated with an ACE2 binding interface and comprising XBB hallmark mutations.
[0337] Turning to FIG. 4A, an XBB RBD portion was used as a reference antigen for design of engineered antigens. The approach described in this example began with an XBB RBD sequence. A PDB structure including RBD positions 325 to 527 (PED ID 7EAM) was used as a polypeptide model. XBB hallmark mutations, as listed in Table IB, are shown in red in FIG. 4A, while non-mutated regions are shown in gray.
[0338] FIG. 4B shows the XBB RBD model with an identification of an ACE2 interface shown in green, defined as in Lan ct al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and comprising the positions as listed in Example 1, above:
[0339] FIGs. 4A and 4B also show an identified conserved surface in violet coloring. The conserved surface was identified as a continuous non-mutated surface. Amino acid sites identified as belonging to the conserved region in this example are listed below:
• Conserved region: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, 1468, S469, T470, E471, 1472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0340] FIG. 4C shows a colorized version of the XBB structural model. Coloration indicates which residues (amino acid sites) belong to a known neutralizing epitope that is associated with (e.g., targeted) by a neutralizing antibody. One-hundred and thirteen (113) known neutralizing epitopes were identified based on data in the RCSB Protein Databank (PBD) and the Immune Epitope Database and Analysis Resource (IEDB). Each residue was scored according to a number the 113 neutralizing epitopes in which it appears. The colorization in FIG. 4C ranges from yellow to green to blue to indicate a relative frequency of a position belonging to an epitope, with a highest count being 70, colored blue, intermediate counts colored green, and low counts colored yellow (i.e., sites that did nor, or infrequently, belonged to an epitope). FIG. 4C and this present example considers neutralizing epitopes in particular, but subsequent designs have also considered non-neutralizing epitopes.
[0341] The present example aimed to generate engineered versions of XBB that would limit or avoid triggering memory immune responses and, rather, cause creation of new antibodies that were tailored to XBB hallmark mutations. Accordingly, the 113 neutralizing epitopes were compared with the XBB hallmark mutations to identify a subset of twenty-six (26) remaining
epitopes that were not targeted by the XBB hallmark mutations. These conserved, non-mutated, epitopes arc listed in Table 2B, below.
[0342] To disrupt the non-mutated epitopes and conserved surface of the XBB RBD polypeptide model, amino acid modifications were introduced at various positions. Amino acid modifications were generated by selecting from a set of allowable mutations that were known to occur in XBB, XBB sub-lineages, and Omicron. Additional criteria was used to restrict aminoacid modifications introduced to (i) require that modifications cause significant change - e.g., changes from Asp to Glu, Arg to Lys, Ile to Leu, Asn to Gin were not considered; and (ii) prohibit (e.g., exclude) modifications that would break a cysteine bond - for example, modifications to position C391 or C525 were disallowed.
[0343] Subject to this criteria, 200,000 randomly generated variants, each a version of XBB, but with a disrupted version of the conserved region. Each candidate variant was evaluated against various design criteria. In particular, first, candidate designs were required to modify amino acids located within each of the 26 non-mutated epitopes, e.g., to maximize immune escape. Second, candidate designs were evaluated to ensure amino-acid modifications were sufficiently spread across the conserved surface, rather than being clustered together. Lastly, each design was scored using a combination of in-silico structural modeling and a machine learning-based language model. Structural modeling was used to compute an ACE2 binding score, described, for example, in PCT Publication WO 2022/235847 Al , entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022. A machine learning-based language model similar to the model described in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” was used to determine a log-likelihood and semantic change score for each candidate synthetic variant. Candidate designs could then be evaluated on the basis of these criteria.
[0344] FIGs. 4D and 4E show two design examples, Design 1 (FIG. 4D) and Design 2 (FIG. 4E). Both designs introduce amino acid modifications within each non-mutated epitope and also sufficiently spread modifications across the conserved surface. ACE2 binding scores, log-likelihood scores, and semantic change scores were also computed for each design. Design 1
(FIG. 4D) was identified as having a higher ACE2 binding score, while Design 2 (FIG. 4E) was identified as more immune escaping (c.g., based on a higher semantic change score).
[0345] Cyan coloration in FIGs. 4D and 4E identifies positions where amino acid modifications were introduced in the conserved region for Designs 1 and 2, respectively.
Modified amino acid positions for each design are listed below:
• Design 1 (Higher ACE2 Binding) modified positions: L335F Y35 IF N354D S359N A372R L390R Q414K T430I P463S E471Q N481K H519Q;
• Design 2 (More immune escaping) modified positions: L335F A352V A372R
N388K L390R T415I T430I P463L T470N Y473S P479S G482R L518V A520E T523A
Table 2B. Twenty-six Non-Mutated Binding/Neutralizing Epitopes.
iii. Example 3: Engineering a Tailored XBB.1.5-Based Antigen
[0346] This Example describes another example design approach for creating an engineered antigen based on (e.g., using as a reference antigen) an XBB variant of a SARS-Cov2 Spike protein, in particular, XBB.1.5. As with Example 2, engineered antigens described in this example aim to lessen and/or avoid activation of a subject’s, such as an individual being vaccinated, memory immune response, in order to novel B-cell immune responses, with XBB.1.5 as a starting point.
[0347] FIG. 5A illustrates a SARS-CoV 2 virus and its components, including the Spike (S) protein, which is shown in greater detail, together with the ACE2 host receptor to which it binds, on the right-hand side of the figure. FIG. 5B shows a 3D structural representation of an XBB.1.5 RBD portion of a SARS-CoV-2 S protein, including positions 325 to 527 (PED ID 7EAM), which was used as a polypeptide model. XBB.1.5 hallmark mutations, listed below, are shown in red in FIG. 5B, while non-mutated regions are shown in gray.
• XBB.1.5 hallmark mutations: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y5O5H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K
[0348] FIG. 5C shows a 3D structure of a full Spike protein, with XBB.1.5 hallmark mutations in red.
[0349] In order to encourage novel B-cell responses to the XBB.l .5 variant, the design approach in this example began by collecting 1,004 binding and neutralizing B-ccll epitopes from CoV-AbDab and IEDB. These epitopes were compared against XBB.l.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.l.5. This set of conserved epitopes was further curated via manual inspection, including removal of epitopes that were not located on an RBD surface and/or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions in the Spike protein RBD, which are listed in Table 1 above (in Example 1).
[0350] A continuous conserved surface over the XBB.1.5 RBD was also identified, and is shown in violet color in FIG. 5B. The amino acid sites identified as belonging to the conserved region in this example are listed below:
• Uninterrupted conserved surface: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, 1468, S469, T470, E471, 1472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0351] FIG. 5D shows the XBB RBD model with an identification of an ACE2 interface shown in green, defined as in Lan et al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and comprising the positions listed below:
• ACE2 interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505
[0352] FIG. 5E shows a colorized version of the XBB structural model. Coloration indicates which residues (amino acid sites) belong to a neutralizing epitopes that is associated with (e.g., targeted) by a neutralizing antibody. Each residue was scored according to a number of the 1,004 neutralizing and non-neutralizing epitopes, described above, in which it appears. The colorization in FIG. 5E ranges from yellow to green to blue to indicate a relative frequency of a position belonging to an epitope, with a highest count being 287, colored blue, intermediate
counts colored green, and low counts colored yellow (i.e., sites that did nor, or infrequently, belonged to an epitope).
[0353] FIG. 5F illustrates epitope density of the final set of 26 conserved epitopes across the continuous conserved surface. The colorization again ranges from yellow to green to blue indicating a relative frequency of a (amino acid) position belonging to one of the 26 conserved epitopes. Two positions (428 and 518) hit (i.e., were present in) a maximal value of 10 out of the 26 conserved epitopes. Positions located outside of the uninterrupted conserved surface are colored gray.
[0354] In order to disrupt conserved sub-regions of the XBB.1.5 RBD, amino acid modifications were introduced at positions across the conserved surface, which comprised 76 total positions, such that at least amino acid modification to be introduced into each conserved epitope. Amino acid modifications to be introduced were selected from a set of allowed mutations, using sequences of related valiants identified as belonging to XBB, BA, and their sub-lineages. Mutations present within these related lineages (i.e., XBB, BA, and their sublineages) at a frequency of greater than or equal to 1 % were included in the set of allowed mutations. Additional filtering criteria was applied to exclude (i) mutations away from Cys that break a cysteine bond, (ii) mutations between Asp-Glu, Arg-Lys, Ile-Leu and Asn-Gln, and (iii) a set of five mutations due to a large impact on ACE2 binding as indicated by deep mutational scanning (DMS). The final set of allowed mutations (86 in total) used in this example are listed in Table 3, together with an identification of each mutation’s lineage.
[0355] Although the absolute number of possible mutations positions and allowable mutations are relatively low, they yield an extremely large number - approximately 1034 - of possible combinations. For example, consider the following epitope: (511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525). Allowed mutations are: E516Q, L517F/H, L518V, H519L/Q/R/Y, A520E/S/T, P521L/Q/S, T523A/S, so there are: 1 X 2 X 1 X 4 X 3 X 3 X 2 = 144 combinations for this epitope. As such, a number of possible combinations of amino acid modifications within a single conserved epitope can be surprisingly large, with twenty- six total epitopes resulting in the aforementioned large search landscape.
Table 3. Mutations from XBB, BA and their sublineages with a mutation frequency >=1% are enough to escape all 26 remaining epitopes. Mutations away from Cys that break a cysteine bond and mutations between Asp-Glu, Arg-Lys, Ile-Leu and Asn-Gln were not considered. A set of 5 mutations was also removed (marked with strikethrough) due to a large impact on ACE2 binding as indicated by DMS.
[0356] In order to sample the large resultant search space, the design techniques in this example used process 500, shown in FIG. 5G. Process 500 splits creation of candidate engineered variants via introduction of amino-acid modifications into two-steps. In a first step, particular (amino acid) positions to modify were selected 510, to generate multiple sets of position combinations to mutate. In this example, each set of position combinations was generated by selecting positions until positions were selected within each of the 26 conserved epitopes. This process resulted in 200,000 sets of position combinations. As shown in FIG. 5G, a filtering step 520 was also included in the position selection approach, whereby a position spread score was computed for each set of position combinations and the results of step 510 filtered based on the position spread score and number of mutations. Following filtering, 2,000 combinations remained, and were grouped into 20 clusters.
[0357] After arriving at a final set of position combinations, for each position combination in the set, all feasible mutations were generated 530 (using the available options at each position from the set of allowed mutations). Modifications were filtered to ensure they did not introduce a charge flip and to aim to minimize charge changes. This process produced a total of 190,945 candidate variants.
[0358] Four scores, listed below, were then computed 540 for each candidate variant:
• ACE2 binding score;
• Likelihood (log-likelihood) (RBD and full spike);
• Semantic change (RBD and full spike) with respect to XBB.1.5; and
• Mutation co-occurrence;
[0359] The particular approaches used to calculate the ACE2 binding score and mutation co-occurrence scores are described in further detail below. Method for calculating likelihood and semantic change scores are described in detail in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al, entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the content of each of which is incorporated herein by reference in its entirety. The final designs were selected 550
using the Pareto front of all the scores and clustering information, followed by a final manual inspection.
[0360] FIGs. 6A-6F show 3D structures of several of the engineered antigen designs. Tables 4A-4C list, in each row, a particular design with its ID, mutations that were added, and various scores and selection criteria as described herein (identified with bold text in the first column of the tables). Each of FIGs. 6A-6E shows XBB.1.5 hallmark mutations in red and positions where mutations were introduced in yellow, and, except for FIG. 6B, positions not considered for mutation in gray and the conserved surface in violet. FIGs. 6 A and 6B show design S4_l, which was selected to maximize ACE2 binding. FIG. 6C shows design Sl_3, which was selected to minimize a number of mutations. FIG. 6D shows design SI 4, which was selected based on surface distribution of mutations. FIG. 6E shows design S13_3, which was selected on basis of a high log-likelihood score, and FIG. 6F shows design SI 1, which provided a diverse set of mutations. Tables 4A-4C list each final engineered antigen design, indicating the mutations added to XBB.1.5, the total number of mutations, values of each of the four scores computed for the design, and selection criteria.
Table 4A. Vaccine designs with 7-9 mutations. Bolded text in the first column identifies designs shown in FTGs. 6A-6F. Bold text in the list of additional mutations (second column) highlights differences in designs mutating a same set of positions. Scores were ranked among all the designs that altered all 26 epitopes.
Table 4B. Vaccine designs with 10-12 mutations. Bolded text in the first column identifies designs shown in FIGs. 6A-6F. Bold text in the list of additional mutations (second column) highlights differences in designs mutating a same set of positions. Scores were ranked among all the designs that altered all 26 epitopes.
Table 4C. Vaccine designs with 13-15 mutations. Bolded text in the first column identifies designs shown in FIGs. 6A-6F. Bold text in the list of additional mutations (second column) highlights differences in designs mutating a same set of positions. Scores were ranked among all the designs that altered all 26 epitopes.
[0361] FIG. 7 shows radar plots of the scores for each of the final engineered antigen designs, demonstrating a good diversity.
Position Spread Score
[0362] RBD candidates were scored using a position spread score, which estimates how spread the positions are across a (conserved) surface. In this example, position spread score was computed as follows:
where N is the number of mutations in a variant and ri,j is the Cα-Cα distance in the 3D coordinate space between residues number i and j . The coordinates were extracted from the structure 7EAM. The score can be later optimized using the distance over the surface mesh instead of direct 3D distances, as two residues maybe close in 3D space but locates at different faces of the RBD.
Position Clustering
[0363] A clustering algorithm based on one described in Cao et al. 2022 was used to cluster designs, according to the following steps:
Given N set of position combinations,
1. Each sequence was first represented with binary encodings, such that the embedding dimension = number of positions.
2. Each sequence was again represented by the Pearson correlation between it and every other sequence, such that the embedding dimension = number of sequences.
3. The embedding space was reduced to 256 using Multidimensional Scaling (MDS).
Mutation Co-Occurrence Score
[0364] A mutation co-occurrence score was calculated by, for each pair of mutations (mi, mj), calculating, approximately the conditioned frequency that:
[0365] For a given RBD with N mutations (with respect to WT), the mutation cooccurrence score was calculated in this example as the averaged log frequency between all pairs of mutations.
ACE2 Binding Calculation
[0366] The sequences were finally clustered using K-means with MDS embeddings. ACE2 binding score was computed in this example using the deep mutational scanning (DMS) data on ACE2 binding from Starr et al., 2022. This data includes DMS results for positions 331- 531 regarding the RBD:ACE2 binding of 8 variants: Alpha, Beta, Delta, Eta, Omicron BA.1, Omicron BA.2 and two versions of wild-type. The ACE2 binding score sums the “delta bind” (change of logio(KD)) over the mutations and variants to estimate the binding change for any RBD variant. Using non-Omicron data, the ACE2 binding score obtained strong correlation on BA.l and BA.2 data with spearman r = 91.4%, Pearson r = 90.7% and r2=83.5%. Predicted versus experimental change in binding is shown in FIG. 8.
Designed Sequences
[0367] Listings of mutations and sequences of engineered SARS-CoV 2 antigens created via the approaches of the present example.
[0368] Tables 5A and 5B below list the engineered antigen designs created via the approaches described in the present example. Table 5A lists RBD mutations according to the designs in this example. Table 5B provides RBD sequences for each design. As described herein, each sequence is an engineered version of a SARS-CoV 2 Spike protein RBD created using an XBB.1.5 variant RBD as a starting point. FIG. 5C shows metric values computed for the sequences in Tables 5 A and 5B, below. For reference, Table 6 below lists the (natural) XBB.1.5 RBD and the Wuhan RBD sequences. Three versions of the XBB.1.5. RBD sequence are shown in Table 6, allowing for minor variations in the (boundaries of the) particular portion of the spike protein that corresponds to the RBD region.
Table 5A. RBD Mutations for Engineered Antigen Designs.
Table 5B. RBD Sequences for Engineered Antigen Designs.
Table 5C. RBD Mutations with Metrics for Designs in Tables 5A and B.
Table 6. Example XBB.1.5 RBD and Wuhan RBD Portions.
iv. Example 4: Additional Engineered Variants
[0369] FIGs. 9A-9B and 9D-9E show four manually designed RBD engineered antigens. Two constructs designed with a different spread of mutations focusing on distance from XBB.1.5 existing mutations are shown in FIGs. 9A-9B, with Table 7 A listing additional mutations added to XBB.1.5. FIG. 9C is a schematic of a SARS-CoV-2 trimer, adapted from Starr et al., SARS- CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021). https://doi.org/10.1038/s41586-021-03807-6. FIGs. 9D-9E shows two constructs designed boosting broadly conserved epitopes (class 4 and class 5 antibody sites) by disrupting epitopes identified as substantially conserved between SARS CoV 1 and SARS CoV 2 with mutations selected from SARS CoV 1 sequences. Table 7B lists additional mutations that were added to XBB.1.5. Tables 7C and 7D show various scoring metrics computed for each sequence.
Table 7A. Two manual designs focusing on spread of mutations.
Table 7B. Two manual designs for boosting broadly conserved epitopes.
Table 7C. Metrics for Designs in Table 7A.
Table 7D. Metrics for Designs in Table 7B.
v. Example 5: Distributing Introduced Amino Acid Modifications Around Hallmark Mutations
[0370] This example describes an embodiment of antigen engineering techniques described herein in which position spread scores that are used to encourage insertion of new amino acid modifications in a distributed fashion (e.g., as opposed to clustered) across a conserved surface account for, not only other introduced amino acid modifications, but also proximity to existing hallmark mutations (e.g., XBB hallmark mutations). In particular, the antigen engineering methods used in this example proceeded similarly to those in Example 3, above, but also included hallmark mutations in the position spread score (e.g., described in Example 3 above), which was used to evaluate candidate engineered antigen designs. Among other things, this approach is believed to ensure that placing additional mutations directly adjacent to hallmark mutations is avoided. In this manner, the embodiment described in this example maintains unique, e.g., Omicron, epitopes in their unaltered form and preferentially (e.g., only) mutates conserved epitopes. In certain cases, for example, approaches where hallmark mutations are not expressly accounted for in this manner, may present a certain likelihood of altering Omicron epitopes in a manner such that elicited antibodies bind with less affinity to real, desired, target epitopes.
[0371] Tables 8 A and 8B, below list the engineered antigen designs created via the approaches described in the present example (i.e., including XBB hallmark mutations in the position spread score). Table 8A lists RBD mutations according to the designs in this example. Table 8B provides RBD sequences for each design. Table 8C shows various scoring metrics computed for each sequence.
Table 8A. RBD Mutations for Engineered Antigen Designs.
Table 8B. RBD Sequences for Engineered Antigen Designs.
122
RECTIFIED SHEET (RULE 91) ISA/EP
Table 8C. Metrics for Designs in Tables 8A and B.
vi. Example 6: Evolutionary Algorithms and Epitope Alteration Score Versions
[0372] This example describes certain approaches for mutation generation, scoring, and evaluation that may be used, additionally or alternatively, to various approaches described herein in certain embodiments.
[0373] In certain embodiments, evolutionary algorithms (e.g., additionally or alternatively to, or in connection with, the two-step position and type selection technique, described herein) may be used to determine mutations. For example, each solution may be represented as a list of 26 mutations (in the case of SARS-CoV-2 RBD), corresponding to the 26 conserved epitopes described in examples above. In certain embodiments, a pool of solutions can be maintained and updated by swapping the mutations per epitope. Solutions can be scored using position spread, ACE2 binding, and machine learning-based (ML) scores, such as loglikelihood and semantic change.
[0374] In certain embodiments, another version of an epitope alteration score may be used. For example, current versions of the epitope alteration score, described, for example, in PCT Publication WO 2022/235847 Al, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST,” and published November 10, 2022 and PCT Publication WO 2022/235853 Al , entitled “IMMUNOGEN SELECTION,” and also published November 10, 2022, the content of each of which is incorporated herein by reference in its entirety, consider mutations in any single position in an epitope sufficient to ‘evade’ the corresponding antibody. In certain embodiments, an, e.g., more stringent, version of an epitope alteration score may be used that considers position location and distance between a mutation and antibody CDR loops to provide increased granularity and accuracy.
[0375] In certain embodiments, in-silico structural modelling may be used to evaluate complex kinetics. vii. Example 7: Updated Designs Balancing Integrity Risk
[0376] This example describes certain approaches for mutation generation, scoring, and evaluation that may be used, additionally or alternatively, to various approaches described herein
in certain embodiments. In particular, this example provides sequences that were generated to balance risk should certain mutations be detrimental to protein integrity. For example, it was found that certain in-silico sequence generation procedures may over-represent mutations at positions 384, 430, and 463 in a set of suggested designs. Accordingly, the designs presented in this example sought to balance risks (e.g., with respect to diversity of mutations across sets of designs) should certain in-silico suggested mutations be detrimental to RBD integrity.
[0377] Table 9 A, below, lists RBD mutations with respect to (e.g., to be added to) an XBB.1.5 RBD portion of a SARS-CoV-2 S protein (e.g., according to any one of the three XBB.1.5 RBDs provided in Table 6) according to the designs in this example. As explained herein, mutation positions identify positions in the context of (i.e., with reference to) a full length SARS-CoV-2 S protein.
[0378] Table 9B shows scoring metrics computed for each sequence shown in Table 9A.
Table 9A. RBD Mutations for Engineered Antigen Designs.
Table 9B. RBD Mutations with Metrics for Designs in Table 9A.
viii. Example 8: Example In-Silico Design Testing Procedure
[0379] This example describes an example experimental procedure for testing engineered synthetic variants designed via various approaches described herein. The procedure of the present example uses three backbones / constructs for evaluating engineered variants, as follows.
[0380] A pcDNA.3.1-SARS-CoV-2-XBB.1.5-CA19 construct: a mammalian expression plasmid encoding XBB.1.5 spike protein with a truncated cytoplasmic tail, harboring the respective mutated XBB.1.5 RBD. Use of this construct is anticipated to allow for (i) an expression check of the spike protein after transfection in HEK293T cells, (ii) production of pscudovirus, and (iii) analysis of immune escape parameters (c.g., complete escape with no detectable titer).
[0381] A pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM construct: a template for transcription of an mRNA encoding a BNT162b3-like transmembrane-anchored RBD-based vaccine antigen. Use of this construct is anticipated to allow for (i) an expression check of RBD in HEK293T cells, (ii) assessment of antibody-binding using a reference panel of different RBD epitope class binders and/or complex immune serum, and (iii) immunogenicity studies.
[0382] A pcDNA3.4-SARS-CoV-2-XBB.1.5-RBD-his-avi construct: a template for RBD protein production with a HIS tag (allowing for purification) and BAP/avi tag (allowing for assay development). This construct may be used for enzyme-linked immunosorbent assay (ELISA), biolayer interferometry/surface plasmon resonance (BLI/SPR ), and/or protein-protein interaction assays to evaluate, for example, whether known antibodies can bind to the created construct.
[0383] Example procedures for testing designs may include one or more (e.g., up to all) of the below listed steps and/or steps shown in FIG. 10A, which arranges testing steps in a hierarchical fashion:
[0384] Expression and folding check: Expression and folding of various antigens (e.g., each designed antigen) will be evaluated using a flow cytometry based approach in step 1001. An example FACS staining protocol that can be used in connection with this approach is shown below.
[0385] As illustrated in FIG. 10B, following transfection, flow cytometry with hACE2- mFc as a primary binding agent can be used to analyze for intracellular and surface expression of encoded antigen. This approach may use a full length XBB S protein as a reference. Binding affinity of RBD proteins may also be evaluated, along with intracellular and/or extracellular surface expression. Results from this step may be compared with predictions from, e.g., machine learning algorithms used in in-silico design of engineered antigens and fed back in to (e.g., to refine) the in-silico design algorithms.
[0386] Antibody escape - monoclonal and polyclonal: Antibody escape will be evaluated using a flow-cytometry approach with a panel of selected reference antibodies binding to different epitope classes on the RBD (e.g., similar to the assay used for the expression and folding check, but with the reference antibodies used as binding agents instead of hACE2-mFc).
[0387] A variety of reference antibodies may be used and/or selected based on the particular epitope class that they bind to, as well as, additionally or alternatively, a level of affinity to the particular SARS-CoV 2 variant (e.g., reference antigen) that is used as a starting point (i.e., which is mutated) for designing the engineered variants. For example, a panel of reference antibodies may include multiple antibodies that bind to different epitope classes (e.g., A, B, C, D, E, and F). Antibodies that bind to the particular SARS-CoV 2 variant used as a reference antigen that is used to design engineered versions thereof (e.g., with disrupted conserved regions) may be included, along with antibodies that do not bind to that particular SARS-CoV 2 variant (but which bind to other variants).
[0388] For example, Cao et al., Nature, 2021 (doi.org/10.1038/s41586-021-04385-3) (in Supplementary Table 1) provides a listing of 247 neutralizing antibodies, from which a panel can be selected. For example, Table 10A shows a subset of 14 antibodies from the table in Cao et al. that may be used as reference antibodies. FIG. 10C shows example binding assay data for several of the antibodies listed in Table 10A. Antibodies exhibiting XBB.1.5 binding are shown in Table 10B, along with IC50 values for wild-type and BA.l variants (* denotes data provided in Cao et al.). Table 10C shows the antibodies from Table 10A, with XBB.1.5 binding data, highlighting those antibodies that bind to XBB.1.5 in green text. Without wishing to be bound to any particular theory, certain classes of antibodies, such as antibodies that bind to Class A-D epitopes, do not bind to BA.l S protein and, accordingly, are not expected to bind to XBB S
protein. Certain antibodies that bind to class E and F will be checked for binding to XBB S protein. Abrogated/reduced binding of reference antibodies can be evaluated using titration EC50 values in step 1002.
Table 10A. Example Panel of Reference Antibodies.
Table 10B. Antibodies exhibiting XBB.1.5 binding.
Table 10C. Antibodies from Table 10A, with XBB.1.5 binding/assay data.
[0389] Additionally or alternatively to various classes of antibodies listed in Tables 10A- C, panels of antibodies may include various other examples of antibody clones, for example, including (but not limited to) antibodies as listed below and/or similar antibodies:
• WRAIR-2057 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants). WRAIR-2057 is described, for example, in further detail in Dussupt, V. et al., Low-dose in vivo protection and neutralization across SARS-CoV-2 valiants by monoclonal antibody combinations. Nat Immunol 22, 1503-1514 (2021). https://doi.org/10.1038/s41590-021-01068-z.
• COVOX-222, which is included in Supplementary Table 1 of Cao et al., Nature, 2021 (doi.org/10.1038/s41586-021-04385-3);
• COVOX-45 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants). COVOX-45 is described, for example, in further detail, in Dejnirattisai, W. et al., The antigenic anatomy of SARS-CoV-2 receptor binding domain, Cell, 184 (8), 2183-2200. e22, 2021, https://doi.Org/10.1016/j.cell.2021.02.032.
• S2H97 (without wishing to be bound to any particular theory, this antibody is believed to bind at a hitherto highly conserved site at the RBD flank and have high affinity to Omicron variants). S2H97 is described, for example, in further detail in Starr, T.N., et al. SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature 597, 97-102 (2021). https://doi.org/10.1038/s41586-021-03807-6.
• 553-49, which is described, for example, in further detail in Zhan W et al., Structural Study of SARS-CoV-2 Antibodies Identifies a Broad-Spectrum Antibody That Neutralizes the Omicron Variant by Disassembling the Spike Trimer. J Virol. 96(16):e0048022. 2022 . doi: 10.1128/jvi.00480-22.
[0390] Antibody escape may, additionally or alternatively, be evaluated using immune scrum from vaccinatcd/convalcsccnt individuals. Abrogated/reduced binding of complex immune serum can be evaluated using titration EC50 values and comparing against XBB1.5 values in step 1003.
[0391] A third approach for assessing escape may employ a pseudovirus generation protocol and pVNT assay set-up to check for loss of nAb titers. This approach would aim to check for abrogated neutralization of respective pseudovirus in step 1004.
[0392] A final step, for example as shown in FIG. 10 A, may include a dedicated immunogenicity study, for example to evaluate whether scrum from immunized animals neutralizes XBB1.5, and/or one or more other variants. Vaccine compositions based on engineered SARS-CoV 2 antigens according to various embodiments described herein may be administered to vaccine naive animals (e.g., mice) (e.g., to evaluate immune response and its magnitude) and to vaccine experienced animals (e.g., mice), e.g., to evaluate ability to overcome immune imprinting.
[0393] FIG. 10B shows an example protocol, illustrating transfection and flow cytometry steps. As shown in FIG. 10B, pcDNA3.1-SARS-CoV-2-Swt-CA19 and/or pcDNA3.1-SARS- CoV-2-SXBB. 1 .5-CA19 may be transfected into HEK293T/17 cells and binding assays as described above can be used to evaluate hACE2 binding and (neutralizing) antibody binding, for example to carry out various steps in the decision tree shown in FIG. 10A.
[0394] Concentration ranges of monoclonal antibodies, hACE2, and dilution series of human BNT162b23 (triple-vax) polyclonal serum that result in sigmoidal binding curves are identified as appropriate. Based on this criteria, mAbs can be tested in a concentration range of lOpg/mL - 0.128 ng/mL (in 8-step 5-fold dilution series) based on previous apparent 1C50 values for WT and BA.l Spike, for example as shown in FIG. 10C. As shown in FIG. 10D, hACE2-mFc binding may be tested using a concentration rage of 25 pg/mL to 0.32 ng/mL. FIG. 10D also shows that XBB.l .5 Spike displays an approximately 6.7-fold higher apparent affinity towards hACE2 binding in comparison with Wild-Type Spike. As shown in FIG. 10E triple-vax IM post-boost polyclonal serum pool can be tested in a dilution series ranging from 1:20 - 1:1.562.500.
[0395] Table 11 shows several variant-of-concern (VOC) mutations, together with their numerical metrics, which may be used as references.
Table 11. Variant of concern (VOC) mutations and associated scores.
[0396] An example FACS-staining protocol is shown below, with Tables 12A and B, below, showing master mixtures used for secondary antibody labeling steps.
Staining with antibodies, hACE2-mFc or human serum pool
• Detach cells
• Add PBS and centrifuge cells
• Count cells:
• Distribute viable cells/well in 96-round-bottom well plates according to plate layout
• Centrifuge plates with cells
• Discard supernatants and vortex plates to singularize cells
• Add FACS buffer and centrifuge plate
• Add 1st antibody, ACE2-mFc or human serum pool, gently tap plates to mix cells
• Incubate for about 15 min at 4°C in the dark
• Add FACS buffer and centrifuge plate
• Discard supernatants and vortex plates to singularize cells
• Wash twice with FACS buffer and centrifuge plate
• Add Mastermix (MM) with secondary antibody, gently tap plates to mix cells
• Incubate for about 15 min at 4°C in the dark
• Add 200 pL/well FACS buffer (DPBS + 2% FBS hi. + 2 mM EDTA) and centrifuge plate: 5 min, 460 xg, RT
• Discard supernatants and vortex plates to singularize cells
• Wash twice with FACS buffer and centrifuge plate:
• Add Histofix (in PBS) to the cells, resuspend and incubate for 15 min at 2-8°C
• Centrifuge cells (5 min, 450xg)
• Discard supernatants and vortex plates to singularize cells
• Wash twice with FACS buffer and centrifuge plate
• Resuspend cells in FACS buffer
• Store plates in the dark at 4°C until they are measured
Table 12A. 2nd antibody Mastermix I (for primary antibodies and human serum).
Table 12B. 2nd antibody Mastermix II (for hACE2-mFc).
ix. Example 9: Evaluation of In-Silico Designed Constructs via Binding Assays
[0397] This example describe results of various assays in accordance with the example testing procedure described in Example 8, above, carried out for certain constructs described herein.
[0398] A first round of tests were carried out for engineered XBB.1.5 variants expressed in context of a full length SARS-CoV-2 spike (S) protein. In particular, certain constructs described herein were incorporated into a pcDNA.3.1-SARS-CoV-2-XBB.1.5-CA19 construct harboring the respective mutated XBB.l .5 RBD. HEK293T/17 cells were seeded into flasks and incubated for 2 days at 37°C and 7.5% CO2. Constructs were transfected into the HEK293T/17
cells and incubated overnight at 37°C and 7.5% CO2. Flow cytometry was used to (i) evaluate expression levels via an anti-S2 fragment antibody, (ii) check for conserved ACE2 binding capacity (via hACE2 binding) and (iii) assess abrogation of RBD-targeted antibody binding, using the panel of five (5) monoclonal antibodies (mAbs) having demonstrated binding to XBB.1.5, shown in Table 10B.
[0399] Turning to FIGs. 11A-11C, binding tests were carried out for certain variant designs shown in Table 9A, expressed in the context of a full length spike (S) protein. In particular, ACE-2 binding was evaluated using a hACE2-mFc antibody, according to the FACS protocol described in Example 8, above. Binding responses for each of the five mAbs listed in Table 10B were also evaluated via the protocol described in Example 8, above. FIGs. 1 1 A-l 1C show results for the S43 engineered XBB.1.5 variant design (FIG. 11B) and the S48 engineered XBB.1.5 variant design (FIG. 11C), along with a (parental / original) XBB.1.5 reference (FIG. 11 A). Among other things, both the S43 and S48 variants were found to exhibit abrogated binding for all but 1-2 of the five mAbs, and the S43 variant maintained hACE-2 binding.
[0400] Turning to FIGs. 12-14, variant designs shown in Table 9A were also expressed in the context of a BNT162b3-like (trimerized transmembrane-anchored RBD-based) vaccine antigen. HEK293T/17 cells were transfected with BNT162b3-XBB.1.5 and each respective variant’s RNAs. Flow cytometry was then used to assess ACE-2 binding via hACE2-mFc binding and binding of the five mAb antibody panel (listed in Table 10B) as before. Binding curves for polyclonal vaccine serum (triple BNT162b2 vax) were also evaluated to assess whether impact on polyclonal serum dose-response was greater in comparison with hACE-2 dose-response for the engineered variants. Without wishing to be bound to any particular theory, it is believed that a stronger impact on polyclonal serum dose-response compared to hACE-2 dose-response when compared to the parental XBB.1.5 for a particular variant suggests successful ‘masking/mutation’ of conserved epitopes.
[0401] FIGs. 12A-B show dose response curves for an XBB.1.5 reference (FIG. 12A), and the S43 variant design (FIG. 12B). As shown in FIG. 12B, design S43 expressed in the context of a trimerized TM-anchored RBD shows close to unaltered hACE-2 binding doseresponse when compared to XBB.1.5. Importantly, binding of Class A Antibody 2, Class F
Antibody 1 , and Class B Antibody 1is abrogated, with only binding of Class F Antibody 2and Class E Antibody 2conscrvcd.
[0402] FIGs. 13A and 13B show results repeating those of FIGs. 12A and B (FIG. 13A shown reference dose response curves and FIG. 13B showing engineered antigen design S43 dose-response curves) confirming the results shown in FIG. 12B - namely, conservation of hACE-2 binding, as well as Class F Antibody 2and Class E Antibody 2 mAb binding, but abrogation of Class A Antibody 2, Class F Antibody 1, and Class B Antibody Ibinding. Polyclonal serum binding was also compared with hACE2 binding as shown in FIGs. 13C and 13D. While, the shift (engineered XBB.1.5 antigen versus parental XBB.1.5) in the polyclonal serum binding curve is similar to, rather than larger than, the hACE-2 binding, the abrogated mAb binding observed in the multiple sets of results shown in FIGs. 12A and 12B and 13A and 13B provides evidence that the mutations introduced into the engineered S43 design successfully disrupt conserved epitopes.
[0403] FIGs. 14A and 14B show another set of dose response curves for XBB.1.5 reference and for the S48 engineered antigen design from Table 9A expressed in the context of a trimerized TM-anchored RBD. S48 displays conserved hACE-2 and Class F Antibody 2binding, as did S43 (S48 harbors 5/7 mutations also found in S43); Class E Antibody 2shows minimal residual binding, whereas binding of Class A Antibody 2, Class F Antibody land Class B Antibody lis completely abrogated.
[0404] Turning to FIGs. 14C and 14D provide additional data that suggest potential for successful disruption of conserved epitopes in S48. FIGs. 14C and 14D compare (i) shift in hACE2 binding curve (FIG. 14C) for the S48 variant relevant to the parental XBB.1.5 reference with (ii) the shift in binding curve for polyclonal serum (FIG. 14D). The shift in ACE2-binding (approximately 4-fold) is less than the shift observed in serum binding (> 10-fold) when comparing parental XBB.1.5 vs. S48, which hints towards progressive abrogation of binding antibody responses.
[0405] Table 13 summarizes results of the screening data described in this example in a tabular format. The first two rows of the table show binding levels for XBB.1.5 and the ideal, desired engineered antigen. The columns representing binding data are arranged in two groups, corresponding to the different expression contexts analyzed - (i) full length XBB.1.5 spike
(pcDNA.3. l -SARS-CoV-2-XBB.1 .5-CA19), (ii) trimerized transmembrane (TM) anchored RBD design (pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM). As shown in Table 13, the reference antigen, XBB.1.5 exhibits strong ACE-2 binding and binds to all of the five mAbs in the panel (which were selected to be mAbs with known XBB.1.5 binding). An ideal, desired, engineered variant should maintain ACE-2 binding, but escape all five of the mAbs known to bind to XBB.1.5, so as to avoid triggering a memory immune response. As shown in the table, constructs S43 and S48 exhibit behavior close to that of the idealized construct. S43 shows conserved ACE-2 binding when expressed as a full-length S protein as well as in the trimerized TM-anchored RBD format, while abrogating binding to all but 1-2 of the mAb panel, depending on expression context. Construct S48 performed particularly well when expressed as trimerized TM anchored RBD, showing conserved ACE-2 binding and only binding with one of the five mAbs.
Table 13. Summary of Certain Test Data.
x. Example 10: Prophetic Mouse Immunization Study
[0406] This example describes planned procedure for, and anticipated results from, an experiment to evaluate whether RNA compositions encoding engineered antigens comprising SARS-CoV-2 S protein and/or portions thereof (e.g., an RBD domain) as described herein induce an immune response characterized by increased naive B cell activation and/or decreased memory B cell activation in vaccine-experienced subjects (mice in the present example).
[0407] Turning to FIG. 15 A, vaccine candidates comprising RNA encoding engineered antigens designed and screened according to the approaches described herein will be administered to mice previously exposed to full length SARS-CoV-2 S protein. In particular, as shown in FIG. 15 A, vaccine candidates are tested on mice previously administered two doses of RNA encoding a full length SARS-CoV-2 S protein, with each engineered antigen vaccine candidate administered as a third dose (booster).
[0408] Mice will be split into groups, each comprising approximately a same number, x, of members, and administered a dosing regimen as illustrated in FIG. 15A. The total number of groups depends on the number of engineered antigen candidates, as well as different expression formats thereof, that will be tested. For example, FIGs. 15 illustrate an example immunization study where engineered antigen designs S43 and S48 having sequences as described in Table 9A are evaluated in contexts of mRNA’s encoding (i) a full-length S protein, analogous to BNT162b2, but encoding the particular (e.g., S43 or S48) engineered XBB.1.5 variant and (ii) a trimerized TM anchored RBD domain. Other candidate antigens and/or expression contexts may be included similarly. In one example, a number of mice in each group will about 7-8. As
shown in FIG. 15 A, mice in each group are first administered two doses of a monovalent composition comprising RNA encoding a SARS-CoV-2 S protein of a Wuhan strain (BNT162b2). First and second doses of the monovalent vaccine are administered 21 days apart. Five (5) weeks after being administered the first dose, groups of mice are divided into groups based on neutralization titers against the Wuhan strain (e.g., mice are allocated into groups so that average neutralization titers are approximately the same in each group). In some embodiments, mice can be allocated into groups based on pseudo-virus neutralization titers (e.g., as shown in FIG. 15 A).
[0409] Each candidate vaccine will then be administered as a third dose, 18 weeks, after administering a first dose of vaccine (i.e., day 126, as shown in FIG. 15A). In certain embodiments, versions of the immunization study approach illustrated in FIG. 15A may be performed with the third, candidate vaccine, dose being administered on a shorter timeline following the first two doses. Without wishing to be bound to any particular theory, a third, candidate, dose may be administered quickly after a second dose, so long as sufficient time has been allowed for completion of immune reactions, e.g., generation of B-cells, in response to the initial, e.g., Wuhan, antigens. In certain embodiments, 28 days (4 weeks) or more may be sufficient, allowing, for example, for a first dose (of BNT162b2) to be administered on day 0, a second dose (of BNT162b2) administered on day 21, and vaccine candidates to be administered as third (e.g., booster) doses four weeks later, e.g., on day 49 (week 7) or later.
[0410] Various engineered antigen designs and expression context formats may be administered as a third dose and compared. For example, as shown in FIG. 15 A, candidates S43 and S48 may be administered as third doses, as a full spike (S) format - BNT162b2 (S43) and BNT162b2 (S48), as well as in a trimerized transmembrane-anchored RBD format - RBD-TM (S43) and RBD-TM (S48). In certain embodiments, a number of controls and/or reference vaccines may also be used. For example, as shown in FIG. 15 A, a third dose of BNT162b2 may be administered. Additionally or alternatively, full spike (S) and RBD-TM versions of the parental XBB.1.5 variant may be administered for comparison with the engineered valiants thereof. In certain embodiments, one test group may not be administered any third dose. Table 14A, below, lists an example set of vaccine candidates (selected based on screening data described herein, e.g., in the previous example), several of which are also listed in FIG. 15A.
Several options for certain other candidate vaccines, which may be used additionally or alternatively, arc listed in Table 14B. Tables 14A and 14B provide description of vaccine candidates formats and references to exemplary sequences that are included in Tables 14C-14I. Other vaccine candidates comprising engineered antigens (e.g., variants of other VOCs) and reference compositions (vaccine doses) can be prepared and evaluated in a manner analogous to that described herein with respect to XBB.1.5.
Table 14A. Vaccine Candidates for Administering to Vaccine-Experienced Mice.
Table 14B: Additional Vaccine Candidates Options.
Table 14C: Sequences of an Exemplary RNA Construct Encoding a Soluble, Trimerized RBD (SP19-XBB.1.5_RBD- GS_Linker-Fibritin_long).
Table 14D: Sequences of an Exemplary RNA Construct Encoding a Full-Length, Prefusion- Stabilized SARS-CoV-2 S Protein (XBB.1.5_P2).
Table 14E: Sequences of an Exemplary RNA Construct Encoding a Soluble, Trimerized RBD (SP16-XBB.1.5_RBD- GS_Linker-Fibritin_long).
Table 14F: Sequences of an Exemplary RNA Construct Encoding a Soluble Trimerized SI Domain (XBB.1.5_Sl-GS_Linker- Fibritin long).
Table 14G: Sequences of an Exemplary RNA Construct Encoding a Membrane-Tethered Spike Protein Comprising a C- terminal Truncation (Spike_deltal9).
Table 14H: Sequences of an Exemplary RNA Construct Encoding a Trimerized, Membrane- Anchored RBD (SP19- XBB.1.5_RBD-GS_Linker-Fibritin_Short-GS_Linker-TM(deltal9)).
Table 141: Sequences of an Exemplary RNA Construct Encoding a Membrane-Anchored SI Domain (XBB.1.5_S1- GS_Linker-Fibritin_Short-GS_Linker-TM(deltal9)).
[0411] Blood samples will be collected immediately before, and 3 weeks, 5 weeks, 9 weeks, 13 weeks, 17 weeks, 18 weeks, 19 weeks, 22 weeks, 26 weeks, 30 weeks, 34 weeks, 35 weeks, 37 weeks, and 39 weeks after administering a first dose of RNA. Thirty-nine (39) weeks after administering a first dose of RNA, mice will be sacrificed, and final blood, lymph node, and spleen samples collected for analysis.
Blood Sample Analysis
[0412] Blood samples will be screened for titers of antibodies that bind and neutralize various SARS-CoV-2 strains and variants (e.g., using ELISA and pseudovirus assays that described herein). Each spleen sample can be analyzed individually. For lymph node samples, samples from two mice can be combined for analysis.
Spleen and Lymph Node Sample Analysis
[0413] B cells are isolated from spleens and lymph nodes will be phenotyped to determine B cell type (e.g., naive, memory, or plasma) and binding specificity (e.g., specificity to the S proteins of various SARS-CoV-2 strains and variants). Phenotyping can be performed using the FACS-based and depletion assays analogous to those depicted in FIG. 16A and 16B and described in further detail in Quandt and Muik et al., Science Immunol. , 7(75) eabq2427 (2022) (doi/10.1126/sciimmunol.abq2427), the contents of which is incorporated by reference herein in its entirety. Pseudovirus neutralization titers will also be collected for each of the sera samples. BCR repertoire analysis will also be performed for B cells isolated from spleen samples.
[0414] For RBD binding assays, samples from all mice can be grouped together to generate enough sample to perform the methods. Each sample can be screened for negative, Wuhan-specific, XBB.1.5 exclusive binders, and cross reaction between Wuhan and XBB.1.5 binding.
[0415] Among other things, the protocol described in the present example can be used to characterize the binding specificity of B cells in subjects administered a booster vaccine that delivers an antigen of a variant of concern - here, an XBB. 1.5 S protein or an immunogenic portion thereof. Specifically, this experimental protocol can be used to determine the relative number of B cells that are specific to XBB.1.5 and/or what portions of the XBB.1.5 S protein those B cells recognize. These results can be used to assess the impact of immune imprinting on various vaccine candidates, and the ability of certain vaccine candidates to evade immune imprinting and induce a de novo immune response.
[0416] Vaccine candidates that are less susceptible to immune imprinting (i.e., more likely to generate a de novo response) can be characterized by one or more of: (i) an increased proportion of B cells that are specific to a variant of concern (namely, XBB.1.5 or an immunogenic portion thereof) delivered by the vaccine candidate, (ii) increased neutralization titers against a variant of concern encoded by the vaccine candidate, and/or (iii) an increased number of B cell receptors that recognize epitopes that are unique to an antigen encoded by the vaccine candidate (i.e., increased B cell breadth).
[0417] Additional analysis techniques include spleen sample analysis, lymph node analysis, and blood sample characterization as shown in FIG. 17. The spleen sample analysis may involve a preparation of single cell suspension that results in >5xl07 leucocytes/spleen after isolation and red blood cells (RBC) depletion (with 80-90% cell viability). Out of these isolated cells, approximately 1.5xl07 leucocytes/spleen may be used for FACS phenotyping for immunogenicity testing. This analysis may involve, for example, naive, memory, and plasma cell specific for the different S proteins. Out of spleen isolated cells, approximately 1.5x107 leucocytes/spleen may be used for tagging and pooling with subsequent magnetic- activated cell sorting (MACS). Subsequent steps may involve FACS sorting and staining for memory and plasma cell population. Another subsequent step may include BCR repertoire analysis. Remaining leucocytes extracted from spleen may be used for enzyme-linked immunosorbent spot (ELISpot) assays and freezing of remaining samples or freezing and ELISpot. Lymph node analysis may involve a preparation of single cell suspension that results in approximately IxlO6 leucocytes per inguinal (iLN) and pelvic (pLN) lymph nodes. This analysis may involve FACS phenotyping for immunogenicity testing. This analysis may involve, for example, naive, memory, and plasma cell specific for the different S proteins.
Finally, blood sample characterization may involve pseudotyped virus neutralization tests (pVNTs) ELISA.
Immune Response in Vaccine Naive Mice
[0418] Turning to FIG. 15B, in certain embodiments, vaccine candidates, along with references, may be tested in vaccine naive mice. As illustrated in FIG. 15B, tests in vaccine naive mice may be carried out on shorter time scales and can be used to evaluate and/or confirm that vaccine candidates based on the various engineered antigen compounds can elicit meaningful immunogenic response to XBB.1.5 on their own. FIG. 15B shows an example protocol where vaccine candidates and references as described herein are administered to vaccine naive mice in a two-dose format, with the first dose administered on day zero and the second dose administered three weeks later, on day 21.
[0419] Blood samples can be collected immediately before, and 2 weeks, 3 weeks, 4 weeks, and 7 weeks after administering a first dose of RNA. Seven (7) weeks after administering a first dose of RNA, mice will be sacrificed, and final blood, lymph node, and spleen samples collected for analysis. Neutralizing antibody generation can be tested via pVNT assays using XBB.1.5 pseudo-virus, while binding antibodies can be tested via ELISA, using XBB.1.5 RBD as a target. Ideal candidates will successfully maintain neutralizing antibody titers, e.g., showing levels comparable to those for the parental XBB.1.5 (e.g., RBD- TM (XBB.1.5), while, at the same time, exhibit reduced binding antibody titers in the ELISA tests, e.g., due to disruption of conserved regions. xi. Example 11: Assessment and Selection of Individual Engineered RBD Mutations for Engineered Antigen Designs
[0420] This example describes assessment and selection of RBD mutations that were introduced in certain engineered antigens described herein. In particular, mutations present in various constructs were evaluated on an individualized basis and subsets were selected to create additional fine-tuned engineered antigens.
[0421] In particular, analysis of various in-silico designed constructs provided herein led to the identification of certain mutations that were introduced in construct S48 as having impact on expression. As explained herein, construct S48 corresponds to a SARS-CoV-2 S protein RBD with (i) XBB.1.5 hallmark mutations together with (ii) an additional set of introduced mutations that were engineered to disrupt memory triggering conserved regions of a baseline, reference, XBB.1.5 reference antigen. As shown in Table 9A, this additional set of introduced mutations in construct S48 included the following mutations:
• S48 additional mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N.
[0422] Of these S48 additional mutations, a subset was determined to be beneficial for expression of constructs comprising the S48 RBD design, while another subset was selected for further evaluation, e.g., as potentially sub-optimal and/or causing a reduction in expression. Mutations determined to be beneficial include, without limitation, K356T, P384S, T43O1, F464Y, and H519N. Mutations selected for further evaluation included L335F and L390R.
[0423] Based on these identified mutations, an additional round of construct designs was created based on an XBB.1.5 RBD reference antigen (i.e., including XBB.1.5 hallmark mutations) in which (i) the set of mutations determined to be beneficial were retained, (ii) the two mutations selected for further evaluation were excluded from a majority of the constructs (and none included both the L335F and L390R mutations) and (iii) an additional set of new mutations were introduced.
[0424] New mutations were introduced following an approach similar to that described in Example 3 and with reference to process 500 in FIG. 5G.
[0425] FIG. 18 illustrates this additional set of tailored constructs. Construct design labels are shown along the vertical, such that each row corresponds to a particular construct design and a listing of possible mutations is shown along the horizontal. XBB.1.5 hallmark mutations are identified via a black asterisk (“*”) symbol, the mutations determined to be beneficial identified via a green star, the two mutations selected for further evaluation identified via a red downward pointing triangle, and new mutations identified via purple diamonds. Shading (dark green) identifies mutations present in a particular construct. As
illustrated in FIG. 18, all constructs included the XBB.1.5 hallmark mutations as well as the five mutations determined to be beneficial. Constructs S122-S127 and S 128 included the L335F mutation and constructs S145 and S156 included the L390R mutation. New mutations were introduced in various constructs, as illustrated in FIG. 18. Tables 15A and 15B, below lists the RBD mutations (including XBB.1.5 hallmark mutations) and additional, engineered, RBD mutations (i.e., in addition to the reference XBB.1.5 mutations), respectively, for each construct design.
[0426] Construct designs shown in FIG. 18 and identified in Tables 15A and 15B, below, were then tested by incorporating each construct designs set of additional mutations (i.e., as listed in Table 15, below) along with the XBB.1.5 hallmark mutations into a BNT162b3 mRNA for in-vitro testing. In particular, the BNT162b3 construct, shown in FIG. 21A is 1397 base pair mRNA encoding a membrane- anchored RBD and Fibritin domain (F) together with a viral signal peptide. As shown in FIG. 19A, the construct includes a secretory signal (“sec”), an RBD domain (“RBD”), a Fibritin domain (“F”) and a transmembrane anchor (“TM”). For each construct design, the RBD encoding domain was modified to encode the XBB.1.5 hallmark mutations and each construct design’s particular set of additional mutations.
[0427] To evaluate expression, flow cytometry was used to assess ACE-2 binding via hACE2-mFc binding assay, which was carried out for the 22 construct designs shown in FIG. 18 and listed in Table 15, as well as for XBB.1.5 and the S48 construct, as points of reference. Results are shown in FIG. 19B and in Table 16A, below. To evaluate immune escape - e.g., the ability of a construct to avoid antibodies produced via prior Wuhan-induced immune response and, accordingly, potentially trigger a de-novo immune response, binding was against polyclonal vaccine serum from triple BNT162b2 vaccinated patients was measured. Results for the 22 construct designs, along with XBB.1.5 and the S48 construct are shown in FIG. 19C and Table 16B, below. Values in Tables 16A and 16B were obtained by repeating the experiment twice, and provide results in terms of area under the curve (AUC) value and flow cytometry intensity (FC).
[0428] As shown in FIG. 19B and Table 16 A, all 22 construct designs from FIG. 18 and Tables 15A and 15B showed higher binding to ACE2 than the S48 construct design, indicative of higher expression. In terms of immune escape, the S48 construct design
exhibited the most significant decrease in binding to the BNT162b2 triple vaccinated patient serum. Several construct designs listed in Tables 15 A and 15B showed notable decrease in (BNT162 serum) binding. In particular, seven constructs show more than 1.5-fold decrease in serum binding while showing no more than a two (2)-fold decrease in ACE2. These construct designs are identified via an asterisk (*) in Tables 16A and 16B, below.
[0429] FIGs. 20A-C shows the binding affinity for XBB.1.5, S48, and seven construct designs shown in FIG. 18 to various SARS-CoV-2 monoclonal antibodies. FIG. 20A shows results for five antibodies for which the seven construct designs showed a full regain of binding, FIG. 20B shows results for five antibodies for which the seven construct designs showed a partial regain of binding, and FIG. 20C shows results for two antibodies for which the seven constructs exhibited little to no binding. The results appear consistent with the polyclonal serum data, with S48 still showing most significant decrease in binding with the studied monoclonal antibodies.
Table 15A. RBD Mutations for Engineered Antigen Designs.
Table 15B. RBD Mutations for Engineered Antigen Designs.
Table 16A. Expression results, evaluated via an hACE2-mFC binding assay. Area under the curve (AUC) and FC intensity (relative to XBB. 1.5=1.00) values are provided.
Table 16B. Immune escape results, evaluated via a binding assay for serum from triple vaccinated BNT162b3 patients. Area under the curve (AUC) and FC intensity (relative to XBB.1.5=1.00) values are provided.
G. Computer System and Network Environment
[0430] As shown in FIG. 21 an implementation of a network environment 2100 for use in providing systems and methods as described herein is shown and described. In brief overview, referring now to FIG. 21, a block diagram of an exemplary cloud computing environment 2100 is shown and described. The cloud computing environment 2100 may include one or more resource providers 2102a, 2102b, 2102c (collectively, 2102). Each resource provider 2102 may include computing resources. In some implementations, computing resources may include any hardware and/or software used to process data. For
example, computing resources may include hardware and/or software capable of executing algorithms, computer programs, and/or computer applications. In some implementations, exemplary computing resources may include application servers and/or databases with storage and retrieval capabilities. Each resource provider 2102 may be connected to any other resource provider 2102 in the cloud computing environment 2100. In some implementations, the resource providers 2102 may be connected over a computer network 2108. Each resource provider 2102 may be connected to one or more computing device 2104a, 2104b, 2104c (collectively, 2104), over the computer network 2108.
[0431] The cloud computing environment 2100 may include a resource manager 2106. The resource manager 2106 may be connected to the resource providers 2102 and the computing devices 2104 over the computer network 2108. In some implementations, the resource manager 2106 may facilitate the provision of computing resources by one or more resource providers 2102 to one or more computing devices 2104. The resource manager 2106 may receive a request for a computing resource from a particular computing device 2104. The resource manager 2106 may identify one or more resource providers 2102 capable of providing the computing resource requested by the computing device 2104. The resource manager 2106 may select a resource provider 2102 to provide the computing resource. The resource manager 2106 may facilitate a connection between the resource provider 2102 and a particular computing device 2104. In some implementations, the resource manager 2106 may establish a connection between a particular resource provider 2102 and a particular computing device 2104. In some implementations, the resource manager 2106 may redirect a particular computing device 2104 to a particular resource provider 2102 with the requested computing resource.
[0432] FIG. 22 shows an example of a computing device 2200 and a mobile computing device 2250 that can be used to implement the techniques described in this disclosure. The computing device 2200 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 2250 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices. The components
shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.
[0433] The computing device 2200 includes a processor 2202, a memory 2204, a storage device 2206, a high-speed interface 2208 connecting to the memory 2204 and multiple high-speed expansion ports 2210, and a low-speed interface 2212 connecting to a low-speed expansion port 2214 and the storage device 2206. Each of the processor 2202, the memory 2204, the storage device 2206, the high-speed interface 2208, the high-speed expansion ports 2210, and the low-speed interface 2212, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 2202 can process instructions for execution within the computing device 2200, including instructions stored in the memory 2204 or on the storage device 2206 to display graphical information for a GUI on an external input/output device, such as a display 2216 coupled to the high-speed interface 2208. In other implementations, multiple processors and/or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein the plurality of functions are performed by any number of processors (one or more) of any number of computing devices (one or more). Furthermore, where a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (one or more) of any number of computing devices (one or more) (e.g., in a distributed computing system).
[0434] The memory 2204 stores information within the computing device 2200. In some implementations, the memory 2204 is a volatile memory unit or units. In some implementations, the memory 2204 is a non-volatile memory unit or units. The memory 2204 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0435] The storage device 2206 is capable of providing mass storage for the computing device 2200. In some implementations, the storage device 2206 may be or contain a computer-readable medium, such as a floppy disk device, a hard disk device, an
optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 2202), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine-readable mediums (for example, the memory 2204, the storage device 2206, or memory on the processor 2202).
[0436] The high-speed interface 2208 manages bandwidth-intensive operations for the computing device 2200, while the low-speed interface 2212 manages lower bandwidthintensive operations. Such allocation of functions is an example only. In some implementations, the high-speed interface 2208 is coupled to the memory 2204, the display 2216 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 2210, which may accept various expansion cards (not shown). In the implementation, the low-speed interface 2212 is coupled to the storage device 2206 and the low-speed expansion port 2214. The low-speed expansion port 2214, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input/output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.
[0437] The computing device 2200 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 2220, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 2222. It may also be implemented as part of a rack server system 2224. Alternatively, components from the computing device 2200 may be combined with other components in a mobile device (not shown), such as a mobile computing device 2250. Each of such devices may contain one or more of the computing device 2200 and the mobile computing device 2250, and an entire system may be made up of multiple computing devices communicating with each other.
[0438] The mobile computing device 2250 includes a processor 2252, a memory 2264, an input/output device such as a display 2254, a communication interface 2266, and a transceiver 2268, among other components. The mobile computing device 2250 may also be provided with a storage device, such as a micro-drive or other device, to provide additional
storage. Each of the processor 2252, the memory 2264, the display 2254, the communication interface 2266, and the transceiver 2268, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.
[0439] The processor 2252 can execute instructions within the mobile computing device 2250, including instructions stored in the memory 2264. The processor 2252 may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 2252 may provide, for example, for coordination of the other components of the mobile computing device 2250, such as control of user interfaces, applications run by the mobile computing device 2250, and wireless communication by the mobile computing device 2250.
[0440] The processor 2252 may communicate with a user through a control interface 558 and a display interface 2256 coupled to the display 2254. The display 2254 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 2256 may comprise appropriate circuitry for driving the display 2254 to present graphical and other information to a user. The control interface 2258 may receive commands from a user and convert them for submission to the processor 2252. In addition, an external interface 2262 may provide communication with the processor 2252, so as to enable near area communication of the mobile computing device 2250 with other devices. The external interface 2262 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.
[0441] The memory 2264 stores information within the mobile computing device 2250. The memory 2264 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 2274 may also be provided and connected to the mobile computing device 2250 through an expansion interface 2272, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 2274 may provide extra storage space for the mobile computing device 2250, or may also store applications or other information for the mobile computing device 2250. Specifically, the expansion
memory 2274 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 2274 may be provided as a security module for the mobile computing device 2250, and may be programmed with instructions that permit secure use of the mobile computing device 2250. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non- hackable manner.
[0442] The memory may include, for example, flash memory and/or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 2252), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 2264, the expansion memory 2274, or memory on the processor 2252). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 2268 or the external interface 2262.
[0443] The mobile computing device 2250 may communicate wirelessly through the communication interface 2266, which may include digital signal processing circuitry where necessary. The communication interface 2266 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 2268 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 2270 may provide additional navigation- and location-related wireless data to the mobile computing device 2250, which may be used as appropriate by applications running on the mobile computing device 2250.
[0444] The mobile computing device 2250 may also communicate audibly using an audio codec 2260, which may receive spoken information from a user and convert it to usable digital information. The audio codec 2260 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 2250. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 2250.
[0445] The mobile computing device 2250 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 2280. It may also be implemented as part of a smart-phone 2282, personal digital assistant, or other similar mobile device.
[0446] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and/or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and/or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0447] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and/or object-oriented programming language, and/or in assembly/machine language. As used herein, the terms machine -readable medium and computer-readable medium refer to any computer program product, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine- readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and/or data to a programmable processor.
[0448] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0449] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0450] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0451] In some implementations, various modules described herein can be separated, combined or incorporated into single or combined modules. Modules depicted in the figures are not intended to limit the systems described herein to the software architectures shown therein.
EQUIVALENTS
[0452] Elements of different implementations described herein may be combined to form other implementations not specifically set forth above. Elements may be left out of the processes, computer programs, databases, etc. described herein without adversely affecting their operation. In addition, the logic flows depicted in the figures do not require the
particular order shown, or sequential order, to achieve desirable results. Various separate elements may be combined into one or more individual elements to perform the functions described herein.
[0453] Throughout the description, where apparatus and systems are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are apparatus, and systems of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.
[0454] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0455] While the invention has been particularly shown and described with reference to specific preferred embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims.
Claims
1. A method for in-silico design of an engineered antigen, the method comprising:
(a) receiving and/or accessing, by a processor of a computing device, a polypeptide model representing a reference antigen of an infectious agent;
(b) identifying, by the processor, within the polypeptide model, one or more memory-triggering conserved region(s) representing conserved portions of the reference antigen that are determined likely to trigger a memory immune response;
(c) generating, by the processor, one or more amino-acid modifications within at least a portion of the one or more conserved region(s), thereby creating a disrupted polypeptide model representing the engineered antigen; and
(d) storing and/or providing, by the processor, the disrupted polypeptide model for display and/or further processing.
2. The method of claim 1, wherein the reference antigen is or comprises at least a portion of a naturally occurring variant of a viral protein.
3. The method of claim 2, wherein the reference antigen is or comprises at least a portion of a SARS-Cov2 Spike polypeptide.
4. The method of claim 3, wherein the reference antigen is or comprises at least a portion of a particular SARS-CoV-2 variant Spike polypeptide.
5. The method of claim 4, wherein the particular SARS-CoV-2 variant is a member of an Omicron and/or XBB lineage classification.
6. The method of claim 1 or 2, wherein the infectious agent is or comprises an RNA virus and the reference antigen is or comprises at least a portion of a protein thereof.
7. The method of claim 1, wherein the reference antigen is or comprises a bacterial protein.
8. The method of claim 1, wherein the reference antigen is or comprises an antigen of a parasite.
9. The method of any one of the preceding claims, wherein the one or more memorytriggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar to (i) one or more variants thereof and/or (ii) an initial/wild-type strain.
10. The method of any one of the preceding claims, wherein the reference antigen is a particular target SARS-CoV-2 variant S polypeptide or portion thereof and wherein the one or more memory-triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar to corresponding portion(s) of (i) one or more other SARS-CoV- 2 variant polypeptides and/or (ii) Wuhan SARS-CoV-2 polypeptide.
11. The method of any one of the preceding claims, wherein the one or more memorytriggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and un-mutated on the reference antigen.
12. The method of claim 11, wherein step (b) comprises: obtaining, by the processor, data corresponding to a set of known epitopes and identifying, within the reference antigen, each of one or more particular known epitopes of the set;
obtaining, by the processor, an identification of a set of hallmark mutations of the reference antigen; and identifying, by the processor, as the set of conserved epitope regions, those particular known epitopes that correspond to portions of the reference antigen without any hallmark mutations.
13. The method of claim 12, wherein the set of known epitopes comprise one or more of the epitopes listed in Table 2 A.
14. The method of claim 12, wherein the set of known epitopes comprise one or more of the epitopes listed in Table 2B.
15. The method of any one of the preceding claims, wherein the one or more memorytriggering conserved region(s) are or comprise a conserved surface representing a un-mutated surface of the reference antigen.
16. The method of any one of the preceding claims, comprising identifying, and/or accessing an identification of, a set of one or more hallmark mutations of the reference antigen.
17. The method of claim 16, wherein the reference antigen is or comprises a SARS-Cov2 XBB.1.5 Spike protein.
18. The method of any one of the preceding claims, wherein step (c) comprises selecting, by the processor the one or more amino acid modifications from a set of allowed mutations.
19. The method of claim 18, wherein the set of allowed mutations is or comprises a plurality of mutations observed as occurring within a set of related antigens.
20. The method of claim 19, wherein the reference antigen is a SARS-CoV-2 protein of a particular variant and the set of related antigens comprises corresponding proteins of other, related, variants.
21. The method of claim 20, wherein the reference antigen is a member of an Omicron lineage and the set of related antigens comprises observed variants belonging to the Omicron lineage.
22. The method of any one of claims 19-21, wherein the set of allowed mutations comprises at least a portion of mutations listed in Table 3.
23. The method of claim 22, wherein the set of allowed mutations comprises at least a portion of mutations listed in Table 3, excluding one or both of L335F and L390R.
24. The method of any one of claims 19-23, wherein the reference antigen is a SARS- CoV2 protein and the set of related antigens comprises corresponding proteins of other coronavirus.
25. The method of any one of the preceding claims, wherein the one or more memorytriggering conserved regions are or comprise a set of conserved epitope regions and step (c) comprises introducing at least one amino acid modification within each of at least a portion of the conserved epitope regions.
26. The method of any one of the preceding claims, wherein the one or more memorytriggering conserved regions are or comprise a conserved surface and step (c) comprises generating the one or more amino acid modifications at positions distributed throughout I across the conserved surface.
27. The method of any one of the preceding claims, comprising performing steps (b) and (c) repeatedly to generate a plurality of candidate polypeptide models, each representing a candidate engineered variant.
28. The method of claim 27, comprising determining, by the processor, values of one or more performance scores for each of the candidate polypeptide models and selecting a subset of the candidate polypeptide models based at least in part on the determined performance score values.
29. The method of claim 28, wherein the one or more performance scores comprise one or both of:
(a) an immune escape score indicative of a likelihood and/or relative capability of a particular candidate engineered variant to be recognized and neutralized by antibodies, and
(b) a fitness score indicative of a likelihood and/or viability of a particular candidate engineered variant.
30. The method of claim 29, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises using a machine learning model.
31. The method of claim 29 or 30, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises using a 3D structural model of at least a portion of the particular candidate variant.
32. The method of any one of claims 27 to 31, wherein the one or more performance scores comprise(s) a position spread score that measures an extent to which amino acid modifications are evenly distributed across a surface of the candidate engineered variant.
33. The method of any one of claims 27 to 32, wherein the one or more performance scores comprise(s) a mutation co-occurrence score.
34. The method of any one of the preceding claims, comprising identifying, by the processor, within the polypeptide model, one or more target regions representing portions of the reference antigen to be retained and excluding the one or more target regions from the one or more memory-triggering conserved region(s).
35. The method of any one of the preceding claims, comprising causing, by the processor, rendering of the disrupted polypeptide model for graphical display.
36. The method of any one of the preceding claims, comprising generating, from the disrupted polypeptide model, a corresponding RNA sequence.
37. The method of any one of the preceding claims, comprising producing a composition comprising a polypeptide based on the disrupted polypeptide model.
38. The method of any one of the preceding claims, comprising assessing the biological activity of the engineered antigen in vitro.
39. The method of claim 38, wherein the biological activity of the engineered antigen is characterized in that:
the engineered antigen is expressed and folded properly; and/or the engineered antigen does not bind to antibodies that bind to the reference antigen; and/or pseudoviruses loaded with the engineered antigen are able to enter cells; and/or the engineered antigen is immunogenic, and/or the engineered antigen reduces the engineered antigen’s activation of the B cell memory immune response to the reference antigen.
40. The method of any one of the preceding claims, comprising producing a composition comprising a nucleic acid encoding the amino acid sequence represented by the disrupted polypeptide model.
41. A vaccine composition comprising the polypeptide and/or nucleic acid of claims 36 or 37.
42. A method of vaccination comprising administering to a subject or a population of subjects the vaccine of claim 41.
43. A system comprising a processor of a computing device and memory having instructions stored thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 36.
44. A method of manufacturing an immunogenic composition comprising: comparing sequences of viral proteins from different variants of an infectious disease agent to identify residual conserved sites in an antigen of interest; replacing at least one or more of the residual conserved sites with a sequence that is characterized by to generate a new sequence; and
producing a vaccine that delivers at least a portion of the new sequence including at least one of the residual conserved sites replaced.
45. An RNA comprising a nucleotide sequence that encodes an engineered antigen, wherein the engineered antigen corresponds to a particular reference antigen having been altered to introduce one or more amino acid modifications within at least a portion of one or more memory-triggering conserved regions having been identified as portions of the reference antigen that are determined likely to trigger a memory immune response.
46. The RNA of claim 45, wherein the one or more memory-triggering conserved region(s) represent portion(s) of the reference antigen that are substantially similar to one or more variants thereof.
47. The RNA of claim 45 or 46, wherein the one or more memory-triggering conserved region(s) are or comprise a set of conserved epitope regions representing known epitopes that are present and un-mutated on the reference antigen.
48. The RNA of claim 47 wherein the set of known epitopes comprise one or more of the epitopes listed in Tables 2 A and/or 2B.
49. The RNA of any one of claims 45 to 48, wherein the one or more memory-triggering conserved region(s) are or comprise a conserved surface representing a un-mutated surface of the reference antigen.
50. The RNA of claim 49, wherein the reference antigen is or comprises at least a portion of an XBB.l .5 variant of a SARS-Cov2 Spike protein.
51. The RNA of any one of claims 45 to 50, the one or more amino acid modifications are selected from a set of allowed mutations.
52. The RNA of claim 51, wherein the set of allowed mutations is or comprises a plurality of mutations observed as occurring within a set of related polypeptides.
53. The RNA of claim 52, wherein the set of related polypeptides comprises corresponding polypeptides of other, related, SARS-CoV 2 variants.
54. The RNA of claim 53, wherein the target polypeptide is a member of an Omicron lineage and the set of related polypeptides comprises corresponding polypeptides of observed variants belonging to the Omicron lineage.
55. The RNA of any one of claims 53 to 54, wherein the one or more memory-triggering conserved regions are or comprise a set of conserved epitope regions and the engineered antigen has at least one amino acid modification within each of at least a portion of the conserved epitope regions.
56. The RNA of any one of claims 45 to 55, wherein the one or more memory-triggering conserved regions are or comprise a conserved surface and the one or more amino acid modifications occur at positions distributed throughout I across the conserved surface.
57. The method, system, vaccine composition, method of manufacturing, or RNA, of any one of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 5A.
58. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the sequences listed in Table 5B.
59. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 7 A and/or Table 7B.
60. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 8A.
61. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the sequences listed in Table 8B.
62. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 9A.
63. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S43 in Table 9A.
64. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: N360D, P384S, L390R, T430I, F464Y, and H519N.
65. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises the mutation combinations identified as S48 in Table 9A.
66. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N.
67. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises at least a portion of a SARS-CoV-2 S protein with mutations P384S, L390R, T430I, F464Y, and H519N.
68. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises at least a portion of a SARS-CoV-2 S protein with one or more of mutations K356T, P384S, L390R, T430I, F464Y, and H519N.
69. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen does not include one or both of mutations L335F and L390R.
70. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed Table 15A and/or Table 15B.
71. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S 123 in Table 15A and/or Table 15B.
72. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: I332V L335F K356T P384S T430I L452Q F464Y H519N.
73. The method, system, vaccine composition, method of manufacturing, or RNA of claim 72, wherein the engineered antigen comprises at least a portion of the following mutations: I332V L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452Q N460K F464Y S477N T478K E484A F486P F490S Q498R N501 Y Y505H E516Q H519N.
74. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S122 in Table 15A and/or Table 15B.
75. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: L335F K356T P384S T430I L452R F464Y H519N.
76. The method, system, vaccine composition, method of manufacturing, or RNA of claim 75, wherein the engineered antigen comprises at least a portion of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H E516Q H519N T523S.
77. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S 109 in Table 15A and/or Table 15B.
78. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: K356T N360S P384S N388K T430I N450D F464Y H519N.
79. The method, system, vaccine composition, method of manufacturing, or RNA of claim 78, wherein the engineered antigen comprises at least a portion of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S N388K D405N R408S K417N T430I N440K V445P G446S N450D N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N T523S.
80. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S 129 in Table 15A and/or Table 15B.
81. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: K356T L335F P384S D389G T430I N450D F464Y H519N.
82. The method, system, vaccine composition, method of manufacturing, or RNA of claim 81, wherein the engineered antigen comprises at least a portion of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G
D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
83. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S 156 in Table 15A and/or Table 15B.
84. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: K356T P384S L390R T430I N450D F464Y I472V H519N.
85. The method, system, vaccine composition, method of manufacturing, or RNA of claim 84, wherein the engineered antigen comprises at least a portion of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S L390R D405N R408S K417N T430I N440K V445P G446S N450D L452R N460K F464Y I472V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
86. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as SI 12 in Table 15A and/or Table 15B.
87. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: K356T P384S D389G T430I N450D F464Y I468V H519N.
88. The method, system, vaccine composition, method of manufacturing, or RNA of claim 87, wherein the engineered antigen comprises at least a portion of the following mutations: G339H R346T K356T L368I S371F S373P S375F T376A P348S D389G D405N R408S K417N T430I N440K V445P G446S N450D N460K F464Y 1468 V S477N T478K E484A F486P F490S Q498R N501Y Y505H H519N.
89. The method, system, vaccine composition, RNA, or method of manufacturing of any one of the preceding claims, wherein the engineered antigen comprises mutation combinations identified as S125 in Table 15A and/or Table 15B.
90. The method, system, vaccine composition, method of manufacturing, or RNA of any one of the preceding claims, wherein the engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein with at least a portion of the following mutations: K356T L335F P384S T430I F464Y T468V H519N.
91. The method, system, vaccine composition, method of manufacturing, or RNA of claim 90, wherein the engineered antigen comprises at least a portion of the following mutations: L335F G339H R346T K356T L368I S371F S373P S375F T376A P348S D405N R408S K417N T430I N440K V445P G446S N460K F464Y I468V S477N T478K E484A F486P F490S Q498R N501Y Y505H E516Q H519N T523S.
92. A method of manufacturing an RNA comprising a nucleotide sequence that encodes an engineered antigen corresponding to an engineered version of a reference antigen, the method comprising producing an RNA whose nucleotide sequence, when compared with the reference antigen shows difference(s) relative to the reference antigen in one or more memory-triggering conserved regions that are common to the reference antigen and (i) one or more pre-existing variants of the reference antigen and/or (ii) a wild-type strain of the reference antigen.
93. The method of claim 92, wherein the reference antigen is or comprises a particular target variant SARS-CoV-2 S protein RBD.
94. The method of claim 92 or 93, wherein the RNA is or comprises the RNA of any one of claims 45 to 91.
95. The method of any one of claims 92 to 94, comprising producing the RNA via in vitro transcription (IVT).
Applications Claiming Priority (10)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363448215P | 2023-02-24 | 2023-02-24 | |
| US202363448217P | 2023-02-24 | 2023-02-24 | |
| US202363449031P | 2023-02-28 | 2023-02-28 | |
| US202363448987P | 2023-02-28 | 2023-02-28 | |
| US202363449936P | 2023-03-03 | 2023-03-03 | |
| US202363452987P | 2023-03-17 | 2023-03-17 | |
| US202363452989P | 2023-03-17 | 2023-03-17 | |
| US202363514221P | 2023-07-18 | 2023-07-18 | |
| US202363514242P | 2023-07-18 | 2023-07-18 | |
| PCT/US2024/017117 WO2024178355A1 (en) | 2023-02-24 | 2024-02-23 | Systems and methods for engineering synthetic antigens to promote tailored immune responses |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4670167A1 true EP4670167A1 (en) | 2025-12-31 |
Family
ID=90368569
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24713301.0A Pending EP4670167A1 (en) | 2023-02-24 | 2024-02-23 | SYSTEMS AND METHODS FOR MANIPULATING SYNTHETIC ANTIGENS TO PROMOTE TAILORED IMMUNE RESPONSES |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4670167A1 (en) |
| JP (1) | JP2026508268A (en) |
| CN (1) | CN121336262A (en) |
| WO (1) | WO2024178355A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| NZ520445A (en) * | 2000-01-25 | 2004-02-27 | Univ Queensland | Proteins comprising conserved regions of neisseria meningitidis surface antigen NhhA |
| AU2011344126B2 (en) * | 2010-12-13 | 2016-11-10 | The University Of Utah Research Foundation | Vaccine antigens that direct immunity to conserved epitopes |
| WO2021243122A2 (en) | 2020-05-29 | 2021-12-02 | Board Of Regents, The University Of Texas System | Engineered coronavirus spike (s) protein and methods of use thereof |
| AU2022271249A1 (en) | 2021-05-04 | 2023-11-16 | BioNTech SE | Immunogen selection |
-
2024
- 2024-02-23 EP EP24713301.0A patent/EP4670167A1/en active Pending
- 2024-02-23 JP JP2025549694A patent/JP2026508268A/en active Pending
- 2024-02-23 CN CN202480024042.0A patent/CN121336262A/en active Pending
- 2024-02-23 WO PCT/US2024/017117 patent/WO2024178355A1/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| CN121336262A (en) | 2026-01-13 |
| WO2024178355A8 (en) | 2024-10-24 |
| WO2024178355A1 (en) | 2024-08-29 |
| JP2026508268A (en) | 2026-03-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Bagabir et al. | Covid-19 and Artificial Intelligence: Genome sequencing, drug development and vaccine discovery | |
| Bonsignori et al. | Inference of the HIV-1 VRC01 antibody lineage unmutated common ancestor reveals alternative pathways to overcome a key glycan barrier | |
| Pourseif et al. | A domain-based vaccine construct against SARS-CoV-2, the causative agent of COVID-19 pandemic: development of self-amplifying mRNA and peptide vaccines | |
| JP6605643B2 (en) | Anti-dengue virus antibodies and uses thereof | |
| Arimori et al. | Engineering ACE2 decoy receptors to combat viral escapability | |
| US11421018B2 (en) | Full spectrum anti-dengue antibody | |
| Liu et al. | A structure-function analysis shows SARS-CoV-2 BA. 2.86 balances antibody escape and ACE2 affinity | |
| Feng et al. | Structural and functional insights into the evolution of SARS-CoV-2 KP. 3.1. 1 spike protein | |
| Wang et al. | Designed mosaic nanoparticles enhance cross-reactive immune responses in mice | |
| Zhang et al. | SARS-CoV-2 Omicron XBB lineage spike structures, conformations, antigenicity, and receptor recognition | |
| Liguori et al. | NadA3 structures reveal undecad coiled coils and LOX1 binding regions competed by meningococcus B vaccine-elicited human antibodies | |
| Panda et al. | Nanobody-peptide-conjugate (NPC) for passive immunotherapy against SARS-CoV-2 variants of concern (VoC): a prospective pan-coronavirus therapeutics | |
| Hamed et al. | State of the art in epitope mapping and opportunities in COVID-19 | |
| WO2024178355A1 (en) | Systems and methods for engineering synthetic antigens to promote tailored immune responses | |
| Hsueh et al. | Phylodynamic analysis and spike protein mutations in porcine deltacoronavirus with a new variant introduction in Taiwan | |
| EP4690211A1 (en) | Systems and methods for detection, monitoring, and interactive display of circulating infectious diseases and their characteristics | |
| Wei et al. | Design of antibody structure-guided epitope vaccines in silico to induce potent immune responses against emerging viruses | |
| Pourseif et al. | Prophylactic domain-based vaccine against SARS-CoV-2, causative agent of COVID-19 pandemic | |
| Zheng et al. | In silico design of a novel multi-epitope mRNA vaccine candidate for BtHKU5-CoV-2 using immunoinformatics | |
| Ravindran et al. | Immunoinformatic approach to design a vaccine against SARS-COV-2 membrane glycoprotein | |
| Tzarum et al. | Structural and biochemical insights into HCV envelope proteins for germline-targeting vaccines | |
| Rezaei et al. | Immunoinformatics-driven design of a conserved RNA-dependent RNA polymerase-based multi-epitope vaccine against avian infectious bronchitis virus | |
| Biacchesi et al. | In silico reconstruction of a salmonid alphavirus virion reveals distinctive molecular features implicated in virulence | |
| Abduljaleel et al. | Peptides-based vaccine against SARS-nCoV-2 antigenic fragmented synthetic epitopes recognized by T cell and ˇ-cell initiation of specific antibodies to fight the infection | |
| Pourseif et al. | Prophylactic domain-based vaccine against SARS-CoV-2, causative agent of COVID-19 pandemic |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250910 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |