WO2025216910A1 - Selecting a candidate antigen for a vaccine using a biological response machine learning model - Google Patents
Selecting a candidate antigen for a vaccine using a biological response machine learning modelInfo
- Publication number
- WO2025216910A1 WO2025216910A1 PCT/US2025/022331 US2025022331W WO2025216910A1 WO 2025216910 A1 WO2025216910 A1 WO 2025216910A1 US 2025022331 W US2025022331 W US 2025022331W WO 2025216910 A1 WO2025216910 A1 WO 2025216910A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- biological response
- subject
- viral
- amino acid
- immunization
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
- G16H20/10—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to drugs or medications, e.g. for ensuring correct administration to patients
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/20—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/30—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
Definitions
- This specification relates to processing data using machine learning models.
- Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. [0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
- a vaccine includes antigens chosen to trigger the immune system to respond to the disease when a subject is exposed to the pathogen.
- the vaccine can stimulate an adaptive immune response that targets the antigens in the vaccine by producing antibodies that can later recognize the pathogen of a natural infection, thereby providing protection against the disease.
- This specification describes a system implemented as computer programs on one or more computers in one or more locations that can select a candidate antigen for inclusion in a vaccine using a biological response machine learning model.
- the biological response machine learning model can be used to generate a predicted biological response score indicative of the biological response of a subject that has been vaccinated with the candidate antigen against a specific viral strain of a disease.
- a subject refers to an entity, e.g., a cell, or a collection of cells, or animal (e.g., a mouse, rat, dog, pig, chimpanzee, etc.) or a person that can be infected with a disease.
- entity e.g., a cell, or a collection of cells, or animal (e.g., a mouse, rat, dog, pig, chimpanzee, etc.) or a person that can be infected with a disease.
- the subject can be a human participating in a scientific experiment or clinical trial for the development of a particular vaccine for the disease.
- the subject can be a non-human primate, e.g.. chimpanzee, bonobo, macaque, etc., that share genetic and physiological similarities to humans.
- the subject can be a mouse being used to study the immune response to the disease.
- the biological response of the subject to a viral strain can include immune system activation, including antibody production and memory cell formation.
- a biological response score can provide a measure of predicted immune system activation.
- a biological response score can characterize a readout of an immunological assay, e.g., a laboratory procedure used to analyze the presence or concentration of a biological substance.
- the assay can indicate the concentration or amount of a biomolecule associated with infection with the disease, e.g., through the presence of antibodies, antigens, cytokines, nucleic acids, etc. in sera, which can be used as a proxy to assess the level of immunity achieved by vaccinating the subject with the candidate antigen.
- a biological response score can characterize a result of an enzyme-linked immunosorbent assay (ELISA), or a neutralization assay, or a T-cell proliferation assay, and so forth.
- ELISA enzyme-linked immunosorbent assay
- the system can receive a set of immunization history data for a number of subjects, e.g., in a population, and a set of viral strains.
- the immunization history data can provide an immunological record of the subject’s previous exposure to one or more diseases, e.g.. either from a natural infection or via a vaccination.
- the system can augment the immunization history data with one or more candidate antigens and can predict the biological response of the subject for each of the viral strains using the biological response machine learning model.
- the system can process each augmented immunization history with each viral strain to generate a predicted biological response score.
- the system can then select a candidate antigen from the set of candidate antigens for inclusion in a vaccine using the predicted biological response scores.
- the system can filter, aggregate, and rank the biological response scores for each combination of immunization with respect to one or more of the immunization history, candidate antigen, and viral strain to inform the candidate antigen selection.
- a method for receiving immunization history data for a subject wherein the immunization history data defines, for each of a plurality of antigens to w hich the subj ect has been exposed, a respective amino acid sequence associated with the antigen; receiving data defining a viral amino acid sequence of a protein included in a viral strain; processing a model input that includes: (i) the immunization history data, and (ii) the viral amino acid sequence, using a biological response machine learning model and in accordance with values of a set of biological response machine learning model parameters, to generate a predicted biological response score characterizing a predicted biological response of the subject having the immunization history to being exposed to the viral strain; and outputting the predicted biological response score of the subject having the immunization history to being exposed to the viral strain.
- processing the model input that includes: (i) the immunization history data, and (ii) the viral amino acid sequence, using the biological response machine learning model comprises: determining a respective attention score for each of the plurality of antigens included in the immunization history data for the subject, and generating the predicted biological response score of the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history data for the subject.
- determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: determining, for each of the plurality of antigens included in the immunization history data for the subject, the respective attention score for the antigen based on a measure of similarity between: (i) the amino acid sequence associated with the antigen, and (ii) the viral amino acid sequence.
- generating the predicted biological response score for the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history for the subject comprises: selecting an antigen that is associated with a respective attention score from among the plurality' of antigens included in the immunization history data for the subject, processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score.
- processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score comprises: generating one or more embeddings that represent the amino acid sequence associated with the selected antigen and the viral amino acid sequence, and processing the one or more embeddings, in accordance w ith the values of the set of biological response machine learning model parameters, to generate the predicted biological response score.
- processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score comprises: processing the one or more embeddings by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
- determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: generating a respective initial embedding of the respective amino acid sequence of each of the plurality of antigens included in the immunization history data for the subject, generating a respective initial embedding of the viral amino acid sequence, applying a query -key -value (QKV) cross-attention operation to attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the plurality of antigens to generate a combined embedding of the plurality of antigens included in the immunization history data for the subject, and processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subj ect, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score.
- QKV query -key -value
- processing data defining: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subj ect, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score comprises: processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
- processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score comprises: processing: (i) the combined embedding of the plurality of antigens included in the immunization history' data for the subject, and (ii) the initial embedding of the viral amino acid sequence by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
- the method further comprises determining, based at least in part on the predicted biological response score of the subject having the immunization history being exposed to the viral strain, that the subject should receive an additional vaccine to provide protection against the viral strain.
- determining that the subject should receive an additional vaccine to provide protection against the viral strain comprises: determining that the predicted biological response score of the subject is below a threshold.
- the method further comprises administering the additional vaccine to the subject.
- the biological response machine learning model has been trained on a set of training examples by a machine learning training technique, wherein each training example corresponds to a respective training subject and comprises: (i) a training input comprising immunization history data for a training subject and a viral amino acid sequence, and (ii) a target output that defines an actual biological response score of the training subject to being exposed to the viral amino acid sequence, and training the biological response machine learning model to. for each training example, reduce a discrepancy between: (i) a predicted biological response score generated by the biological response machine learning model by processing the training input of the training example, and (ii) the actual biological response score specified by the training example.
- the immunization history data for the subject comprises: (i) one or more antigens associated with vaccines that have been administered to the subject, and (ii) one or more antigens associated with natural infections of the subject.
- the immunization history data for the subject further comprises an indication of whether each antigen in the immune history data is associated with a vaccine or a natural infection and a date at which the subject was exposed to the antigen.
- the immunization history data for the subject further comprises an age of the subj ect.
- a method for selecting a candidate antigen from a set of candidate antigens for inclusion in a vaccine comprising: obtaining data defining a set of immunization histories; obtaining data defining a set of viral strains, generating a plurality of predicted biological response scores, wherein: each predicted biological response score defines a predicted biological response of a subject to being exposed to a corresponding viral strain from the set of viral strains if the subject has a corresponding immunization history from the set of immunization histories and has additionally been vaccinated with a vaccine that includes a corresponding candidate antigen from the set of candidate antigens, and generating each predicted biological response score comprises: generating an augmented immunization history by adding the corresponding candidate antigen to the corresponding immunization history, and processing a model input that includes: (i) a respective amino acid sequence of each antigen included in the augmented immunization history, and (ii) a viral amino acid sequence
- the set of immunization histories comprises a set of immunization histories compiled from a population of a plurality of subjects.
- each of the immunization histories in the set of immunization histories compiled from the population defines a respective sequence of antigens, and wherein each antigen comprises a respective amino acid sequence associated with the respective antigen, to which the subject has been exposed.
- the set of candidate antigens comprises antigens associated with wildtype infections observed to have infected at least one subject of the population.
- the set of candidate antigens comprises antigens associated with wildtype viral strains which have infected the population during a period of time defined with respect to the first infection of a particular wildty pe viral strain.
- the set of candidate antigens comprises antigens associated with wildtype viral strains which have infected the population within a first location during the period of time.
- the set of candidate antigens comprises a set of antigens associated with synthesized viral strains, wherein each synthesized viral strain comprises a viral amino acid sequence including an amino acid at a position in the sequence defined by combinations of amino acids observed in a set of wildtype viral strains at that position.
- the method further comprising determining the set of viral strains as a proper subset of a set of possible viral strains.
- determining the set of viral strains as a proper subset of a set of possible viral strains comprises: determining the set of viral strains as a result of an optimization that encourages: (i) decreased similarity between viral strains in the set of viral strains, and (ii) an increase in similarity between each viral strain in a set of possible viral strains and at least one viral strain in the set of viral strains.
- processing the model input that includes: (i) the respective amino acid sequence of each antigen included in the augmented immunization history, and (ii) the viral amino acid sequence of the corresponding viral strain, using the biological response machine learning model comprises: determining a respective attention score for each of the plurality of antigens included in the immunization history data for the subject, and generating the predicted biological response score of a subject that has the augmented immunization history based at least in part on the attention scores for the plurality of antigens included in the augmented immunization history- data for the subject.
- determining a respective attention score for each of the plurality of antigens included in the augmented immunization history data for the subject comprises: determining, for each of the plurality of antigens included in the augmented immunization history’ data for the subject, the respective attention score for the antigen based on a measure of similarity between: (i) the amino acid sequence associated with the antigen, and (ii) the viral amino acid sequence.
- using the biological response machine learning model to generate the predicted biological response score further comprises: selecting an antigen that is associated with a respective attention score from among the plurality of antigens included in the augmented immunization history data for the subject, processing data defining: (i) the amino acid sequence associated with the selected antigen from the augmented immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score.
- processing data defining: (i) the amino acid sequence associated with the selected antigen from the augmented immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score comprises: generating one or more embeddings that represent the amino acid sequence associated with the selected antigen and the viral amino acid sequence, and processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score.
- determining a respective attention score for each of the plurality of antigens included in the augmented immunization history data for the subject comprising: generating a respective initial embedding of the respective amino acid sequence of each of the plurality of antigens included in the augmented immunization history data for the subject, generating a respective initial embedding of the viral amino acid sequence, applying a query-key-value (QKV) cross-attention operation to attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the plurality of antigens to generate a combined embedding of the plurality' of antigens included in the augmented immunization history' data for the subject, and processing: (i) the combined embedding of the plurality of antigens included in the augmented immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score.
- QKV query-key-value
- the predicted biological response score for each of the viral amino acid sequences comprises a predicted dilution titer readout value determining a measure of a level of protection achieved by a sera inoculation for the augmented immunization history.
- selecting a candidate antigen from the set of candidate antigens for inclusion in a vaccine based at least in part on the plurality of predicted biological response scores comprises: identifying a proper subset of the set of candidate antigens as protection-providing candidate antigens, comprising, for each protection-providing candidate antigen: determining that, for at least a first threshold number of immunization histories augmented to include the protection-providing candidate antigen, the predicted biological response score for at least a second threshold number of viral strains satisfies a threshold score criteria, and selecting the candidate antigen for inclusion in the vaccine from the identified proper subset of protection-providing candidate antigens.
- selecting the candidate antigen for inclusion in the vaccine from the identified proper subset of protection-providing candidate antigens comprises: generating, for each candidate antigen in the identified proper subset of protection-providing candidate antigens, an aggregated biological response score by aggregating the predicted biological response scores associated with the candidate antigen, and selecting the candidate antigen for inclusion in the vaccine based at least in part on the aggregated biological response scores for the candidate antigens included in the identified proper subset of protectionproviding candidate antigens.
- generating the aggregated biological response score for the candidate antigen comprises: generating the aggregated biological response score for the candidate antigen as a linear combination of the predicted biological response scores associated with the candidate antigen.
- generating the aggregated biological response score for the candidate antigen comprises, for each predicted biological response score associated with the candidate antigen: weighting the predicted biological response score based on a likelihood of the immunization history associated with the predicted biological response score, wherein for each immunization history, the likelihood the immunization history is based at least in part on a predicted frequency of occurrence of the immunization history among a population of subjects.
- the method further comprising physically synthesizing the vaccine that includes the selected candidate antigen.
- the method further comprising performing physical experiments to determine one or more properties of the physically synthesized vaccine that includes the selected candidate antigen.
- the method further comprising administering the physically synthesized vaccine to a subject.
- the method further comprising determining that a subject should receive a vaccine that includes the selected candidate antigen.
- the techniques described enable the prediction of a biological response to a disease using a subject’s immunization history' and a viral strain of the disease.
- the ability of the system to process the entire immunization history of the subject presents a more nuanced and detailed view of the adaptive immunity of the subject with respect to the viral strain to the model, thereby increasing the accuracy of the predicted biological response score.
- the system can process the immunization history and viral strain, determine that the predicted biological response score is below a threshold, and recommend administering a vaccine to the subject using the immunization history data.
- the system can flexibly employ machine learning techniques to select candidate antigens from the immunization history for further processing.
- the system can employ either a hard- or soft-attention mechanism to select the candidate antigens.
- the system can employ a computationally-efficient hard-attention mechanism (which can reduce the number of trainable parameters of the system and reduce the amount of training data required for training), or can employ a soft-attention mechanism that allows for adaptively weighting the relative importance of various antigens in the immunization history using continuous-valued attention weights.
- the techniques described enable the selection of a candidate antigen for vaccine development from a set of candidate antigens.
- the system can process a set of immunization histories augmented with one or more candidate antigens and a set of viral strains in order to generate a corresponding number of biological response scores that can be further analyzed for vaccine design.
- the system can aggregate the biological response scores over various population-level characteristics, e.g., the varying immunization histories of subjects in the population, in order to inform the selection of a candidate antigen that can achieve a certain level of protection of the population, e.g., by exceeding a threshold biological response score over a number of the viral strains.
- the techniques of this specification enable an optimized selection of viral strains to test the candidate antigen against, e.g., in order to assure the candidate antigen has been tested against a representative set of viral strains.
- the system can actively leam which regions of viral strain space provide representative viral strains for specific diseases.
- active learning refers to learning from experimental results to identify regions of viral strain space for further training and then obtaining data that relates to the identified regions, e.g.. by suggesting further experimentation that pertains to the identified regions of viral strain space.
- the ability of the system to rely on actively learning the space of viral strains can replace the need for extensive literature searches to identify a representative set of viral strains for a disease.
- using the biological response machine learning model to generate a predicted biological response score can replace the need to perform a large number of assays during early-stage research and development of a vaccine.
- researchers aim to identify and validate potential vaccine targets, e.g., from a set of candidate antigens. In some cases, this involves performing assays for potential candidate antigens to assess the biological response of a subject to a viral strain given a vaccination with the candidate antigen.
- the biological response machine learning model can be used to reduce an amount of the assays that would be needed to fully evaluate the candidate antigens, e.g., the model can predict biological response scores that can be used to funnel the set of candidate antigens down to a smaller number of viral strains for which the researchers can perform physical assays, thereby reducing the time and resources necessary to select a candidate antigen.
- the biological response machine learning model can be used to reject a number of candidate antigens from the set of candidate antigens whose biological response scores fall below a certain threshold value.
- FIG. 1 is a system diagram of an example vaccine design system.
- FIG. 2 illustrates example aggregation methods an example vaccine design system can use to select the candidate antigen using the biological response scores.
- FIG. 3 is a flow diagram of an example process for selecting a candidate antigen using predicted biological response scores.
- FIG. 4 is a flow diagram of an example process for generating a predicted biological response score for a subject.
- FIG. 5 is a flow diagram of an example process for generating a predicted biological response score using a hard-attention mechanism.
- FIG. 6 is a flow diagram of an example process for generating a predicted biological response score using a soft-attention mechanism.
- FIG. 1 shows an example vaccine design system 100.
- the vaccine design system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
- the vaccine design system 100 can be used to select a candidate antigen for inclusion in a vaccine using a biological response machine learning model.
- the system 100 can receive a set of immunization history data for a number of subjects and a set of viral strains and can predict the biological response of each subject for each of the viral strains given the subject was vaccinated with one or more candidate antigens using the biological response machine learning model. More specifically, the system can select a candidate antigen from the set of candidate antigens by predicting the biological responses for each of the immunization histories to viral strains of a particular disease in order to inform vaccine development for the particular disease, e.g., by providing a predicted level of protection against each of the viral strains.
- the vaccine design system 100 can obtain a set of immunization history data 110.
- the immunization history data 110 can include one or more subjects’ immunization histories, e.g., an immunological record of previous exposure to one or more disease, e.g., including a particular disease, as described below, as well as related data, e.g., demographic data or temporal data that can be used to characterize a previous natural infection.
- the system can obtain the set of immunization history data 110 from a biological record database, e.g., an electronic health records database.
- Each immunization history included in the immunization history data 1 10 can provide an immunological record of the subject’s previous exposure to one or more diseases, e.g., either from a natural infection or via a vaccination.
- the immunization history can include a set of antigens, e.g., foreign molecules, e.g.. proteins, peptide chains, carbohydrates, etc. that the subject has been exposed to environmentally or through one or more previous infections by one or more pathogens.
- the immunization history can include a set of antibodies, associated with specific antigens, and developed in response to vaccinations, natural infections, or both, and detected in the subject’s sera or inferred based on a time and place of infection, as will be described in more detail below.
- Each antigen e.g., either detected or inferred based on a time of infection, can be used to distinguish a previous infection with a particular strain of the disease.
- the antigen can be a unique protein or peptide chain composed of a corresponding respective sequence of amino acids, and each sequence of amino acids can be different for each previous infection included in the immunization history data 110.
- the immunization history data 110 can include antigens from previous natural infections, e.g., the natural infection antigens 114, and antigens from vaccinations against the disease, e.g., vaccine antigens 112.
- the sequences of one or all of the antigens included in the immunization history can be determined by performing an assay, e.g., a protein fragment complementation assay, using a specific primer and polymerase chain reaction (PCR) to identify previous exposures to the disease, or with next generation sequencing.
- an assay e.g., a protein fragment complementation assay
- PCR polymerase chain reaction
- the sequences of one or more antigens can be inferred from detecting specific antibodies in a subject’s sera, e.g., as certain antibodies are known to be induced by specific antigens.
- the sequences of one or more antigens can be inferred from sequencing of B-cells, T-cells, or both detected in the subject’s sample.
- the sequences of one or more antigens can be inferred from available information about circulating viral strains at a particular location and during a particular time.
- the sequences of the antigens can be known by virtue of having been developed and administered in a controlled vaccination setting.
- the immunization history data 110 can additionally include an indication of whether or not each antigen in the immunization history is associated with a natural infection, e.g., the natural infection antigens 114, or a vaccination, e.g., the vaccine antigens 112.
- the immunization history data 110 can include a date or estimated timeframe within which the subject was exposed to the antigen.
- the immunization history data 110 can additionally include demographic information about the subject, e.g., age, age range, gender, an indication of whether the subject lives in a rural or urban area, etc.
- the immunization history data 110 can be synthesized from available information.
- a given subject’s immunization history can include known antigens, e.g., vaccine antigens 112 from one or more medical records, and an established consensus of a dominant variants for natural infections based on location and timing.
- the system 100 can evaluate available information about circulating viral strains at a particular location and during a particular time to provide a speculation of the particular antigen corresponding with a natural infection, e.g., which the system 100 can incorporate as part of the natural infection antigens 114.
- the system 100 can compile the immunization histories 110 from a population of subjects.
- the immunization histories 110 can be compiled from a population in a particular location.
- the immunization histories 110 can be randomly sampled from a population across locations at a particular time. More specifically, the system 100 can seek to determine a candidate antigen for a vaccine that targets a particular disease in a particular population of subjects and can assemble the relevant immunization histories 110 of those subjects in order to select an appropriate candidate antigen.
- the system 100 can receive the set of candidate antigens to evaluate for inclusion in a vaccine.
- the set of candidate antigens can be provided as an input to the system 100, e.g., by a user.
- the system 100 can determine the set of candidate antigens as a proper subset of the set of all possible antigens.
- the system 100 can restrict the set of all possible antigens to be antigens associated with the particular disease.
- the system 100 can further restrict the set of antigens associated with the particular disease to be antigens associated with wildtype infections, e.g., antigens physically observed to have previously infected the population.
- the system can restrict the set of antigens to a set with certain immunological importance, e.g.. recent circulating strains, standard of care strains, and variants of concern (VOC).
- VOC variants of concern
- the candidate antigens 132 can include a set of antigens associated with wildtype viral strains, e.g., viral strains observed to have infected at least one subject of the population represented by the immunization history data 110.
- wildtype strains can include strains known to have infected the population during a particular period of time, at a particular location, or both.
- the wildtype strains can include wildtype viral strains which infected the population with respect to the first noted infection of a particular viral strain of a disease.
- the wildtype strains can include wildtype viral strains that were observed after a first infection, e.g., from the original Wuhan strain, the Alpha, Beta, Gamma, Delta wave, etc.
- the wildtype strains can include wildtype viral strains that were observed after a first infection by the H INI or the H3N2 subtype.
- the system 100 can generate the candidate antigens 132 using the wildtype strains of a particular disease.
- the system 100 can analyze each position along the sequence of viral amino acids and compile a set of observed amino acid types at one or more given positions, or at each given position, of the viral amino acid sequence for a particular disease.
- the compiled sets of amino acids for each position can include: (alanine, cysteine, and lysine) at the first position, (arginine and methionine) at the second position, (valine, isoleucine, and proline) at the third position, and (leucine, valine, isoleucine, and proline) at the fourth position.
- the system 100 can generate combinations of candidate antigens using the amino acids observed in the wildtype strains, e.g., the system 100 can generate 72 combinations of amino acids, e.g., as the candidate antigens 132, from the compiled sets for the example case of a four amino acid sequence.
- the system 100 can augment the immunization history data 110 with the set of candidate antigens 132.
- the system 100 can process the set of candidate antigens 132 and the immunization history data 110 using an augmentation engine 130 that can add each candidate antigen to the immunization history data 110, e.g., in order to augment the immunity represented by the vaccine antigens 112 and natural infection antigens 114 included in the immunization history data 110.
- the system 100 can assess the immunity each candidate antigen in the set 132 can confer to the population by augmenting the immunization history data 110.
- the augmentation engine 130 can add the amino acid sequence associated with each candidate antigen in the set 132 to each immunization history 110, e.g.. to generate corresponding augmented immunization history data 134.
- the system 100 can generate N copies of each immunization history' 110, where N is the number of candidate antigens in the set of candidate antigens 132 being evaluated, and include each candidate antigen within a respective immunization history’ to generate the augmented immunization history data 134.
- the augmentation engine 130 can generate N distinct versions of each immunization history included in the data 110, such that each immunization history has been augmented with each of the candidate antigens 132, e.g., has received a synthetic immunity that can be conferred by additionally vaccinating the subject with each of the candidate antigens.
- the immunization history 7 data 110 includes a set of eight immunization histories, e.g.. for eight different subjects, and a set of four candidate antigens
- the augmentation engine 130 can generate a set of 32 different immunization histories, e.g.. four augmented immunization histories for each of the eight subjects, as the augmented immunization history' data 134.
- the system 100 can also obtain a set of viral strains 120 representative of the particular disease.
- Each viral strain can be represented by its respective amino acid sequence, and differences between the sequences can be used to characterize the different viral strains.
- the system 100 can determine the set of viral strains 120 as a result of an optimization process, e.g., a random optimization process performed to determine the set of viral strains 120 from all possible viral strains, e.g., in a viral strain space.
- an optimization process e.g., a random optimization process performed to determine the set of viral strains 120 from all possible viral strains, e.g., in a viral strain space.
- the system 100 can use a simple iterative greedy algorithm to select the set of viral strains 120 from all viral strains to ensure the selected viral strains 120 are dissimilar enough from each other to provide a representative sampling of potential viral strains to test the immunity conferred by each candidate antigen 132.
- the system 100 can perform hierarchical clustering with prototype linkage or archetypal analysis over the viral strain space to characterize the underlying structure of the set of relevant viral strains associated with the particular disease. For example, the system 100 can apply archetypal analysis to select a pool of viral strains that define the boundary 7 of viral strain space, e.g., the boundary relevant to the particular disease that the viral strains can be identified from. In particular, the system 100 can perform a hierarchical clustering to create clusters of viral strains in order to provide a threshold level of coverage of the viral strain space.
- the system 100 can use an optimization algorithm to select viral strains that are far away from each other in accordance with a threshold of goodness, e.g., by optimizing for selecting viral strains that are dissimilar (e.g., sequentially, antigenically. etc.).
- the system 100 can perform clustering to encourage decreased similarity 7 between viral strains in the set of viral strains 120 and an increase in similarity between each viral strain in a set of possible viral strains and at least one viral strain in the set of viral strains.
- the system 100 can receive experimental data to implement an active learning algorithm to iteratively sample the space of available viral strains to reduce the uncertainty of biological response score prediction model 150 in regions of viral strain space with high uncertainty, e.g., by indicating experiments be performed to augment training with respect to the region of viral strain space with high uncertainty. In this way, the viral strains can be selected to optimize learning of the biological response score prediction model 1 0.
- the system 100 can evaluate the identified set of viral strains 120 against a set of determined representative viral strains. The system 100 can select a pool of representative viral strains using a variety of methods. For example, the system 100 can use an additional model or can run a representative set algorithm to identify the representative viral strains.
- the representative set algorithm can be used to select explorative viral strains, e g., in order to learn more about a diverse set of viral strains for a particular disease.
- the system 100 can inform the selection of explorative viral strains using active learning to reduce uncertainty of the model 150, as described above.
- the system 100 can then evaluate the identified set of viral strains against the set of representative viral strains, e.g., to ensure proper coverage of the relevant viral strain space.
- the viral strains selected can be further optimized such that the system 100 does not repeatedly test the augmented immunization histories 134 against the same or similar viral strains, which would be an unnecessary use of computational resources.
- the system 100 can process the augmented immunization history data 134 and the viral strains 120 using a biological response score subsystem 140.
- the system 100 can generate a predicted biological response score 155 indicative of a predicted biological response for a subject being exposed to a corresponding viral strain from the set of viral strains 120 if the subject had been vaccinated with a vaccine that included the candidate antigen from the corresponding candidate antigen used to augment the immunization history.
- the subsystem 140 can process each of the augmented immunization histories and each of the viral strains 120 using a biological response machine learning model 150 to generate a predicted biological response score 155.
- the system can generate 32 augmented immunization histories, which can then be tested against the 72 viral strains to generate 2304 predictions, i.e.. 32 x 72 predictions.
- the biological response machine learning model 150 can have any appropriate machine learning architecture configured to process one or more immunization histories and one or more viral strains to generate a predicted biological response score 155 indicative of an immunity against the viral strain.
- the biological response machine learning model 150 can be implemented in part or in whole as a neural network.
- the neural network can include any appropriate number of neural network layers (e.g., 1 layer. 5 layers, or 10 layers) of any appropriate type (e.g., fully connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers or as a directed graph of layers).
- the biological response machine learning model 150 can include one or more of a random forest model, decision tree model, support vector machine model, linear regression model, or any other regression model that can predict continuous values.
- the biological response score 155 can be determined from an augmented immunization history' and a viral strain to generate a predicted readout of an assay indicating the concentration or amount of a proxy biomolecule associated with infection with the viral strain, e.g., antibodies, antigens, cytokines, proteins, nucleic acids, etc., as the biological response score 155.
- the predicted biological response score 155 can include a predicted dilution titer readout value that provides a measure of a level of protection achieved by vaccination for the augmented immunization history’.
- the biological response machine learning model 150 can be trained with biological experiment data from actual previously performed assays.
- the model 150 can be trained using a machine learning training technique, e.g., to regress to the readout of an experimental assay using a number of training examples.
- the model 150 can be trained by calculating and backpropagating gradients of an objective function to update parameter values of the model, e.g., using the update rule of any appropriate gradient descent optimization algorithm, e.g.. RMSprop or Adam.
- each training example can correspond to a respective training subject and include the subject’s immunization history, a viral amino acid sequence, and a target output that defines the actual biological response of the training subject to being exposed to the viral amino acid sequence, e g., as provided by the assay.
- the biological response machine learning model 150 can be trained using an objective function to reduce a discrepancy between the predicted biological response score generated by the model 150 and the actual biological response specified by the training example.
- the biological response machine learning model 150 can be trained using an LI or L2 loss, mean absolute error, Huber loss, Hinge loss, etc.
- the biological response machine learning model 150 can be trained using data from a serological assay, e.g., an enzyme-linked immunosorbent (ELISA) assay, Western Blot, hemagglutination inhibition assay (HAI) or radioimmunoassay (RIA).
- ELISA enzyme-linked immunosorbent
- HAI hemagglutination inhibition assay
- RIA radioimmunoassay
- the biological response machine learning model can be trained using data from a hematological assay or a cellular immunity assay, e.g., a T-cell proliferation assay or cytokine assay.
- the biological response machine learning model 150 can be trained using a functional assay, e.g., a complement or neutralization assay.
- the biological response score subsystem 140 can process an augmented immunization history and viral strain pair, e.g., a particular (augmented immunization history, viral strain) pair, using the biological response machine learning model 150 to generate a biological response score 155, e.g., the predicted readout response for a subject with the augmented immunization history' to the viral strain.
- the subsystem 140 can repeat processing for each of the viral strains in the set of viral strains 120 and each of the augmented immunization histories in the augmented immunization history 134 to generate respective biological response scores 155, as will be described in further detail below.
- the biological response score subsystem 140 can determine a respective attention score 152 for each of the antigens included in the augmented immunization history as part of processing the augmented immunization history 134 and viral strain pair 120 using the model 150.
- the model 150 can employ either a hard- or soft- attention mechanism to identify the attention score 152 as a measure of relevancy for each of the previous natural infection 112 or vaccination antigens 114 in the augmented immunization history 134.
- the model 150 can distill the information provided by the input data, e.g., the augmented immunization history 134 and the viral strain 120, with respect to the attention score 152 as will be described in further detail below.
- the model 150 can determine a respective attention score 152 for each antigen in the augmented immunization history to focus on particular combinations of antigens from the augmented immunization history 134 and the viral strain 120, e.g., those that are most similar, as well as any additional information that can inform the score 152, e.g., temporal and demographic information that can be used to characterize the response of the augmented immunization history’ to the viral strain 120.
- the model 150 can determine a measure of similarity between the amino acid sequence associated with each antigen in the augmented immunization history' 134 and the viral amino acid sequence 120 using a hard-attention mechanism in order to select promising candidate antigens, e.g., relevant antigens from the augmented immunization history with respect to the viral strain sequence 120 and any associated temporal and demographic data, for further processing.
- the measure of similarity can characterize a distance between the amino acid sequences, e.g., the number of amino acid mismatches, e.g., pointwise across each position.
- the system can use a Hamming distance measurement or BLOSUM score.
- the measure of similarity can take the biophysical and biochemical differences between amino acids into account to ensure that chemically-similar amino acids are considered closer than chemically -dissimilar amino acids.
- the measure of similarity can include information about geometric distances between and mutual orientation of the amino acids in the protein structure.
- the measure of similarity can pertain to the whole amino acid sequence, can focus on more relevant regions of the sequences, e.g., epitopes, or both.
- the model 150 can employ machine learning techniques, e.g.. an attention-mechanism, to focus on combinations of antigens from the augmented immunization history that are similar to the set of viral strains 120, e.g., a representative set of viral strains for the disease, thereby adhering to a common biological hypothesis that immunity' to a specific antigen confers some level of immunity to a similar antigen.
- machine learning techniques e.g.. an attention-mechanism
- the model 150 can employ an attention-mechanism to focus on the earliest exposures in the immunization history, thereby adhering to the concept of original antigenic sin, e.g., that a subject will produce antibodies against a dominant variant from a previous exposure, even when subsequent infection is a result of a new dominant antigen.
- the biological response machine learning model 150 can select an antigen that is associated with a particular attention score, e.g., from the maximum attention score, and can process data defining the amino acid sequence associated with the selected antigen and the viral amino acid sequence associated with the viral strain to generate the predicted biological response score 155.
- the model 150 can generate one or more embeddings representing the amino acid sequence associated with the selected antigen and viral amino acid and process the embeddings using the biological response machine learning model 150 to generate the corresponding biological response score 155.
- the embeddings representing the amino acid sequences can be generated with any kind of embedding method.
- the subsystem 140 can maintain predefined embeddings for each of the 20 amino acids, e.g., using one-hot encoding or physical chemistry-informed embeddings, and can combine each embedding corresponding to each amino acid in the amino acid sequence in the specified order.
- the model 150 can generate the amino acid sequence embeddings in a non-leamed way, e.g., by combining the predefined embeddings using concatenation, summation or averaging, element-wise multiplication, etc.
- the model 150 can generate the amino acid sequence embedding in a learned way, e.g., by passing the respective embeddings through a feedforward neural network and combining the outputs or training another model to adaptively combine embeddings.
- the subsystem 140 can then process the embedding representing the selected antigen amino acid sequence and the embedding representing the viral amino acid sequence using the biological response machine learning model 150.
- the biological response machine learning model 150 can be a feedforward neural network with a number of neural network layers that can process the embeddings to generate the predicted biological response score 155.
- the model 150 can process a concatenation of the selected antigen embedding and the initial viral amino acid sequence embedding.
- the model 150 can process a difference of the selected antigen embedding and the viral amino acid sequence embedding, e.g., by pointwise subtracting the values of the two embeddings or pointwise comparing to indicate which values are different across the two embeddings.
- the model 150 can process a difference of the selected antigen embedding and the viral amino acid sequence embedding and aggregate those differences over regions, e.g., identified important regions, e.g., epitopes.
- the model 150 can determine attention scores 152 and generate a combined embedding of the amino acids in the augmented immunization history using a soft- attention mechanism, which relies on a weighted average of input elements with respect to attention scores.
- the model 150 can generate initial embeddings for each of the antigens in the augmented immunization history' and the viral amino acid sequence of the viral strain and can apply a query-key-value (QKV) cross-attention mechanism to generate a combined embedding of the antigens in the augmented immunization history.
- QKV query-key-value
- the model 150 can derive a set of query, key, and value inputs from the input data, e.g., by projecting the augmented immunization history 7 into key and value spaces and the viral amino acid sequence into a query 7 space using linear transformations.
- the model 150 can then calculate the dot product between the queries, e.g., the viral strain, and the keys, e.g., the key embedding of the augmented immunization history, to represent the similarity between each query and key as the attention score for each query-key.
- the model 150 can attend the augmented immunization history' 134 and viral amino acid sequence 120 to generate a combined embedding of the antigens in the augmented immunization history.
- the model 150 can attend the augmented immunization history and viral amino acid sequence using the first block of a transformer model, e.g., a neural network that uses self-attention mechanisms to capture contextual relationships.
- the model 150 can then process the combined embedding of the antigens of the augmented immunization history data and the initial embedding of the viral amino acid sequence to generate the predicted biological response score 155.
- the model 150 can process the combined embedding and the viral amino acid sequence embedding using a feedforward neural network with a number of neural network layers to generate the predicted biological response score 155.
- the model 150 can process a concatenation of the combined embedding and the initial viral amino acid sequence embedding.
- the model 150 can process a difference of the combined embedding and the ammo acid sequence embedding.
- the subsystem 140 can generate the biological response scores 155 for each augmented immunization history' viral strain pair using the model 150 and further process the scores 155 using an aggregation engine 160 to inform the selection of the candidate antigen 170.
- the aggregation engine 160 can aggregate the biological response scores, e.g.. over viral strain types or immunization histories, in order to determine a subset of candidate antigens.
- the aggregation engine 160 can filter the biological response scores, e.g., to determine a subset of protection-providing candidate antigens, that provide a certain level of protection against each viral strain based on the predicted biological response score satisfying a threshold score criteria.
- the functionality of an example aggregation engine 1 0 will be covered in more detail in FIG. 2.
- the system 100 can select the candidate antigen 170 using the aggregation engine 160. In particular, the system 100 can rank the candidate antigens 170 using the aggregated biological response scores. As an example, the system 100 can select a candidate antigen that provides a certain level of protection for a majority of the viral strains based on the biological response scores 155. As another example, the system 100 can select a candidate antigen that provides protection for a certain number of augmented immunization histories.
- the system 100 can provide the candidate antigen 170 as output, e.g., for use in a laboratory setting, e.g., as part of the drug discovery process.
- one or more researchers can synthesize a vaccine that includes the selected candidate antigen and perform physical experiments to evaluate one or more properties of the vaccine, e.g., the immunogenicity, antigenicity, toxicity, absorption, distribution, metabolism, excretion, etc., in order to vet the vaccine’s actual conferred immunity and any side effects.
- the vaccine can be administered to a subject.
- the system 100 can determine, based in part on the predicted biological response score 155 for a subject, e.g., the score being below a threshold, that the subject should receive an additional vaccine to provide protection against the viral strain.
- FIG. 2 illustrates example aggregation methods that an example vaccine design system can use to select the candidate antigen using the biological response scores.
- the aggregation engine of FIG. 1 can perform the aggregation methods 210 and 260 to inform the selection of the candidate antigen.
- the aggregation engine can first receive and arrange the biological response scores output from the biological response machine learning model, e.g., the biological response machine learning model 150.
- the biological response machine learning model 150 outputs a biological response score for each augmented immunization viral strain pair
- the aggregation engine 160 can populate one or more data structures, e.g., tables, graphs, etc., to organize the biological response scores for one or more of the candidate antigens being evaluated.
- the system has received and processed a set of eight immunization histories and five viral strains to yield 40 biological response scores for each candidate antigen. Since the system 100 augmented the immunization histories with a set of N candidate antigens, the aggregation engine 160 has received 40*N biological response scores, which can be simplified for further processing by being organized and stored in data structures, e.g., data structures organized by candidate antigen.
- the biological response scores are organized in tables by candidate antigen, e.g., the biological response score tables 200, 250, 270. and 290.
- Each table can provide the biological response scores for a particular candidate antigen, e.g., the biological response score table for candidate antigen 1 200, the biological response score table for candidate antigen 2 250, the biological response score table for candidate antigen 3 270, the biological response score table for candidate antigen N 290, etc.
- the table 200 provides the biological response scores for each augmented immunization history that includes the candidate antigen 1 across the set of viral strains.
- the table 250 provides the biological response scores for each augmented immunization history that includes candidate antigen 2 250 across the set of viral strains.
- the aggregation engine 160 can perform a filtering step before aggregation.
- the engine 160 can identify a proper subset of the set of candidate antigens as protection-providing candidate antigens based on the biological response scores associated with each candidate antigen. More specifically, the system can determine that a threshold number of augmented immunization histories including the protection-providing candidate antigen have biological response scores that satisfy a threshold score criteria for at least a threshold number of viral strains.
- the system can determine a subset of the candidate antigens that satisfy a rule, e.g., that at least 5 of the 8 augmented immunization histories including the protection-providing candidate antigen have biological response scores above 60 for 3 of the 5 viral strains.
- the system can calculate a Z-score to determine which candidate antigens provide protection for at least x% of the immunization histories, e.g., 50%, 70%, or 85%.
- removing candidate antigens that do not provide a desired level of protection can streamline the aggregation process by reducing the computational resources necessary to aggregate the biological response scores.
- the aggregation engine 160 can assess the biological response scores by performing one or more aggregation methods in order to evaluate the protection provided by a particular candidate antigen.
- the aggregation engine 160 can perform a viral-aggregation 210 to aggregate the biological response scores for each augmented immunization history that includes the candidate antigen across the set of viral strains. This can provide a measure of the level of protection provided by the candidate antigen for a given augmented immunization history, e.g., the augmented immunization history 1 as is depicted in the viral-aggregation 210 being performed on table 200.
- the engine 160 can compute a viral-aggregated score by aggregating the biological response scores across viral strains for a particular set of augmented immunization histories, e.g., the set of immunization history’ 1 augmented with each of the candidate antigens, and compare the viral-aggregated score to inform selection of the candidate antigen.
- the viral-aggregated scores can provide a measure of the immunity conferred by each of the candidate antigens for the particular immunization history.
- the system can compute viral-aggregated scores for each of the immunization histories corresponding with the particular candidate antigen.
- the system can repeat the viral-aggregation over each candidate antigen data structure to compare viral- aggregated responses per augmented immunization history across the set of candidate antigens to evaluate the level of protection provided by each of the candidate antigens across the viral strains per immunization history type.
- the engine 160 can compute a population-aggregated score by aggregating the biological response scores across augmented immunization histories per viral strain.
- the augmentation engine 160 can aggregate each of the biological response scores for viral strain 2 across the immunization histories that have been augmented with candidate antigen 2.
- the population-aggregated scores can provide a measure of immunity conferred by the candidate antigen to the population, e.g., based on the immunization histories present in the population, against the viral strain.
- the system can compute population-aggregated scores for each of the viral strains in the set of viral strains to evaluate the level of protection provided by the candidate antigen for the population against each viral strain. Furthermore, the system can repeat the population-aggregation over each candidate antigen data structure to compare population-aggregated responses per viral strain across the set of candidate antigens.
- the aggregation for either the viral-aggregation 210 or the population-aggregation method can be calculated a variety' of ways.
- the aggregation can be an average or median.
- the aggregation can be a weighted average, e.g., the aggregation can be calculated as a linear combination of the predicted biological response scores associated with the candidate antigen.
- the weights can be frequency-based, e.g., based on a likelihood of the immunization history' occurring in the population, or the prevalence of a viral strain.
- the system can then rank each aggregated biological response score to select the candidate antigen.
- the system can take the maximum scoring candidate antigen according to either the population-aggregated or viral-aggregated biological response score.
- the system can implement a rules-based system to determine the candidate antigen selection using the aggregated biological response score.
- FIG. 3 is a flow diagram of an example process for selecting a candidate antigen using predicted biological response scores.
- the process 300 will be described as being performed by a system of one or more computers located in one or more locations.
- a vaccine design system e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300.
- the system can obtain a set of immunization histories (step 310).
- Each immunization history’ can include an immunological record of the subject's previous exposure to a particular disease, e.g., either from a natural infection or via a vaccination.
- the immunization history 7 can include antigens from previous natural infections, e.g., the natural infection antigens, and antigens from vaccinations against the disease, e g., vaccine antigens.
- the set of immunization histories can be from a population of subjects.
- the system can obtain a set of viral strains (step 320), e.g.. a set of viral strains representative of the particular disease.
- Each viral strain can be represented by a respective viral amino acid sequence.
- the system can obtain the set of viral strains as a result of an optimization process, e.g., a random optimization, clustering algorithm, or through active learning.
- the system can ensure that the set of viral strains is a representative subset of the set of all possible viral strains, e.g., to ensure proper coverage of the relevant viral strain space for a particular disease.
- the system can then augment each set of immunization histories with a set of candidate antigens (step 330).
- the system can combine the antigens in each immunization history with each candidate antigen to generate a set of augmented immunization histories indicative of if the subject had been vaccinated with a vaccine that included the candidate antigen from the corresponding candidate antigen used to augment the immunization history 7 .
- the system can determine the set of candidate antigens associated with wildtype infections with the disease, e.g.. by selecting the candidate antigens associated with wildtype viral strains which have infected the population at a particular location, during a particular time, or both.
- the system can then predict biological response scores for the set of augmented immunization histories to each of the viral strains (step 340).
- the biological response score can provide a measure of immune system activation, e.g., the biological response score can be a predicted readout of an assay 7 indicating the concentration, potency, or amount of a biomolecule produced in response to infection with the disease given the subject was vaccinated with the candidate antigen in the augmented immunization history.
- the system can predict a biological response score for each augmented immunization history 7 viral strain pair using a biological response machine learning model.
- the system can repeat processing with the biological response machine learning model until each of the viral strains in the set of viral strains has been evaluated against each augmented immunization history 7 .
- the system can determine a respective attention score for the antigens included in the immunization history of the subject as part of using the biological response machine learning model to generate the predicted biological response score of the subject.
- An example of predicting a biological response score using a biological response machine learning model will be covered in more detail in FIG. 4, and example attention mechanisms will be covered in further detail in FIG. 5 and 6.
- the system can then select the candidate antigen using the predicted biological response scores (step 350).
- the system can aggregate the biological response scores, e.g., with respect to the immunization histories and the viral strains, to inform selection of the candidate antigen.
- the system can compute a viral -aggregated score by aggregating the biological response scores across viral strains for a particular set of augmented immunization histories.
- the system can compute a population-aggregated score by aggregating the biological response scores across immunization histories for a particular viral strain.
- the system can then rank the aggregated biological response scores to select the candidate antigen.
- the system can output the candidate antigen for use in a laboratory setting, e.g., as part of the drug discovery process.
- a vaccine that includes the candidate antigen can be synthesized, tested, validated, and administered to subjects.
- the system can additionally filter the set of candidate antigens to determine a subset of protection-providing candidate antigens, e.g., antigens that provide a certain level of protection against each viral strain based on the predicted biological response score satisfying a first threshold score criteria, before performing the aggregation.
- the system can determine one or more protection-providing candidate antigens that confer a threshold level of protection across the immunization histories and viral strains, e.g., the protection-providing candidate antigen can protect a threshold number of immunization histories for at least a second threshold number of viral strains.
- FIG. 4 is a flow diagram of an example process for generating a predicted biological response score for a subject using a biological response machine learning model, e.g., a biological response machine learning model.
- a biological response machine learning model e.g., a biological response machine learning model.
- the process 400 will be described as being performed by a system of one or more computers located in one or more locations.
- a vaccine design system e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
- the process 400 can be performed to predict biological response scores for the set of augmented immunization histories to each of the viral strains, e.g., step 340 in the process 300 depicted in FIG. 3.
- the system can receive immunization history data for a subject (step 410) and data defining a set of viral strains (step 420).
- the immunization history' data can include a set of antigens that the subject has been previously exposed to, e.g., via either natural infection or vaccination.
- the immunization history can be additionally augmented with a set of one or more candidate antigens to synthetically represent the subject having been previously exposed to the candidate antigens.
- the system can then process the immunization history’ and viral strain data using a biological response machine learning model (step 430), e.g., in accordance with values of a set of biological response machine learning model parameters of the biological response machine learning model, to generate a predicted biological response of the subject to the viral strain (step 440). More specifically, the system can compute attention scores to identify a measure of relevancy for each of the previous natural infection or vaccination antigens in the immunization history. The system can then use the attention scores to generate the predicted biological response score.
- the system can compute the set of attention scores using a hard-attention mechanism based on a calculated similarity between the amino acid sequences of the amino acid and the viral amino acid sequences of the viral strain.
- the system can select an antigen associated with a respective attention score, e.g., the highest attention score, and process data defining the amino acid sequence associated with the selected antigen and the viral amino acid sequence to generate the predicted biological response score.
- the system can generate embeddings of the amino acid sequences in the selected antigen and viral strain, e g., by combining a set of predefined embeddings for each of the 20 amino acid sequences and can process the embeddings, e.g., using a feedforward neural network, to generate the predicted biological response score.
- An example of computing the attention scores using a hard-attention mechanism will be covered in FIG. 5.
- the system can compute the attention score using a query-key-value attention mechanism.
- the system can initialize a set of embeddings representing the amino acid sequences in the immunization history and the viral strain and can attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the antigens to generate a combined embedding of the amino acid sequences in the immunization history.
- the system can then process the combined embedding with the viral amino acid sequence embedding, e.g., using a feedforward neural network, to generate the predicted biological response score.
- FIG. 6 An example of computing the attention scores using a soft-attention mechanism will be covered in FIG. 6.
- the system can output the predicted biological response score (step 450), e.g., to characterize the predicted biological response of the subject having the immunization history being exposed to the viral strain.
- the biological response score can be used to determine that the subject should receive an additional vaccine to provide protection against a particular viral strain, e.g., by determining that the predicted biological response score of the subject is below a threshold, and furthermore administering the additional vaccine to the subject.
- the biological response score can be used to determine the most effective candidate antigen for the subject, e.g., given the subject’s immunization history'. The most effective candidate antigen can then be included in a vaccine that is synthesized, tested, validated, and administered to the subject.
- FIG. 5 is a flow diagram of an example process for generating a predicted biological response score using a hard-attention mechanism.
- the process 500 will be described as being performed by a system of one or more computers located in one or more locations.
- a vaccine design system e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.
- the process 500 can be performed to process the immunization history and viral amino acid data using a biological response machine learning model, e.g., step 430 in the process 400 depicted in FIG. 4.
- the system can determine a measure of similarity- between the amino acid sequence associated with each antigen in the immunization history and each viral amino acid sequence in the set of viral strains (step 510).
- the measure of similarity can characterize a distance between the amino acid sequences, e.g., by characterizing the number of amino acid mismatches, e.g., pointwise across each position.
- the system can use a Hamming distance measurement or BLOSUM score.
- the measure of similarity can take the biophysical and biochemical differences between amino acids into account, e.g., to ensure that chemically-similar amino acids are considered closer than chemically-dissimilar amino acids.
- the system can then select the antigen with a particular attention score (step 520) and generate one or more embeddings representing the amino acid sequence of the selected antigen and viral amino acid sequence (step 530). More specifically, the system can select an antigen from the immunization history with the highest measure of similarity' to the viral amino acid sequence and embed each sequence. In some cases, the system can identify and combine predefined embeddings of each of the 20 amino acids in the respective amino acid sequence specified by the antigen and the viral amino acid sequence. In other cases, the system can generate the embedding using an additional model, e.g.. an encoder model.
- an additional model e.g. an encoder model.
- the system can then process the one or more embeddings using a biological response machine learning model (step 540), e.g., the biological response machine learning model, and output a predicted biological response score for the selected antigen and the viral strain (step 550).
- a biological response machine learning model e.g., the biological response machine learning model
- the system can process the embeddings using a feedforward neural network to generate the predicted biological response score.
- FIG. 6 is a flow diagram of an example process for generating a predicted biological response score using a soft-attention mechanism.
- the process 600 will be described as being performed by a system of one or more computers located in one or more locations.
- a vaccine design system e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 600.
- the process 600 can be performed to process the immunization history and viral amino acid data using a biological response machine learning model, e.g., step 430 in the process 400 depicted in FIG. 4.
- the system can generate an initial embedding of the amino acid sequence for each antigen included in the immunization history (step 610) and an initial embedding of each of the viral amino acid sequences included in the viral strain (step 620).
- the system can generate the initial embeddings using an encoder, e.g.. the first block of a transformer.
- the system can then apply a query-key -value cross attention operation (step 630). More specifically, the system can attend the initial embeddings of the viral sequence over the initial embeddings of the amino acid sequences of the antigens to generate a combined embedding of the antigens included in the immunization history (step 640).
- the system can then process the combined embedding and initial viral amino acid embedding using a biological response machine learning model (step 650), e.g., using the biological response machine learning model, to output the predicted biological response score (step 660).
- a biological response machine learning model e.g., using the biological response machine learning model
- the system can process the embeddings using a feedforward neural network to generate the predicted biological response score.
- Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus.
- the computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
- the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
- data processing apparatus refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers.
- the apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
- the apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
- a computer program which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
- a program may, but need not, correspond to a file in a file system.
- a program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code.
- a computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
- engine is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions.
- an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
- the processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output.
- the processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
- Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit.
- a central processing unit will receive instructions and data from a read-only memory or a random access memory or both.
- the essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data.
- the central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
- a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.
- a computer need not have such devices.
- a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e g., a universal serial bus (USB) flash drive, to name just a few.
- PDA personal digital assistant
- GPS Global Positioning System
- USB universal serial bus
- Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way' of example semiconductor memory' devices, e g., EPROM, EEPROM, and flash memory' devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
- semiconductor memory' devices e g., EPROM, EEPROM, and flash memory' devices
- magnetic disks e.g., internal hard disks or removable disks
- magneto-optical disks e.g., CD-ROM and DVD-ROM disks.
- embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g.. a mouse or a trackball, by which the user can provide input to the computer.
- a display device e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor
- keyboard and a pointing device e.g.. a mouse or a trackball
- Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.
- a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser.
- a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
- Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
- Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
- a machine learning framework e.g., a TensorFlow framework, or a Jax framework.
- Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components.
- the components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
- LAN local area network
- WAN wide area network
- the computing system can include clients and servers.
- a client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
- a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client.
- Data generated at the user device e.g., a result of the user interaction, can be received at the server from the device.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Medical Informatics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Theoretical Computer Science (AREA)
- Bioinformatics & Computational Biology (AREA)
- Data Mining & Analysis (AREA)
- Spectroscopy & Molecular Physics (AREA)
- General Health & Medical Sciences (AREA)
- Evolutionary Biology (AREA)
- Biophysics (AREA)
- Biotechnology (AREA)
- Bioethics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Software Systems (AREA)
- Public Health (AREA)
- Evolutionary Computation (AREA)
- Epidemiology (AREA)
- Databases & Information Systems (AREA)
- Artificial Intelligence (AREA)
- Chemical & Material Sciences (AREA)
- Analytical Chemistry (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Medicines Containing Antibodies Or Antigens For Use As Internal Diagnostic Agents (AREA)
- Peptides Or Proteins (AREA)
Abstract
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for selecting a candidate antigen for inclusion in a vaccine using a biological response machine learning model. In one aspect, a method comprises receiving immunization history data for a subject, receiving viral strain data, using a biological response machine learning model to generate a predicted biological response score characterizing a predicted biological response of the subject having the immunization history to being exposed to the viral strain, and outputting the predicted biological response score.
Description
SELECTING A CANDIDATE ANTIGEN FOR A VACCINE USING A BIOLOGICAL
RESPONSE MACHINE LEARNING MODEL
BACKGROUND
[0001] This specification relates to processing data using machine learning models.
[0002] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model. [0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
[0004] This specification also relates to vaccine development for a disease caused by a pathogen. A vaccine includes antigens chosen to trigger the immune system to respond to the disease when a subject is exposed to the pathogen. In particular, the vaccine can stimulate an adaptive immune response that targets the antigens in the vaccine by producing antibodies that can later recognize the pathogen of a natural infection, thereby providing protection against the disease.
SUMMARY
[0005] This specification describes a system implemented as computer programs on one or more computers in one or more locations that can select a candidate antigen for inclusion in a vaccine using a biological response machine learning model. In particular, the biological response machine learning model can be used to generate a predicted biological response score indicative of the biological response of a subject that has been vaccinated with the candidate antigen against a specific viral strain of a disease.
[0006] In this specification, a subject refers to an entity, e.g., a cell, or a collection of cells, or animal (e.g., a mouse, rat, dog, pig, chimpanzee, etc.) or a person that can be infected with a disease.
[0007] In particular, the subject can be a human participating in a scientific experiment or clinical trial for the development of a particular vaccine for the disease. As another example, the subject can be a non-human primate, e.g.. chimpanzee, bonobo, macaque, etc., that share genetic and physiological similarities to humans. As yet another example, the subject can be a mouse being used to study the immune response to the disease.
[0008] The biological response of the subject to a viral strain can include immune system activation, including antibody production and memory cell formation. A biological response score can provide a measure of predicted immune system activation. More specifically, a biological response score can characterize a readout of an immunological assay, e.g., a laboratory procedure used to analyze the presence or concentration of a biological substance. For instance, the assay can indicate the concentration or amount of a biomolecule associated with infection with the disease, e.g., through the presence of antibodies, antigens, cytokines, nucleic acids, etc. in sera, which can be used as a proxy to assess the level of immunity achieved by vaccinating the subject with the candidate antigen. In particular cases, a biological response score can characterize a result of an enzyme-linked immunosorbent assay (ELISA), or a neutralization assay, or a T-cell proliferation assay, and so forth.
[0009] In some cases, the system can receive a set of immunization history data for a number of subjects, e.g., in a population, and a set of viral strains. In particular, the immunization history data can provide an immunological record of the subject’s previous exposure to one or more diseases, e.g.. either from a natural infection or via a vaccination. The system can augment the immunization history data with one or more candidate antigens and can predict the biological response of the subject for each of the viral strains using the biological response machine learning model. In particular, the system can process each augmented immunization history with each viral strain to generate a predicted biological response score. The system can then select a candidate antigen from the set of candidate antigens for inclusion in a vaccine using the predicted biological response scores. As an example, the system can filter, aggregate, and rank the biological response scores for each combination of immunization with respect to one or more of the immunization history, candidate antigen, and viral strain to inform the candidate antigen selection.
[0010] According to a first aspect, there is provided a method for receiving immunization history data for a subject, wherein the immunization history data defines, for each of a plurality of antigens to w hich the subj ect has been exposed, a respective amino acid sequence associated with the antigen; receiving data defining a viral amino acid sequence of a protein included in a viral strain; processing a model input that includes: (i) the immunization history data, and (ii) the viral amino acid sequence, using a biological response machine learning model and in accordance with values of a set of biological response machine learning model parameters, to generate a predicted biological response score characterizing a predicted biological response of the subject having the immunization history to being exposed to the viral strain; and
outputting the predicted biological response score of the subject having the immunization history to being exposed to the viral strain.
[0011] In some implementations, processing the model input that includes: (i) the immunization history data, and (ii) the viral amino acid sequence, using the biological response machine learning model comprises: determining a respective attention score for each of the plurality of antigens included in the immunization history data for the subject, and generating the predicted biological response score of the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history data for the subject.
[0012] In some implementations, determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: determining, for each of the plurality of antigens included in the immunization history data for the subject, the respective attention score for the antigen based on a measure of similarity between: (i) the amino acid sequence associated with the antigen, and (ii) the viral amino acid sequence.
[0013] In some implementations, generating the predicted biological response score for the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history for the subject comprises: selecting an antigen that is associated with a respective attention score from among the plurality' of antigens included in the immunization history data for the subject, processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score.
[0014] In some implementations, processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score comprises: generating one or more embeddings that represent the amino acid sequence associated with the selected antigen and the viral amino acid sequence, and processing the one or more embeddings, in accordance w ith the values of the set of biological response machine learning model parameters, to generate the predicted biological response score.
[0015] In some implementations, processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score comprises: processing the one or more embeddings by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
[0016] In some implementations, determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: generating a respective initial embedding of the respective amino acid sequence of each of the plurality of antigens included in the immunization history data for the subject, generating a respective initial embedding of the viral amino acid sequence, applying a query -key -value (QKV) cross-attention operation to attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the plurality of antigens to generate a combined embedding of the plurality of antigens included in the immunization history data for the subject, and processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subj ect, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score.
[0017] In some implementations, processing data defining: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subj ect, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score comprises: processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
[0018] In some implementations, processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score comprises: processing: (i) the combined embedding of the plurality of antigens included in the immunization history' data for the subject, and (ii) the initial embedding of the viral amino acid sequence by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
[0019] In some implementations, the method further comprises determining, based at least in part on the predicted biological response score of the subject having the immunization history being exposed to the viral strain, that the subject should receive an additional vaccine to provide protection against the viral strain.
[0020] In some implementations, determining that the subject should receive an additional vaccine to provide protection against the viral strain comprises: determining that the predicted biological response score of the subject is below a threshold.
[0021] In some implementations, the method further comprises administering the additional vaccine to the subject.
[0022] In some implementations, the biological response machine learning model has been trained on a set of training examples by a machine learning training technique, wherein each training example corresponds to a respective training subject and comprises: (i) a training input comprising immunization history data for a training subject and a viral amino acid sequence, and (ii) a target output that defines an actual biological response score of the training subject to being exposed to the viral amino acid sequence, and training the biological response machine learning model to. for each training example, reduce a discrepancy between: (i) a predicted biological response score generated by the biological response machine learning model by processing the training input of the training example, and (ii) the actual biological response score specified by the training example.
[0023] In some implementations, the immunization history data for the subject comprises: (i) one or more antigens associated with vaccines that have been administered to the subject, and (ii) one or more antigens associated with natural infections of the subject.
[0024] In some implementations, the immunization history data for the subject further comprises an indication of whether each antigen in the immune history data is associated with a vaccine or a natural infection and a date at which the subject was exposed to the antigen.
[0025] In some implementations, the immunization history data for the subject further comprises an age of the subj ect.
[0026] According to a second aspect, there is provided a method for selecting a candidate antigen from a set of candidate antigens for inclusion in a vaccine, the method comprising: obtaining data defining a set of immunization histories; obtaining data defining a set of viral strains, generating a plurality of predicted biological response scores, wherein: each predicted biological response score defines a predicted biological response of a subject to being exposed to a corresponding viral strain from the set of viral strains if the subject has a corresponding immunization history from the set of immunization histories and has additionally been vaccinated with a vaccine that includes a corresponding candidate antigen from the set of candidate antigens, and generating each predicted biological response score comprises: generating an augmented immunization history by adding the corresponding candidate antigen to the corresponding immunization history, and processing a model input that includes: (i) a respective amino acid sequence of each antigen included in the augmented immunization history, and (ii) a viral amino acid sequence of the corresponding viral strain, using a biological response machine learning model and in accordance with values of a set of biological response machine learning model parameters to generate the predicted biological response score, and
selecting a candidate antigen from the set of candidate antigens for inclusion in a vaccine based at least in part on the plurality of predicted biological response scores.
[0027] In some implementations, the set of immunization histories comprises a set of immunization histories compiled from a population of a plurality of subjects.
[0028] In some implementations, each of the immunization histories in the set of immunization histories compiled from the population defines a respective sequence of antigens, and wherein each antigen comprises a respective amino acid sequence associated with the respective antigen, to which the subject has been exposed.
[0029] In some implementations, the set of candidate antigens comprises antigens associated with wildtype infections observed to have infected at least one subject of the population.
[0030] In some implementations, the set of candidate antigens comprises antigens associated with wildtype viral strains which have infected the population during a period of time defined with respect to the first infection of a particular wildty pe viral strain.
[0031] In some implementations, the set of candidate antigens comprises antigens associated with wildtype viral strains which have infected the population within a first location during the period of time.
[0032] In some implementations, the set of candidate antigens comprises a set of antigens associated with synthesized viral strains, wherein each synthesized viral strain comprises a viral amino acid sequence including an amino acid at a position in the sequence defined by combinations of amino acids observed in a set of wildtype viral strains at that position.
[0033] In some implementations, the method further comprising determining the set of viral strains as a proper subset of a set of possible viral strains.
[0034] In some implementations, determining the set of viral strains as a proper subset of a set of possible viral strains comprises: determining the set of viral strains as a result of an optimization that encourages: (i) decreased similarity between viral strains in the set of viral strains, and (ii) an increase in similarity between each viral strain in a set of possible viral strains and at least one viral strain in the set of viral strains.
[0035] In some implementations, processing the model input that includes: (i) the respective amino acid sequence of each antigen included in the augmented immunization history, and (ii) the viral amino acid sequence of the corresponding viral strain, using the biological response machine learning model comprises: determining a respective attention score for each of the plurality of antigens included in the immunization history data for the subject, and generating the predicted biological response score of a subject that has the augmented immunization
history based at least in part on the attention scores for the plurality of antigens included in the augmented immunization history- data for the subject.
[0036] In some implementations, determining a respective attention score for each of the plurality of antigens included in the augmented immunization history data for the subject comprises: determining, for each of the plurality of antigens included in the augmented immunization history’ data for the subject, the respective attention score for the antigen based on a measure of similarity between: (i) the amino acid sequence associated with the antigen, and (ii) the viral amino acid sequence.
[0037] In some implementations, using the biological response machine learning model to generate the predicted biological response score, further comprises: selecting an antigen that is associated with a respective attention score from among the plurality of antigens included in the augmented immunization history data for the subject, processing data defining: (i) the amino acid sequence associated with the selected antigen from the augmented immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score.
[0038] In some implementations, processing data defining: (i) the amino acid sequence associated with the selected antigen from the augmented immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score comprises: generating one or more embeddings that represent the amino acid sequence associated with the selected antigen and the viral amino acid sequence, and processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score.
[0039] In some implementations, determining a respective attention score for each of the plurality of antigens included in the augmented immunization history data for the subject comprising: generating a respective initial embedding of the respective amino acid sequence of each of the plurality of antigens included in the augmented immunization history data for the subject, generating a respective initial embedding of the viral amino acid sequence, applying a query-key-value (QKV) cross-attention operation to attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the plurality of antigens to generate a combined embedding of the plurality' of antigens included in the augmented immunization history' data for the subject, and processing: (i) the combined embedding of the plurality of antigens included in the augmented immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score.
[0040] In some implementations, the predicted biological response score for each of the viral amino acid sequences comprises a predicted dilution titer readout value determining a measure of a level of protection achieved by a sera inoculation for the augmented immunization history. [0041] In some implementations, selecting a candidate antigen from the set of candidate antigens for inclusion in a vaccine based at least in part on the plurality of predicted biological response scores comprises: identifying a proper subset of the set of candidate antigens as protection-providing candidate antigens, comprising, for each protection-providing candidate antigen: determining that, for at least a first threshold number of immunization histories augmented to include the protection-providing candidate antigen, the predicted biological response score for at least a second threshold number of viral strains satisfies a threshold score criteria, and selecting the candidate antigen for inclusion in the vaccine from the identified proper subset of protection-providing candidate antigens.
[0042] In some implementations, selecting the candidate antigen for inclusion in the vaccine from the identified proper subset of protection-providing candidate antigens comprises: generating, for each candidate antigen in the identified proper subset of protection-providing candidate antigens, an aggregated biological response score by aggregating the predicted biological response scores associated with the candidate antigen, and selecting the candidate antigen for inclusion in the vaccine based at least in part on the aggregated biological response scores for the candidate antigens included in the identified proper subset of protectionproviding candidate antigens.
[0043] In some implementations, for each candidate antigen in the identified proper subset of protection-providing candidate antigens, generating the aggregated biological response score for the candidate antigen comprises: generating the aggregated biological response score for the candidate antigen as a linear combination of the predicted biological response scores associated with the candidate antigen.
[0044] In some implementations, for each candidate antigen in the identified proper subset of protection-providing candidate antigens, generating the aggregated biological response score for the candidate antigen comprises, for each predicted biological response score associated with the candidate antigen: weighting the predicted biological response score based on a likelihood of the immunization history associated with the predicted biological response score, wherein for each immunization history, the likelihood the immunization history is based at least in part on a predicted frequency of occurrence of the immunization history among a population of subjects.
[0045] In some implementations, the method further comprising physically synthesizing the vaccine that includes the selected candidate antigen.
[0046] In some implementations, the method further comprising performing physical experiments to determine one or more properties of the physically synthesized vaccine that includes the selected candidate antigen.
[0047] In some implementations, the method further comprising administering the physically synthesized vaccine to a subject.
[0048] In some implementations, the method further comprising determining that a subject should receive a vaccine that includes the selected candidate antigen.
[0049] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.
[0050] The techniques described enable the prediction of a biological response to a disease using a subject’s immunization history' and a viral strain of the disease. In particular, the ability of the system to process the entire immunization history of the subject presents a more nuanced and detailed view of the adaptive immunity of the subject with respect to the viral strain to the model, thereby increasing the accuracy of the predicted biological response score. Additionally, instead of performing a laboratory procedure, e.g., taking a sample from the subject to evaluate the subject’s immunization history' with respect to the viral strain, the system can process the immunization history and viral strain, determine that the predicted biological response score is below a threshold, and recommend administering a vaccine to the subject using the immunization history data.
[0051] Additionally, the system can flexibly employ machine learning techniques to select candidate antigens from the immunization history for further processing. In particular, the system can employ either a hard- or soft-attention mechanism to select the candidate antigens. More specifically, the system can employ a computationally-efficient hard-attention mechanism (which can reduce the number of trainable parameters of the system and reduce the amount of training data required for training), or can employ a soft-attention mechanism that allows for adaptively weighting the relative importance of various antigens in the immunization history using continuous-valued attention weights.
[0052] Moreover, the techniques described enable the selection of a candidate antigen for vaccine development from a set of candidate antigens. The system can process a set of immunization histories augmented with one or more candidate antigens and a set of viral strains in order to generate a corresponding number of biological response scores that can be further analyzed for vaccine design. In particular, the system can aggregate the biological response
scores over various population-level characteristics, e.g., the varying immunization histories of subjects in the population, in order to inform the selection of a candidate antigen that can achieve a certain level of protection of the population, e.g., by exceeding a threshold biological response score over a number of the viral strains.
[0053] Additionally, the techniques of this specification enable an optimized selection of viral strains to test the candidate antigen against, e.g., in order to assure the candidate antigen has been tested against a representative set of viral strains. More specifically, the system can actively leam which regions of viral strain space provide representative viral strains for specific diseases. In this specification, active learning refers to learning from experimental results to identify regions of viral strain space for further training and then obtaining data that relates to the identified regions, e.g.. by suggesting further experimentation that pertains to the identified regions of viral strain space. The ability of the system to rely on actively learning the space of viral strains can replace the need for extensive literature searches to identify a representative set of viral strains for a disease.
[0054] Furthermore, using the biological response machine learning model to generate a predicted biological response score can replace the need to perform a large number of assays during early-stage research and development of a vaccine. During the discovery stage of drug development, researchers aim to identify and validate potential vaccine targets, e.g., from a set of candidate antigens. In some cases, this involves performing assays for potential candidate antigens to assess the biological response of a subject to a viral strain given a vaccination with the candidate antigen. The biological response machine learning model can be used to reduce an amount of the assays that would be needed to fully evaluate the candidate antigens, e.g., the model can predict biological response scores that can be used to funnel the set of candidate antigens down to a smaller number of viral strains for which the researchers can perform physical assays, thereby reducing the time and resources necessary to select a candidate antigen. In particular, the biological response machine learning model can be used to reject a number of candidate antigens from the set of candidate antigens whose biological response scores fall below a certain threshold value.
[0055] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0056] FIG. 1 is a system diagram of an example vaccine design system.
[0057] FIG. 2 illustrates example aggregation methods an example vaccine design system can use to select the candidate antigen using the biological response scores.
[0058] FIG. 3 is a flow diagram of an example process for selecting a candidate antigen using predicted biological response scores.
[0059] FIG. 4 is a flow diagram of an example process for generating a predicted biological response score for a subject.
[0060] FIG. 5 is a flow diagram of an example process for generating a predicted biological response score using a hard-attention mechanism.
[0061] FIG. 6 is a flow diagram of an example process for generating a predicted biological response score using a soft-attention mechanism.
[0062] Like reference numbers and designations in the various drawings indicate like elements.
DETAILED DESCRIPTION
[0063] FIG. 1 shows an example vaccine design system 100. The vaccine design system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.
|0064| The vaccine design system 100 can be used to select a candidate antigen for inclusion in a vaccine using a biological response machine learning model. In particular, the system 100 can receive a set of immunization history data for a number of subjects and a set of viral strains and can predict the biological response of each subject for each of the viral strains given the subject was vaccinated with one or more candidate antigens using the biological response machine learning model. More specifically, the system can select a candidate antigen from the set of candidate antigens by predicting the biological responses for each of the immunization histories to viral strains of a particular disease in order to inform vaccine development for the particular disease, e.g., by providing a predicted level of protection against each of the viral strains.
[0065] In particular, the vaccine design system 100 can obtain a set of immunization history data 110. The immunization history data 110 can include one or more subjects’ immunization histories, e.g., an immunological record of previous exposure to one or more disease, e.g., including a particular disease, as described below, as well as related data, e.g., demographic data or temporal data that can be used to characterize a previous natural infection. As an
example, the system can obtain the set of immunization history data 110 from a biological record database, e.g., an electronic health records database.
[0066] Each immunization history included in the immunization history data 1 10 can provide an immunological record of the subject’s previous exposure to one or more diseases, e.g., either from a natural infection or via a vaccination. As an example, the immunization history can include a set of antigens, e.g., foreign molecules, e.g.. proteins, peptide chains, carbohydrates, etc. that the subject has been exposed to environmentally or through one or more previous infections by one or more pathogens. In particular, the immunization history can include a set of antibodies, associated with specific antigens, and developed in response to vaccinations, natural infections, or both, and detected in the subject’s sera or inferred based on a time and place of infection, as will be described in more detail below.
[0067] Each antigen, e.g., either detected or inferred based on a time of infection, can be used to distinguish a previous infection with a particular strain of the disease. As an example, the antigen can be a unique protein or peptide chain composed of a corresponding respective sequence of amino acids, and each sequence of amino acids can be different for each previous infection included in the immunization history data 110.
[0068] More specifically, the immunization history data 110 can include antigens from previous natural infections, e.g., the natural infection antigens 114, and antigens from vaccinations against the disease, e.g., vaccine antigens 112. In some cases, the sequences of one or all of the antigens included in the immunization history can be determined by performing an assay, e.g., a protein fragment complementation assay, using a specific primer and polymerase chain reaction (PCR) to identify previous exposures to the disease, or with next generation sequencing.
[0069] In some cases, the sequences of one or more antigens can be inferred from detecting specific antibodies in a subject’s sera, e.g., as certain antibodies are known to be induced by specific antigens. As an example, the sequences of one or more antigens can be inferred from sequencing of B-cells, T-cells, or both detected in the subject’s sample. In some cases, the sequences of one or more antigens can be inferred from available information about circulating viral strains at a particular location and during a particular time. In other cases, e.g.. in the case of known vaccine antigens 112, the sequences of the antigens can be known by virtue of having been developed and administered in a controlled vaccination setting.
[0070] In some cases, the immunization history data 110 can additionally include an indication of whether or not each antigen in the immunization history is associated with a natural infection, e.g., the natural infection antigens 114, or a vaccination, e.g., the vaccine antigens
112. As another example, the immunization history data 110 can include a date or estimated timeframe within which the subject was exposed to the antigen.
[0071] In some cases, the immunization history data 110 can additionally include demographic information about the subject, e.g., age, age range, gender, an indication of whether the subject lives in a rural or urban area, etc.
[0072] In some cases, the immunization history data 110 can be synthesized from available information. For example, a given subject’s immunization history can include known antigens, e.g., vaccine antigens 112 from one or more medical records, and an established consensus of a dominant variants for natural infections based on location and timing. In particular, the system 100 can evaluate available information about circulating viral strains at a particular location and during a particular time to provide a speculation of the particular antigen corresponding with a natural infection, e.g., which the system 100 can incorporate as part of the natural infection antigens 114.
[0073] In some cases, the system 100 can compile the immunization histories 110 from a population of subjects. As an example, the immunization histories 110 can be compiled from a population in a particular location. As another example, the immunization histories 110 can be randomly sampled from a population across locations at a particular time. More specifically, the system 100 can seek to determine a candidate antigen for a vaccine that targets a particular disease in a particular population of subjects and can assemble the relevant immunization histories 110 of those subjects in order to select an appropriate candidate antigen.
[0074] In some cases, the system 100 can receive the set of candidate antigens to evaluate for inclusion in a vaccine. In particular, the set of candidate antigens can be provided as an input to the system 100, e.g., by a user. In other cases, the system 100 can determine the set of candidate antigens as a proper subset of the set of all possible antigens. As an example, the system 100 can restrict the set of all possible antigens to be antigens associated with the particular disease. The system 100 can further restrict the set of antigens associated with the particular disease to be antigens associated with wildtype infections, e.g., antigens physically observed to have previously infected the population. As another example, the system can restrict the set of antigens to a set with certain immunological importance, e.g.. recent circulating strains, standard of care strains, and variants of concern (VOC).
[0075] In some cases, the candidate antigens 132 can include a set of antigens associated with wildtype viral strains, e.g., viral strains observed to have infected at least one subject of the population represented by the immunization history data 110. In particular, the wildtype strains can include strains known to have infected the population during a particular period of time, at
a particular location, or both. For example, the wildtype strains can include wildtype viral strains which infected the population with respect to the first noted infection of a particular viral strain of a disease. In the case of SARS-CoV-2, the wildtype strains can include wildtype viral strains that were observed after a first infection, e.g., from the original Wuhan strain, the Alpha, Beta, Gamma, Delta wave, etc. In the case of Influenza A, the wildtype strains can include wildtype viral strains that were observed after a first infection by the H INI or the H3N2 subtype.
[0076] In some cases, the system 100 can generate the candidate antigens 132 using the wildtype strains of a particular disease. In particular, the system 100 can analyze each position along the sequence of viral amino acids and compile a set of observed amino acid types at one or more given positions, or at each given position, of the viral amino acid sequence for a particular disease. For example, in the case of a viral amino acid sequence composed of four amino acids, the compiled sets of amino acids for each position can include: (alanine, cysteine, and lysine) at the first position, (arginine and methionine) at the second position, (valine, isoleucine, and proline) at the third position, and (leucine, valine, isoleucine, and proline) at the fourth position. In particular, the system 100 can generate combinations of candidate antigens using the amino acids observed in the wildtype strains, e.g., the system 100 can generate 72 combinations of amino acids, e.g., as the candidate antigens 132, from the compiled sets for the example case of a four amino acid sequence.
|0077| The system 100 can augment the immunization history data 110 with the set of candidate antigens 132. In particular, the system 100 can process the set of candidate antigens 132 and the immunization history data 110 using an augmentation engine 130 that can add each candidate antigen to the immunization history data 110, e.g., in order to augment the immunity represented by the vaccine antigens 112 and natural infection antigens 114 included in the immunization history data 110. In particular, the system 100 can assess the immunity each candidate antigen in the set 132 can confer to the population by augmenting the immunization history data 110.
[0078] More specifically, the augmentation engine 130 can add the amino acid sequence associated with each candidate antigen in the set 132 to each immunization history 110, e.g.. to generate corresponding augmented immunization history data 134. For example, the system 100 can generate N copies of each immunization history' 110, where N is the number of candidate antigens in the set of candidate antigens 132 being evaluated, and include each candidate antigen within a respective immunization history’ to generate the augmented immunization history data 134.
[0079] In particular, the augmentation engine 130 can generate N distinct versions of each immunization history included in the data 110, such that each immunization history has been augmented with each of the candidate antigens 132, e.g., has received a synthetic immunity that can be conferred by additionally vaccinating the subject with each of the candidate antigens. In an example in which the immunization history7 data 110 includes a set of eight immunization histories, e.g.. for eight different subjects, and a set of four candidate antigens, the augmentation engine 130 can generate a set of 32 different immunization histories, e.g.. four augmented immunization histories for each of the eight subjects, as the augmented immunization history' data 134.
[0080] The system 100 can also obtain a set of viral strains 120 representative of the particular disease. Each viral strain can be represented by its respective amino acid sequence, and differences between the sequences can be used to characterize the different viral strains.
[0081] In some cases, the system 100 can determine the set of viral strains 120 as a result of an optimization process, e.g., a random optimization process performed to determine the set of viral strains 120 from all possible viral strains, e.g., in a viral strain space. As an example, the system 100 can use a simple iterative greedy algorithm to select the set of viral strains 120 from all viral strains to ensure the selected viral strains 120 are dissimilar enough from each other to provide a representative sampling of potential viral strains to test the immunity conferred by each candidate antigen 132.
|0082| In other cases, the system 100 can perform hierarchical clustering with prototype linkage or archetypal analysis over the viral strain space to characterize the underlying structure of the set of relevant viral strains associated with the particular disease. For example, the system 100 can apply archetypal analysis to select a pool of viral strains that define the boundary7 of viral strain space, e.g., the boundary relevant to the particular disease that the viral strains can be identified from. In particular, the system 100 can perform a hierarchical clustering to create clusters of viral strains in order to provide a threshold level of coverage of the viral strain space. [0083] In yet another case, the system 100 can use an optimization algorithm to select viral strains that are far away from each other in accordance with a threshold of goodness, e.g., by optimizing for selecting viral strains that are dissimilar (e.g., sequentially, antigenically. etc.). In particular the system 100 can perform clustering to encourage decreased similarity7 between viral strains in the set of viral strains 120 and an increase in similarity between each viral strain in a set of possible viral strains and at least one viral strain in the set of viral strains.
[0084] In a further case, the system 100 can receive experimental data to implement an active learning algorithm to iteratively sample the space of available viral strains to reduce the
uncertainty of biological response score prediction model 150 in regions of viral strain space with high uncertainty, e.g., by indicating experiments be performed to augment training with respect to the region of viral strain space with high uncertainty. In this way, the viral strains can be selected to optimize learning of the biological response score prediction model 1 0. [0085] Furthermore, the system 100 can evaluate the identified set of viral strains 120 against a set of determined representative viral strains. The system 100 can select a pool of representative viral strains using a variety of methods. For example, the system 100 can use an additional model or can run a representative set algorithm to identify the representative viral strains. In some cases, the representative set algorithm can be used to select explorative viral strains, e g., in order to learn more about a diverse set of viral strains for a particular disease. In this case, the system 100 can inform the selection of explorative viral strains using active learning to reduce uncertainty of the model 150, as described above.
[0086] The system 100 can then evaluate the identified set of viral strains against the set of representative viral strains, e.g., to ensure proper coverage of the relevant viral strain space. In particular, the viral strains selected can be further optimized such that the system 100 does not repeatedly test the augmented immunization histories 134 against the same or similar viral strains, which would be an unnecessary use of computational resources.
[0087] The system 100 can process the augmented immunization history data 134 and the viral strains 120 using a biological response score subsystem 140. In particular, the system 100 can generate a predicted biological response score 155 indicative of a predicted biological response for a subject being exposed to a corresponding viral strain from the set of viral strains 120 if the subject had been vaccinated with a vaccine that included the candidate antigen from the corresponding candidate antigen used to augment the immunization history. More specifically, the subsystem 140 can process each of the augmented immunization histories and each of the viral strains 120 using a biological response machine learning model 150 to generate a predicted biological response score 155. For example, continuing from the example in paragraph 85 in a case with eight subjects and four candidate antigens, the system can generate 32 augmented immunization histories, which can then be tested against the 72 viral strains to generate 2304 predictions, i.e.. 32 x 72 predictions.
[0088] The biological response machine learning model 150 can have any appropriate machine learning architecture configured to process one or more immunization histories and one or more viral strains to generate a predicted biological response score 155 indicative of an immunity against the viral strain. As an example, the biological response machine learning model 150 can be implemented in part or in whole as a neural network. In the case that the biological
response machine learning model 150 is or includes a neural network, the neural network can include any appropriate number of neural network layers (e.g., 1 layer. 5 layers, or 10 layers) of any appropriate type (e.g., fully connected layers, attention layers, convolutional layers, etc.) connected in any appropriate configuration (e.g., as a linear sequence of layers or as a directed graph of layers). As another example, the biological response machine learning model 150 can include one or more of a random forest model, decision tree model, support vector machine model, linear regression model, or any other regression model that can predict continuous values.
[0089] As an example, the biological response score 155 can be determined from an augmented immunization history' and a viral strain to generate a predicted readout of an assay indicating the concentration or amount of a proxy biomolecule associated with infection with the viral strain, e.g., antibodies, antigens, cytokines, proteins, nucleic acids, etc., as the biological response score 155. In particular, the predicted biological response score 155 can include a predicted dilution titer readout value that provides a measure of a level of protection achieved by vaccination for the augmented immunization history’. In this case, the biological response machine learning model 150 can be trained with biological experiment data from actual previously performed assays.
[0090] The model 150 can be trained using a machine learning training technique, e.g., to regress to the readout of an experimental assay using a number of training examples. In the case that the biological response machine learning model 150 is a neural network, the model 150 can be trained by calculating and backpropagating gradients of an objective function to update parameter values of the model, e.g., using the update rule of any appropriate gradient descent optimization algorithm, e.g.. RMSprop or Adam.
[0091] In particular, each training example can correspond to a respective training subject and include the subject’s immunization history, a viral amino acid sequence, and a target output that defines the actual biological response of the training subject to being exposed to the viral amino acid sequence, e g., as provided by the assay. In this case, the biological response machine learning model 150 can be trained using an objective function to reduce a discrepancy between the predicted biological response score generated by the model 150 and the actual biological response specified by the training example. As an example, the biological response machine learning model 150 can be trained using an LI or L2 loss, mean absolute error, Huber loss, Hinge loss, etc.
[0092] As a particular example, the biological response machine learning model 150 can be trained using data from a serological assay, e.g., an enzyme-linked immunosorbent (ELISA)
assay, Western Blot, hemagglutination inhibition assay (HAI) or radioimmunoassay (RIA). As another example, the biological response machine learning model can be trained using data from a hematological assay or a cellular immunity assay, e.g., a T-cell proliferation assay or cytokine assay. As yet another example, the biological response machine learning model 150 can be trained using a functional assay, e.g., a complement or neutralization assay.
[0093] More specifically, the biological response score subsystem 140 can process an augmented immunization history and viral strain pair, e.g., a particular (augmented immunization history, viral strain) pair, using the biological response machine learning model 150 to generate a biological response score 155, e.g., the predicted readout response for a subject with the augmented immunization history' to the viral strain. The subsystem 140 can repeat processing for each of the viral strains in the set of viral strains 120 and each of the augmented immunization histories in the augmented immunization history 134 to generate respective biological response scores 155, as will be described in further detail below.
[0094] For example, the biological response score subsystem 140 can determine a respective attention score 152 for each of the antigens included in the augmented immunization history as part of processing the augmented immunization history 134 and viral strain pair 120 using the model 150. In particular, the model 150 can employ either a hard- or soft- attention mechanism to identify the attention score 152 as a measure of relevancy for each of the previous natural infection 112 or vaccination antigens 114 in the augmented immunization history 134. More specifically, the model 150 can distill the information provided by the input data, e.g., the augmented immunization history 134 and the viral strain 120, with respect to the attention score 152 as will be described in further detail below.
[0095] In particular, the model 150 can determine a respective attention score 152 for each antigen in the augmented immunization history to focus on particular combinations of antigens from the augmented immunization history 134 and the viral strain 120, e.g., those that are most similar, as well as any additional information that can inform the score 152, e.g., temporal and demographic information that can be used to characterize the response of the augmented immunization history’ to the viral strain 120.
[0096] As an example, the model 150 can determine a measure of similarity between the amino acid sequence associated with each antigen in the augmented immunization history' 134 and the viral amino acid sequence 120 using a hard-attention mechanism in order to select promising candidate antigens, e.g., relevant antigens from the augmented immunization history with respect to the viral strain sequence 120 and any associated temporal and demographic data, for further processing.
[0097] In some cases, the measure of similarity can characterize a distance between the amino acid sequences, e.g., the number of amino acid mismatches, e.g., pointwise across each position. In particular, the system can use a Hamming distance measurement or BLOSUM score. In other cases, the measure of similarity can take the biophysical and biochemical differences between amino acids into account to ensure that chemically-similar amino acids are considered closer than chemically -dissimilar amino acids. In yet another case, the measure of similarity can include information about geometric distances between and mutual orientation of the amino acids in the protein structure. In this case, the measure of similarity can pertain to the whole amino acid sequence, can focus on more relevant regions of the sequences, e.g., epitopes, or both.
[0098] More specifically, the model 150 can employ machine learning techniques, e.g.. an attention-mechanism, to focus on combinations of antigens from the augmented immunization history that are similar to the set of viral strains 120, e.g., a representative set of viral strains for the disease, thereby adhering to a common biological hypothesis that immunity' to a specific antigen confers some level of immunity to a similar antigen. Similarly, the model 150 can employ an attention-mechanism to focus on the earliest exposures in the immunization history, thereby adhering to the concept of original antigenic sin, e.g., that a subject will produce antibodies against a dominant variant from a previous exposure, even when subsequent infection is a result of a new dominant antigen.
|0099| In the case that the biological response machine learning model 150 employs a hard- attention mechanism to calculate the attention score(s) 152, the model 150 can select an antigen that is associated with a particular attention score, e.g., from the maximum attention score, and can process data defining the amino acid sequence associated with the selected antigen and the viral amino acid sequence associated with the viral strain to generate the predicted biological response score 155. In particular, the model 150 can generate one or more embeddings representing the amino acid sequence associated with the selected antigen and viral amino acid and process the embeddings using the biological response machine learning model 150 to generate the corresponding biological response score 155.
[0100] The embeddings representing the amino acid sequences can be generated with any kind of embedding method. As an example, the subsystem 140 can maintain predefined embeddings for each of the 20 amino acids, e.g., using one-hot encoding or physical chemistry-informed embeddings, and can combine each embedding corresponding to each amino acid in the amino acid sequence in the specified order. In some cases, the model 150 can generate the amino acid sequence embeddings in a non-leamed way, e.g., by combining the predefined embeddings
using concatenation, summation or averaging, element-wise multiplication, etc. As another example, the model 150 can generate the amino acid sequence embedding in a learned way, e.g., by passing the respective embeddings through a feedforward neural network and combining the outputs or training another model to adaptively combine embeddings.
[0101] The subsystem 140 can then process the embedding representing the selected antigen amino acid sequence and the embedding representing the viral amino acid sequence using the biological response machine learning model 150. As an example, the biological response machine learning model 150 can be a feedforward neural network with a number of neural network layers that can process the embeddings to generate the predicted biological response score 155. In some cases, the model 150 can process a concatenation of the selected antigen embedding and the initial viral amino acid sequence embedding. In other cases, the model 150 can process a difference of the selected antigen embedding and the viral amino acid sequence embedding, e.g., by pointwise subtracting the values of the two embeddings or pointwise comparing to indicate which values are different across the two embeddings. In even other cases, the model 150 can process a difference of the selected antigen embedding and the viral amino acid sequence embedding and aggregate those differences over regions, e.g., identified important regions, e.g., epitopes.
[0102] As another example, the model 150 can determine attention scores 152 and generate a combined embedding of the amino acids in the augmented immunization history using a soft- attention mechanism, which relies on a weighted average of input elements with respect to attention scores. In this case, the model 150 can generate initial embeddings for each of the antigens in the augmented immunization history' and the viral amino acid sequence of the viral strain and can apply a query-key-value (QKV) cross-attention mechanism to generate a combined embedding of the antigens in the augmented immunization history.
[0103] In particular, the model 150 can derive a set of query, key, and value inputs from the input data, e.g., by projecting the augmented immunization history7 into key and value spaces and the viral amino acid sequence into a query7 space using linear transformations. The model 150 can then calculate the dot product between the queries, e.g., the viral strain, and the keys, e.g., the key embedding of the augmented immunization history, to represent the similarity between each query and key as the attention score for each query-key. The subsystem 140 can then normalize weights across query-keys by applying a softmax function and compute the weighted sum as the attention score: Attention Q, K, T) = V * softmax ^Y where d is the dimensionality of the keys and V is the value embedding of the augmented immunization
history. In particular, the model 150 can attend the augmented immunization history' 134 and viral amino acid sequence 120 to generate a combined embedding of the antigens in the augmented immunization history. In some cases, the model 150 can attend the augmented immunization history and viral amino acid sequence using the first block of a transformer model, e.g., a neural network that uses self-attention mechanisms to capture contextual relationships.
[0104] In the case that the biological response machine learning model 150 employs a soft- attention mechanism to calculate the attention score(s) 152, the model 150 can then process the combined embedding of the antigens of the augmented immunization history data and the initial embedding of the viral amino acid sequence to generate the predicted biological response score 155. In particular, the model 150 can process the combined embedding and the viral amino acid sequence embedding using a feedforward neural network with a number of neural network layers to generate the predicted biological response score 155. In some cases, the model 150 can process a concatenation of the combined embedding and the initial viral amino acid sequence embedding. In other cases, the model 150 can process a difference of the combined embedding and the ammo acid sequence embedding.
[0105] The subsystem 140 can generate the biological response scores 155 for each augmented immunization history' viral strain pair using the model 150 and further process the scores 155 using an aggregation engine 160 to inform the selection of the candidate antigen 170. In particular the aggregation engine 160 can aggregate the biological response scores, e.g.. over viral strain types or immunization histories, in order to determine a subset of candidate antigens. As another example, the aggregation engine 160 can filter the biological response scores, e.g., to determine a subset of protection-providing candidate antigens, that provide a certain level of protection against each viral strain based on the predicted biological response score satisfying a threshold score criteria. The functionality of an example aggregation engine 1 0 will be covered in more detail in FIG. 2.
[0106] The system 100 can select the candidate antigen 170 using the aggregation engine 160. In particular, the system 100 can rank the candidate antigens 170 using the aggregated biological response scores. As an example, the system 100 can select a candidate antigen that provides a certain level of protection for a majority of the viral strains based on the biological response scores 155. As another example, the system 100 can select a candidate antigen that provides protection for a certain number of augmented immunization histories.
[0107] After selecting the candidate antigen 170, the system 100 can provide the candidate antigen 170 as output, e.g., for use in a laboratory setting, e.g., as part of the drug discovery
process. In particular, one or more researchers can synthesize a vaccine that includes the selected candidate antigen and perform physical experiments to evaluate one or more properties of the vaccine, e.g., the immunogenicity, antigenicity, toxicity, absorption, distribution, metabolism, excretion, etc., in order to vet the vaccine’s actual conferred immunity and any side effects.
[0108] Once the vaccine has been properly evaluated, e.g., via physical experiments, clinical trials, or both, the vaccine can be administered to a subject. In some cases, the system 100 can determine, based in part on the predicted biological response score 155 for a subject, e.g., the score being below a threshold, that the subject should receive an additional vaccine to provide protection against the viral strain.
[0109] FIG. 2 illustrates example aggregation methods that an example vaccine design system can use to select the candidate antigen using the biological response scores. As an example, the aggregation engine of FIG. 1 can perform the aggregation methods 210 and 260 to inform the selection of the candidate antigen.
[0110] The aggregation engine can first receive and arrange the biological response scores output from the biological response machine learning model, e.g., the biological response machine learning model 150. In the case that the biological response machine learning model 150 outputs a biological response score for each augmented immunization viral strain pair, the aggregation engine 160 can populate one or more data structures, e.g., tables, graphs, etc., to organize the biological response scores for one or more of the candidate antigens being evaluated.
[oni] In the particular example depicted, the system has received and processed a set of eight immunization histories and five viral strains to yield 40 biological response scores for each candidate antigen. Since the system 100 augmented the immunization histories with a set of N candidate antigens, the aggregation engine 160 has received 40*N biological response scores, which can be simplified for further processing by being organized and stored in data structures, e.g., data structures organized by candidate antigen.
[0112] In the particular example depicted, the biological response scores are organized in tables by candidate antigen, e.g., the biological response score tables 200, 250, 270. and 290. Each table can provide the biological response scores for a particular candidate antigen, e.g., the biological response score table for candidate antigen 1 200, the biological response score table for candidate antigen 2 250, the biological response score table for candidate antigen 3 270, the biological response score table for candidate antigen N 290, etc. As an example, the table 200 provides the biological response scores for each augmented immunization history
that includes the candidate antigen 1 across the set of viral strains. As another example, the table 250 provides the biological response scores for each augmented immunization history that includes candidate antigen 2 250 across the set of viral strains.
[0113] In some cases, the aggregation engine 160 can perform a filtering step before aggregation. In particular, the engine 160 can identify a proper subset of the set of candidate antigens as protection-providing candidate antigens based on the biological response scores associated with each candidate antigen. More specifically, the system can determine that a threshold number of augmented immunization histories including the protection-providing candidate antigen have biological response scores that satisfy a threshold score criteria for at least a threshold number of viral strains. As an example, the system can determine a subset of the candidate antigens that satisfy a rule, e.g., that at least 5 of the 8 augmented immunization histories including the protection-providing candidate antigen have biological response scores above 60 for 3 of the 5 viral strains. As another example, the system can calculate a Z-score to determine which candidate antigens provide protection for at least x% of the immunization histories, e.g., 50%, 70%, or 85%. In particular, removing candidate antigens that do not provide a desired level of protection can streamline the aggregation process by reducing the computational resources necessary to aggregate the biological response scores.
[0114] The aggregation engine 160 can assess the biological response scores by performing one or more aggregation methods in order to evaluate the protection provided by a particular candidate antigen. In particular, the aggregation engine 160 can perform a viral-aggregation 210 to aggregate the biological response scores for each augmented immunization history that includes the candidate antigen across the set of viral strains. This can provide a measure of the level of protection provided by the candidate antigen for a given augmented immunization history, e.g., the augmented immunization history 1 as is depicted in the viral-aggregation 210 being performed on table 200.
[0115] As an example, the engine 160 can compute a viral-aggregated score by aggregating the biological response scores across viral strains for a particular set of augmented immunization histories, e.g., the set of immunization history’ 1 augmented with each of the candidate antigens, and compare the viral-aggregated score to inform selection of the candidate antigen. In this case, the viral-aggregated scores can provide a measure of the immunity conferred by each of the candidate antigens for the particular immunization history. In particular, the system can compute viral-aggregated scores for each of the immunization histories corresponding with the particular candidate antigen. Furthermore, the system can repeat the viral-aggregation over each candidate antigen data structure to compare viral-
aggregated responses per augmented immunization history across the set of candidate antigens to evaluate the level of protection provided by each of the candidate antigens across the viral strains per immunization history type.
[0116] As another example, the engine 160 can compute a population-aggregated score by aggregating the biological response scores across augmented immunization histories per viral strain. In the particular example depicted in the population-aggregation method 260 being performed on table 260. the augmentation engine 160 can aggregate each of the biological response scores for viral strain 2 across the immunization histories that have been augmented with candidate antigen 2. In this case, the population-aggregated scores can provide a measure of immunity conferred by the candidate antigen to the population, e.g., based on the immunization histories present in the population, against the viral strain. In particular, the system can compute population-aggregated scores for each of the viral strains in the set of viral strains to evaluate the level of protection provided by the candidate antigen for the population against each viral strain. Furthermore, the system can repeat the population-aggregation over each candidate antigen data structure to compare population-aggregated responses per viral strain across the set of candidate antigens.
[0117] The aggregation for either the viral-aggregation 210 or the population-aggregation method can be calculated a variety' of ways. In particular, the aggregation can be an average or median. In particular, the aggregation can be a weighted average, e.g., the aggregation can be calculated as a linear combination of the predicted biological response scores associated with the candidate antigen. In this case, the weights can be frequency-based, e.g., based on a likelihood of the immunization history' occurring in the population, or the prevalence of a viral strain.
[0118] The system can then rank each aggregated biological response score to select the candidate antigen. As an example, the system can take the maximum scoring candidate antigen according to either the population-aggregated or viral-aggregated biological response score. As another example, the system can implement a rules-based system to determine the candidate antigen selection using the aggregated biological response score.
[0119] FIG. 3 is a flow diagram of an example process for selecting a candidate antigen using predicted biological response scores. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a vaccine design system, e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300.
[0120] In particular, the system can obtain a set of immunization histories (step 310). Each immunization history’ can include an immunological record of the subject's previous exposure to a particular disease, e.g., either from a natural infection or via a vaccination. In particular, the immunization history7 can include antigens from previous natural infections, e.g., the natural infection antigens, and antigens from vaccinations against the disease, e g., vaccine antigens. In some cases, the set of immunization histories can be from a population of subjects.
[0121] The system can obtain a set of viral strains (step 320), e.g.. a set of viral strains representative of the particular disease. Each viral strain can be represented by a respective viral amino acid sequence. In some cases, the system can obtain the set of viral strains as a result of an optimization process, e.g., a random optimization, clustering algorithm, or through active learning. In particular, the system can ensure that the set of viral strains is a representative subset of the set of all possible viral strains, e.g., to ensure proper coverage of the relevant viral strain space for a particular disease.
[0122] The system can then augment each set of immunization histories with a set of candidate antigens (step 330). In particular, the system can combine the antigens in each immunization history with each candidate antigen to generate a set of augmented immunization histories indicative of if the subject had been vaccinated with a vaccine that included the candidate antigen from the corresponding candidate antigen used to augment the immunization history7. As an example, the system can determine the set of candidate antigens associated with wildtype infections with the disease, e.g.. by selecting the candidate antigens associated with wildtype viral strains which have infected the population at a particular location, during a particular time, or both.
[0123] The system can then predict biological response scores for the set of augmented immunization histories to each of the viral strains (step 340). The biological response score can provide a measure of immune system activation, e.g., the biological response score can be a predicted readout of an assay7 indicating the concentration, potency, or amount of a biomolecule produced in response to infection with the disease given the subject was vaccinated with the candidate antigen in the augmented immunization history.
[0124] In particular, the system can predict a biological response score for each augmented immunization history7 viral strain pair using a biological response machine learning model. The system can repeat processing with the biological response machine learning model until each of the viral strains in the set of viral strains has been evaluated against each augmented immunization history7. More specifically, the system can determine a respective attention score for the antigens included in the immunization history of the subject as part of using the
biological response machine learning model to generate the predicted biological response score of the subject. An example of predicting a biological response score using a biological response machine learning model will be covered in more detail in FIG. 4, and example attention mechanisms will be covered in further detail in FIG. 5 and 6.
[0125] The system can then select the candidate antigen using the predicted biological response scores (step 350). In particular, the system can aggregate the biological response scores, e.g., with respect to the immunization histories and the viral strains, to inform selection of the candidate antigen. As an example, the system can compute a viral -aggregated score by aggregating the biological response scores across viral strains for a particular set of augmented immunization histories. As another example, the system can compute a population-aggregated score by aggregating the biological response scores across immunization histories for a particular viral strain. The system can then rank the aggregated biological response scores to select the candidate antigen. After selecting the candidate antigen, the system can output the candidate antigen for use in a laboratory setting, e.g., as part of the drug discovery process. In particular, a vaccine that includes the candidate antigen can be synthesized, tested, validated, and administered to subjects.
[0126] In some cases, the system can additionally filter the set of candidate antigens to determine a subset of protection-providing candidate antigens, e.g., antigens that provide a certain level of protection against each viral strain based on the predicted biological response score satisfying a first threshold score criteria, before performing the aggregation. In particular, the system can determine one or more protection-providing candidate antigens that confer a threshold level of protection across the immunization histories and viral strains, e.g., the protection-providing candidate antigen can protect a threshold number of immunization histories for at least a second threshold number of viral strains.
[0127] FIG. 4 is a flow diagram of an example process for generating a predicted biological response score for a subject using a biological response machine learning model, e.g., a biological response machine learning model. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a vaccine design system, e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.
[0128] The process 400 can be performed to predict biological response scores for the set of augmented immunization histories to each of the viral strains, e.g., step 340 in the process 300 depicted in FIG. 3. In particular, the system can receive immunization history data for a subject
(step 410) and data defining a set of viral strains (step 420). The immunization history' data can include a set of antigens that the subject has been previously exposed to, e.g., via either natural infection or vaccination. In some cases, the immunization history can be additionally augmented with a set of one or more candidate antigens to synthetically represent the subject having been previously exposed to the candidate antigens.
[0129] The system can then process the immunization history’ and viral strain data using a biological response machine learning model (step 430), e.g., in accordance with values of a set of biological response machine learning model parameters of the biological response machine learning model, to generate a predicted biological response of the subject to the viral strain (step 440). More specifically, the system can compute attention scores to identify a measure of relevancy for each of the previous natural infection or vaccination antigens in the immunization history. The system can then use the attention scores to generate the predicted biological response score.
[0130] In some cases, the system can compute the set of attention scores using a hard-attention mechanism based on a calculated similarity between the amino acid sequences of the amino acid and the viral amino acid sequences of the viral strain. In the case of employing a hard- attention mechanism to generate the attention score, the system can select an antigen associated with a respective attention score, e.g., the highest attention score, and process data defining the amino acid sequence associated with the selected antigen and the viral amino acid sequence to generate the predicted biological response score. In particular, the system can generate embeddings of the amino acid sequences in the selected antigen and viral strain, e g., by combining a set of predefined embeddings for each of the 20 amino acid sequences and can process the embeddings, e.g., using a feedforward neural network, to generate the predicted biological response score. An example of computing the attention scores using a hard-attention mechanism will be covered in FIG. 5.
[0131] In the case of employing a soft-attention mechanism, the system can compute the attention score using a query-key-value attention mechanism. In particular, the system can initialize a set of embeddings representing the amino acid sequences in the immunization history and the viral strain and can attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the antigens to generate a combined embedding of the amino acid sequences in the immunization history. The system can then process the combined embedding with the viral amino acid sequence embedding, e.g., using a feedforward neural network, to generate the predicted biological response score. An example of computing the attention scores using a soft-attention mechanism will be covered in FIG. 6.
[0132] The system can output the predicted biological response score (step 450), e.g., to characterize the predicted biological response of the subject having the immunization history being exposed to the viral strain. As an example, the biological response score can be used to determine that the subject should receive an additional vaccine to provide protection against a particular viral strain, e.g., by determining that the predicted biological response score of the subject is below a threshold, and furthermore administering the additional vaccine to the subject. As another example, in the case that the immunization history has been augmented with a set of one or more candidate antigens, the biological response score can be used to determine the most effective candidate antigen for the subject, e.g., given the subject’s immunization history'. The most effective candidate antigen can then be included in a vaccine that is synthesized, tested, validated, and administered to the subject.
[0133] FIG. 5 is a flow diagram of an example process for generating a predicted biological response score using a hard-attention mechanism. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a vaccine design system, e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500. The process 500 can be performed to process the immunization history and viral amino acid data using a biological response machine learning model, e.g., step 430 in the process 400 depicted in FIG. 4.
|0134] In particular, the system can determine a measure of similarity- between the amino acid sequence associated with each antigen in the immunization history and each viral amino acid sequence in the set of viral strains (step 510). As an example, the measure of similarity can characterize a distance between the amino acid sequences, e.g., by characterizing the number of amino acid mismatches, e.g., pointwise across each position. In particular, the system can use a Hamming distance measurement or BLOSUM score. As another example, the measure of similarity can take the biophysical and biochemical differences between amino acids into account, e.g., to ensure that chemically-similar amino acids are considered closer than chemically-dissimilar amino acids.
[0135] The system can then select the antigen with a particular attention score (step 520) and generate one or more embeddings representing the amino acid sequence of the selected antigen and viral amino acid sequence (step 530). More specifically, the system can select an antigen from the immunization history with the highest measure of similarity' to the viral amino acid sequence and embed each sequence. In some cases, the system can identify and combine predefined embeddings of each of the 20 amino acids in the respective amino acid sequence
specified by the antigen and the viral amino acid sequence. In other cases, the system can generate the embedding using an additional model, e.g.. an encoder model.
[0136] The system can then process the one or more embeddings using a biological response machine learning model (step 540), e.g., the biological response machine learning model, and output a predicted biological response score for the selected antigen and the viral strain (step 550). In particular, the system can process the embeddings using a feedforward neural network to generate the predicted biological response score.
[0137] FIG. 6 is a flow diagram of an example process for generating a predicted biological response score using a soft-attention mechanism. For convenience, the process 600 will be described as being performed by a system of one or more computers located in one or more locations. For example, a vaccine design system, e.g., the vaccine design system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 600. The process 600 can be performed to process the immunization history and viral amino acid data using a biological response machine learning model, e.g., step 430 in the process 400 depicted in FIG. 4.
[0138] In particular, the system can generate an initial embedding of the amino acid sequence for each antigen included in the immunization history (step 610) and an initial embedding of each of the viral amino acid sequences included in the viral strain (step 620). As an example, the system can generate the initial embeddings using an encoder, e.g.. the first block of a transformer.
[0139] The system can then apply a query-key -value cross attention operation (step 630). More specifically, the system can attend the initial embeddings of the viral sequence over the initial embeddings of the amino acid sequences of the antigens to generate a combined embedding of the antigens included in the immunization history (step 640).
[0140] The system can then process the combined embedding and initial viral amino acid embedding using a biological response machine learning model (step 650), e.g., using the biological response machine learning model, to output the predicted biological response score (step 660). In particular, the system can process the embeddings using a feedforward neural network to generate the predicted biological response score.
[0141] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular
operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. [0142] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0143] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0144] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules,
sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0145] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0146] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0147] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e g., a universal serial bus (USB) flash drive, to name just a few.
[0148] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way' of example semiconductor memory' devices, e g., EPROM, EEPROM, and flash memory' devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0149] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT
(cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g.. a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0150] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.
[0151] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, or a Jax framework.
[0152] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0153] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which
acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0154] While this specification contains many specific implementation details, these should not be constmed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0155] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0156] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0157] What is claimed is:
Claims
1. A method performed by one or more computers, the method comprising: receiving immunization history' data for a subject, wherein the immunization history data defines, for each of a plurality of antigens to which the subject has been exposed, a respective amino acid sequence associated with the antigen; receiving data defining a viral amino acid sequence of a protein included in a viral strain; processing a model input that includes: (i) the immunization history data, and (ii) the viral amino acid sequence, using a biological response machine learning model and in accordance with values of a set of biological response machine learning model parameters, to generate a predicted biological response score characterizing a predicted biological response of the subject having the immunization history' to being exposed to the viral strain; and outputting the predicted biological response score of the subject having the immunization history to being exposed to the viral strain.
2. The method of claim 1, wherein processing the model input that includes: (i) the immunization history’ data, and (ii) the viral amino acid sequence, using the biological response machine learning model comprises: determining a respective attention score for each of the plurality of antigens included in the immunization history' data for the subject; and generating the predicted biological response score of the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history’ data for the subject.
3. The method of claim 2, wherein determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: determining, for each of the plurality' of antigens included in the immunization history data for the subject, the respective attention score for the antigen based on a measure of similarity between: (i) the amino acid sequence associated with the antigen, and (ii) the viral amino acid sequence.
4. The method of any one of claims 2-3. wherein generating the predicted biological response score for the subject based at least in part on the attention scores for the plurality of antigens included in the immunization history for the subject comprises:
selecting an antigen that is associated with a respective attention score from among the plurality of antigens included in the immunization history data for the subject; processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization his ton data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score.
5. The method of claim 4, wherein processing data defining: (i) the amino acid sequence associated with the selected antigen from the immunization history data for the subject, and (ii) the viral amino acid sequence, to generate the predicted biological response score comprises: generating one or more embeddings that represent the amino acid sequence associated with the selected antigen and the viral amino acid sequence; and processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score.
6. The method of claim 5. wherein processing the one or more embeddings, in accordance with the values of the set of biological response machine learning model parameters, to generate the predicted biological response score comprises: processing the one or more embeddings by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
7. The method of claim 2, wherein determining the respective attention score for each of the plurality of antigens included in the immunization history data for the subject comprises: generating a respective initial embedding of the respective amino acid sequence of each of the plurality of antigens included in the immunization history data for the subject; generating a respective initial embedding of the viral amino acid sequence; applying a query-key-value (QKV) cross-attention operation to attend the initial embedding of the viral sequence over the initial embeddings of the amino acid sequences of the plurality of antigens to generate a combined embedding of the plurality of antigens included in the immunization history data for the subject; and processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score.
8. The method of claim 7, wherein processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence, to generate the predicted biological response score comprises: processing: (i) the combined embedding of the plurality of antigens included in the immunization history data for the subject, and (ii) the initial embedding of the viral amino acid sequence by a plurality of neural network layers of the biological response machine learning model to generate the predicted biological response score.
9. The method of any preceding claim, further comprising determining, based at least in part on the predicted biological response score of the subject having the immunization history being exposed to the viral strain, that the subject should receive an additional vaccine to provide protection against the viral strain.
10. The method of claim 9, wherein determining that the subject should receive an additional vaccine to provide protection against the viral strain comprises: determining that the predicted biological response score of the subject is below a threshold.
1 1 . The method of any one of claims 9-10, further comprising administering the additional vaccine to the subject.
12. The method of any preceding claim, wherein the biological response machine learning model has been trained on a set of training examples by a machine learning training technique; wherein each training example corresponds to a respective training subject and comprises: (i) a training input comprising immunization history data for a training subject and a viral amino acid sequence, and (ii) a target output that defines an actual biological response score of the training subject to being exposed to the viral amino acid sequence; and training the biological response machine learning model to, for each training example, reduce a discrepancy between: (i) a predicted biological response score generated by the biological response machine learning model by processing the training input of the training example, and (ii) the actual biological response score specified by the training example.
13. The method of any preceding claim, wherein the immunization history data for the subject comprises: (i) one or more antigens associated with vaccines that have been administered to the subject, and (ii) one or more antigens associated with natural infections of the subj ect.
14. The method of claim 13, wherein the immunization history data for the subject further comprises an indication of whether each antigen in the immune history data is associated with a vaccine or a natural infection and a date at which the subject was exposed to the antigen.
15. The method of claim 14, wherein the immunization history data for the subject further comprises an age of the subject.
16. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-10 or 12-15.
17. A sy stem compri sing : one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-10 or 12-15.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463631175P | 2024-04-08 | 2024-04-08 | |
| US63/631,175 | 2024-04-08 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025216910A1 true WO2025216910A1 (en) | 2025-10-16 |
Family
ID=95446593
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/022331 Pending WO2025216910A1 (en) | 2024-04-08 | 2025-03-31 | Selecting a candidate antigen for a vaccine using a biological response machine learning model |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025216910A1 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210118575A1 (en) * | 2019-10-21 | 2021-04-22 | Sanofi Pasteur, Inc. | Systems and Methods for Designing Vaccines |
| US20220122690A1 (en) * | 2020-07-17 | 2022-04-21 | Genentech, Inc. | Attention-based neural network to predict peptide binding, presentation, and immunogenicity |
| WO2022175683A1 (en) * | 2021-02-22 | 2022-08-25 | Fluidic Analytics Limited | Improvements in or relating to immunity profiling |
-
2025
- 2025-03-31 WO PCT/US2025/022331 patent/WO2025216910A1/en active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210118575A1 (en) * | 2019-10-21 | 2021-04-22 | Sanofi Pasteur, Inc. | Systems and Methods for Designing Vaccines |
| US20220122690A1 (en) * | 2020-07-17 | 2022-04-21 | Genentech, Inc. | Attention-based neural network to predict peptide binding, presentation, and immunogenicity |
| WO2022175683A1 (en) * | 2021-02-22 | 2022-08-25 | Fluidic Analytics Limited | Improvements in or relating to immunity profiling |
Non-Patent Citations (1)
| Title |
|---|
| YONGQUN HE ET AL: "Emerging Vaccine Informatics", JOURNAL OF BIOMEDICINE AND BIOTECHNOLOGY, vol. 2010, 1 January 2010 (2010-01-01), pages 1 - 26, XP055002103, ISSN: 1110-7243, DOI: 10.1155/2010/218590 * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Yao et al. | SVMTriP: a method to predict antigenic epitopes using support vector machine to integrate tri-peptide similarity and propensity | |
| US20250308710A1 (en) | Systems and Methods for Designing Vaccines | |
| Lian et al. | EPMLR: sequence-based linear B-cell epitope prediction method using multiple linear regression | |
| Ito et al. | A protein language model for exploring viral fitness landscapes | |
| Zhou et al. | TEMPO: A transformer-based mutation prediction framework for SARS-CoV-2 evolution | |
| Quadeer et al. | Identifying immunologically-vulnerable regions of the HCV E2 glycoprotein and broadly neutralizing antibodies that target them | |
| Vaishnav et al. | Overview of computational vaccinology: vaccine development through information technology | |
| US20240177799A1 (en) | T-cell receptor optimization with reinforcement learning and mutation policies for precision immunotherapy | |
| Gokhale et al. | Disentangling the causes of mumps reemergence in the United States | |
| Gardner et al. | Design, statistical analysis and reporting standards for test accuracy studies for infectious diseases in animals: progress, challenges and recommendations | |
| Ronel et al. | The clonal structure and dynamics of the human T cell response to an organic chemical hapten | |
| US11011253B1 (en) | Escape profiling for therapeutic and vaccine development | |
| Charoenkwan et al. | TROLLOPE: A novel sequence-based stacked approach for the accelerated discovery of linear T-cell epitopes of hepatitis C virus | |
| Bing et al. | Comparison of empirical and dynamic models for HIV viral load rebound after treatment interruption | |
| Corbeil | Iryonlp at mediqa-corr 2024: Tackling the medical error detection & correction task on the shoulders of medical agents | |
| Barkan et al. | Leveraging large language models to predict antibody biological activity against influenza A hemagglutinin | |
| Huang et al. | Prediction of linear B-cell epitopes of hepatitis C virus for vaccine development | |
| WO2025216910A1 (en) | Selecting a candidate antigen for a vaccine using a biological response machine learning model | |
| WO2025216911A1 (en) | Selecting a candidate antigen for a vaccine using a biological response machine learning model | |
| Liu et al. | PLM-IL4: Enhancing IL-4-inducing peptide prediction with protein language model | |
| Kain et al. | Rethinking statistical approaches for serological data analysis for viral surveillance | |
| Bukhari et al. | Prediction of antigenic peptides of SARS-CoV-2 pathogen using machine learning | |
| Tasnim et al. | Next mutation prediction of sars-cov-2 spike protein sequence using encoder-decoder based long short term memory (lstm) method | |
| Li et al. | ctP 2 ISP: Protein–Protein Interaction Sites Prediction Using Convolution and Transformer With Data Augmentation | |
| Slabodkin et al. | Weakly supervised identification and generation of adaptive immune receptor sequences associated with immune disease status |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25719593 Country of ref document: EP Kind code of ref document: A1 |