WO2025199168A1 - MACHINE LEARNING-BASED METHODS FOR MODELING pMHC CONFORMERS - Google Patents

MACHINE LEARNING-BASED METHODS FOR MODELING pMHC CONFORMERS

Info

Publication number
WO2025199168A1
WO2025199168A1 PCT/US2025/020470 US2025020470W WO2025199168A1 WO 2025199168 A1 WO2025199168 A1 WO 2025199168A1 US 2025020470 W US2025020470 W US 2025020470W WO 2025199168 A1 WO2025199168 A1 WO 2025199168A1
Authority
WO
WIPO (PCT)
Prior art keywords
peptide
protein
data
trained
predicted
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/020470
Other languages
French (fr)
Inventor
Andrew Martin WATKINS
Santrupti NERLI
Ji Won Park
Changpeng LU
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Genentech Inc
Original Assignee
Genentech Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Genentech Inc filed Critical Genentech Inc
Publication of WO2025199168A1 publication Critical patent/WO2025199168A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B15/00ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
    • G16B15/30Drug targeting using structural data; Docking or binding prediction
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • G16B40/20Supervised data analysis
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • G16B40/30Unsupervised data analysis

Definitions

  • FIELD FIELD
  • This disclosure relates generally to machine learning-based methods for modeling conformational states of peptide-protein complexes that account for the dynamic nature of three-dimensional peptide and protein structures, and that may be used to predict features of the peptide-protein complexes.
  • the present disclosure relates more specifically to the application of the disclosed methods to generate conformer ensembles for peptide-Major Histocompatibility Complex (pMHC) complexes that may be used in downstream predictions of peptide binding and/or other features of peptide-MHC interactions.
  • pMHC peptide-Major Histocompatibility Complex
  • Binding interactions that drive the formation of peptide-protein complexes underlie a variety of biochemical processes of critical importance to living organisms.
  • HLAs Human leukocyte antigens
  • Chr 6 chromosome 6
  • the human major histocompatibility complex is a linked set of genetic loci comprising the HLA sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 2 of 108 genes that play a fundamental role in, e.g., the cell-mediated immune response to infectious disease, the acceptance of transplanted tissues, etc.
  • the MHC complex encodes the ⁇ -chains of the MHC class I molecules HLA-A, HLA-B, and HLA-C (alleles) and the ⁇ - and ⁇ -chains of the MHC class II molecules HLA-DR, HLA-DP, and HLA-DQ (allotypes), all of which are expressed in a co-dominant fashion.
  • Cytotoxic T lymphocytes are activated upon recognition of peptides bound to the MHC class I molecules (see, e.g., Ito et al. (2010), “Regulation of the Induction and Function of Cytotoxic T Lymphocytes by Natural Killer T Cell”, Journal of Biomedicine and Biotechnology 2010:641757).
  • Activated cytotoxic T cells secrete essential cytolytic mediators (e.g., perforin, granzyme, etc.) and induce apoptosis in target cells (tumor cells, viral infected cells, etc.).
  • cytotoxic T cells secrete cytokines such as interferon gamma (IFN- ⁇ ) and tumor necrosis factor alpha (TNF- ⁇ ), which enhance antigen presentation and mediate antipathogenic effects.
  • IFN- ⁇ interferon gamma
  • TNF- ⁇ tumor necrosis factor alpha
  • IL2 interleukin-2
  • IFN- ⁇ tumor necrosis factor alpha
  • Peptide-MHC binding affinity is primarily determined by the amino acid sequence of the peptide binding core (also referred to as “minimal epitope”, referring to the part of a peptide extending between the “anchor residues” of the peptide that fit into pockets of the binding groove of the MHC sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 3 of 108 molecule in the peptide-MHC complex, typically having a median length of about nine amino acid residues); however, the amino-terminal flanking (N-flank) sequence and/or the carboxy- terminal flanking (C-flank) sequence (i.e.
  • ML machine learning
  • the disclosed methods can be used to generate an ensemble (i.e., a group) of allowed peptide-protein conformers.
  • the allowed peptide-protein conformers are allowed three-dimensional structures for a peptide- protein complex that can exist at any given point in time, and that impact the function of the peptide-protein complex at any given point in time.
  • the ensemble of allowed peptide-protein conformers represent an allowed distribution of structural parameters for the peptide-protein complex and can be used to more accurately predict features and functional properties of the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 4 of 108 peptide-protein complex (e.g., complementary interface features, complementary surface fingerprint features, and/or binding of the peptide-protein complex to other proteins or protein complexes).
  • the improvements in prediction accuracy enabled by the disclosed methods and systems are based on the use of embedded representations of peptide and protein sequences as input, machine-learning based prediction of structural constraints for the peptide-protein complex (e.g., predicted posterior distributions for pairwise distance data for amino acid residues in the peptide and the protein, and predicted posterior distributions for dihedral angle data for the peptide in the peptide-protein complex), and the use of the predicted structural constraints to select an initial template structure for the peptide-protein complex and generate a plurality of compatible structures that captures the dynamic behavior of the peptide-protein complex in solution.
  • prior methods have failed to account for the dynamic properties and range of allowed conformational states of peptide-protein complexes in solution.
  • the present inventors have shown through extensive benchmarking against ensembles from molecular dynamics (MD) simulations that the approach is able to recover a majority of the backbones sampled by the molecular dynamics simulations for peptide-HLA complexes involving peptide lengths up to 13 amino acids.
  • the inventors further demonstrated that the approach was able to capture the dynamic nature of the wild-type and mutating peptide-HLA complexes recognized by TCRs.
  • the methods described herein are the first methods able to capture the flexibility of peptides (even those up to 13 amino acids in lengths) in a matter of minutes without relying on time-intensive MD simulations.
  • the disclosed machine learning-based methods can be used to generate conformer ensembles for pMHC complexes that represent an allowed distribution of structural parameters for a given pMHC complex, and that can be used in downstream predictions of the features and functional properties of the pMHC complex (e.g., predictions of complementary interface features, complementary surface fingerprint features, and/or binding of the pMHC complex to other proteins or protein complexes. More accurate predictions of pMHC features and functional properties can, for example, facilitate the identification of epitopes (e.g., peptide fragments derived from neoantigens associated with a tumor) that, when presented on the surface of antigen-presenting cells, provoke a robust immune response.
  • epitopes e.g., peptide fragments derived from neoantigens associated with a tumor
  • Improved representation of the dynamic nature of peptide-HLA complexes also improves the ability to capture conformations that can be present when the peptide-MHC complex is bound by a TCR.
  • Dynamic allostery is believed to play an important part in TCR sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 5 of 108 recognition of peptide-HLA complexes and correctly capturing this is important to be able to accurately predict TCR cross-reactivity and specificity.
  • the disclosed methods can thus aid in selecting peptide fragments of neoantigens for inclusion in or development of therapeutic cancer vaccines.
  • neoantigens associated with a tumor can be patient- specific, and the disclosed methods can aid in selecting peptide fragments of the patient- specific neoantigens for inclusion in or development of personalized cancer vaccines.
  • the disclosed methods can also aid in designing and/or characterizing any TCR-based therapeutic or other tumor antigen recognizing therapy.
  • the disclosed methods can aid in design and/or characterizing of Bi-specific T-cell engagers (BiTE), T cell receptor-engineered T cell (TCR-T) therapy, chimeric antigen receptor (CAR) T cell therapy (CAR-T), etc.
  • systems configured to perform any of the methods disclosed herein, and computer-readable storage media comprising one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform any of the methods disclosed herein.
  • the systems may comprise one or more processors and one or more computer-readable storage media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods disclosed herein.
  • Disclosed herein are computer-implemented methods for generating a plurality of compatible structures for a peptide-protein complex comprising: inputting peptide sequence data for at least one peptide and protein sequence data for at least one protein into a trained machine learning model to determine: predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance data and the predicted dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data.
  • a plurality of compatible structures may also be referred to as “a conformer ensemble”.
  • An initial structure may also be referred to as “template structure”.
  • the trained machine learning model may be a machine learning model that has been trained to take as input peptide sequence data for at least one peptide and protein sequence data for at least one protein, and produce as output predicted sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 6 of 108 pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex.
  • the machine learning model may have been trained using structural data for a plurality of training peptide-protein complexes.
  • the structural data may comprise, for each of the plurality of training peptide-protein complexes, values of: pairwise distances between at least one peptide amino acid residue of the peptide and at least one protein amino acid residue of protein in the training peptide-protein complex, and one or more dihedral angles for at least one peptide amino acid residue in the peptide in the training peptide-protein complex.
  • the values may be experimentally determined and/or predicted using a structural prediction algorithm.
  • the computer-implemented method further comprises predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex.
  • each compatible structure for the peptide-protein complex conforms with one or more posterior distributions of the predicted pairwise distance data and with one or more posterior distributions of the predicted dihedral angle data output by the trained machine learning model.
  • the predicted pairwise distance data may comprise, for each of one or more pairs of at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein, a predicted posterior distribution of pairwise distances between said peptide amino acid residue and said protein amino acid residue in the peptide-protein complex.
  • the predicted dihedral angle data may comprise, for each of one or more peptide amino acid residues in the at least one peptide, a predicted posterior distribution of one or more dihedral angles of the peptide amino acid residue in the peptide-protein complex.
  • Predicted data comprising a predicted posterior distribution may refer to the predicted data comprising predicted values of parameters that characterize said distribution.
  • the posterior distributions may be used as constraints in a peptide docking protocol, i.e. in a physics engine for modeling peptide-protein complexes.
  • the trained machine learning model has been trained at least in part to predict at least one auxiliary features of input protein-peptide complexes.
  • each compatible structure for the peptide-protein complex conforms with the sequence data for the at least one peptide and the at least one protein.
  • sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 7 of 108 generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises generating a template structure that conforms with the sequence data for the at least one peptide and the at least one protein using the initial structure and a homology modeling method.
  • the at least one auxiliary feature comprises a prediction of binding of the at least one peptide to the at least one protein.
  • the at least one auxiliary feature comprises a prediction of at least one interface feature for the peptide- protein complex.
  • the peptide sequence data comprises a peptide embedded representation for the at least one peptide generated using a trained peptide embedding machine learning model.
  • the protein sequence data comprises a protein embedded representation for the at least one protein generated using a trained protein embedding machine learning model.
  • the trained peptide embedding machine learning model may take as input a peptide sequence and produce as output a peptide embedded representation.
  • the trained protein embedding machine learning model may take as input a protein sequence and produce as output a protein embedded representation.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model may be machine learning models that have been trained in a self-supervised manner to learn a peptide/protein embedding from an input sequence.
  • the peptide embedding machine learning model and/or the trained protein embedding machine learning model may have been trained using unlabeled data (e.g. protein and/or peptide sequence data) and tasks such as masked language modelling, causal language modelling, etc.
  • the trained peptide embedding machine learning model and the trained protein embedding machine learning model are the same model.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise foundational models.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise protein language models. In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise large language models (LLM). In some embodiments, the trained peptide embedding machine learning model and/or the trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 8 of 108 protein embedding machine learning model comprise an Evolutionary Scale Model (ESM). In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are integrated with, and re-trained simultaneously with, the trained machine learning model.
  • EMM Evolutionary Scale Model
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are fine- tuned (e.g. after self-supervised training) for a peptide-protein complex related prediction task. Such training may be supervised training. In embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are fine- tuned for peptide-protein binding prediction. Fine tuning of a model (e.g. a foundation model) for a particular task may include incorporating the model as part of a machine learning model trained to perform the particular task, and at least partially retraining the parameters of the model for this task.
  • a model e.g. a foundation model
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model may be deep learning models, e.g. deep neural networks. Any deep learning architecture known in the art suitable for use in language modeling may be used.
  • the trained machine learning model comprises a trained neural network.
  • the trained machine learning model comprises a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained state space model.
  • MLP multilayer perceptron
  • the trained machine learning model is trained using a training data set that includes incomplete structural data for one or more peptide-protein complexes.
  • the trained machine learning model is trained using a training data set that includes, for each of a plurality of training peptide-protein complexes, values of pairwise distances between peptide amino acid residues and protein amino acid residues in the training peptide-protein complex, and values of dihedral angles for peptide amino acid residues in the training peptide-protein complex, wherein the training data does not comprise all pairwise distances for one or more training peptide-protein complexes and/or does not comprise both dihedral angles ⁇ and ⁇ for one or more peptide amino acid residues in one or more training peptide-protein complexes.
  • such incomplete training peptide- protein complexes structural data may still have been used for training the machine learning model.
  • the trained machine learning model has been trained using training sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 9 of 108 data including values of dihedral angles for peptide amino acid residues in training peptide- protein complexes for only one of the dihedral angles ⁇ and ⁇ .
  • the trained machine learning model is further configured to determine an uncertainty in the predicted pairwise distance data and the predicted dihedral angle data.
  • the trained machine learning model is configured to predict parameters characterizing: a distribution of pairwise distance for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and a joint distribution of dihedral angles ⁇ and ⁇ for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex.
  • the trained machine learning model is configured to predict parameters characterizing: a distribution of pairwise distance for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; a distribution of dihedral angle ⁇ for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; and a distribution of dihedral angle ⁇ for the at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex.
  • the parameters characterizing a distribution may comprise a parameter that characterizes the maximum likelihood value of the distribution (e.g.
  • the machine learning model is trained or has been trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide- protein combinations; training a classifier to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and appending a regressor head to at least a portion of the trained classifier and training the regressor head to predict pairwise distance data and dihedral angle data (e.g.
  • pairwise distance data and dihedral angle data probabilistic distributions of pairwise distance data and dihedral angle data) for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from experimental (e.g. crystal) structure data for a set of example peptide- protein complexes.
  • experimental (e.g. crystal) structure data for a set of example peptide- protein complexes.
  • the pairwise distance data and dihedral angle data derived from the experimental (e.g. crystal) structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide- sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 10 of 108 protein complex is presented as a separate training instance during training of the regressor head.
  • the machine learning model has been trained in at least two steps.
  • the machine learning model may have been trained to predict peptide-protein binding for an input peptide sequence data and an input protein sequence data, using training data comprising peptide sequence data for a plurality of peptides; protein sequence data for a plurality of proteins; and peptide-protein binding data for a plurality of peptide-protein combinations.
  • Predicting peptide-protein binding may comprise classifying peptide-protein combinations as binding / not binding to form a peptide-protein complex.
  • Such training data may comprise peptide and protein sequence data for a plurality of peptides and proteins that are known to bind to each other to form a peptide-protein complex, and peptide and protein sequence data for a plurality of peptides and proteins that are not known or expected to bind to each other to form a peptide-protein complex.
  • the machine learning model may comprise a peptide sequence embedding model, a protein sequence embedding model and a classification head. Such a model may be referred to as “classifier”.
  • a second step at least a part of the machine learning model (e.g.
  • the peptide sequence embedding model and the protein sequence embedding model may have been trained to predict the predicted pairwise distance data and the predicted dihedral angle data (e.g. probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes).
  • the training in the second step may be based on pairwise distance data and dihedral angle data derived from experimental (e.g. crystal) structure data for a set of example (i.e. training) peptide-protein complexes.
  • the machine learning model may comprise the peptide sequence embedding model, the protein sequence embedding model and a regression head. Such a model may be referred to as “regressor”.
  • the machine learning model (e.g. regressor) is further trained or has been further trained on augmented pairwise distance data and dihedral angle data.
  • Augmented pairwise distance data and dihedral angle data may refer to data that has not been experimentally determined.
  • Such data may have been obtained for example using a structural modeling method, such as e.g., using a trained structure prediction machine learning model, e.g. a model of the AlphaFold family.
  • Such data may have been obtained by obtaining a predicted crystal structure for a peptide and protein sequence known to bind to each other to form a peptide-protein complex using a trained structure prediction machine learning model.
  • the augmented pairwise distance data and dihedral angle data is derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 11 of 108 least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide- protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide- protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regress
  • identifying the initial structure for the peptide-protein complex based on the predicted pairwise distance data and predicted dihedral angle data comprises: obtaining a plurality of experimental (e.g., crystal) structures for peptide-protein complexes from a protein structure database; and selecting an experimental (e.g. crystal) structure from the plurality of experimental (e.g. crystal) structures based on a sequence and/or structural similarity between the experimental structure and the peptide-protein complex.
  • a structural similarity may be a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model.
  • the experimental structure is selected based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and a maximum likelihood estimate for the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model.
  • selecting an experimental structure based on a structural similarity comprises: (i) for each example peptide-protein experimental structure in the plurality of experimental structures, determining a likelihood (e.g.
  • the protein structure database comprises the Protein Data Bank (PDB).
  • the initial structure comprises a plurality of initial structures.
  • generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises: inputting the initial structure for the peptide- protein complex into a structural modeling software package; inputting the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, wherein compatible structures for the peptide-protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function.
  • the computer-implemented method further comprises iterating over a plurality of candidate peptides to identify at least one peptide that has a maximum likelihood for formation of the peptide-protein complex.
  • the at least one peptide comprises at least a portion of a peptide that is expressed in normal cells in a subject.
  • the at least one peptide comprises at least a portion of a neoantigen expressed in tumor cells from a cancer patient.
  • the at least one peptide comprises at least a portion of a peptide that is expressed in a virus or bacterium.
  • the at least one protein comprises a Major Histocompatibility Complex (MHC) class I protein.
  • the at least one protein comprises a Major Histocompatibility Complex (MHC) class II protein.
  • the computer-implemented method further comprises iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in normal cells from the subject for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex.
  • pMHC-I peptide-MHC class I
  • pMHC-II peptide-MHC class II
  • the computer-implemented method further comprises iterating over a plurality of candidate neoantigen sequences, or portions thereof, expressed in tumor cells from the cancer patient for at least one MHC class I or class II protein to rank order the neoantigen sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC- II) complex.
  • pMHC-I peptide-MHC class I
  • pMHC- II peptide-MHC class II
  • the computer-implemented method further comprises iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in a virus or bacterium for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex.
  • the computer-implemented method (or a method associated with the computer-implemented method) further comprises developing a personalized anti- cancer vaccine based on the rank ordering of the neoantigen sequences, or portions thereof.
  • the computer-implemented method further comprises developing an anti-viral vaccine, an anti- bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof.
  • the computer-implemented method further comprises predicting one or more characteristics of a complex comprising the peptide-protein complex and a further protein.
  • the further protein may be e.g. a T cell receptor, antigen recognition molecule, T cell sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 14 of 108 engaging molecule, etc.
  • the predicted characteristics may include one or more of: the likelihood of formation of the complex, the stability of the complex, the binding energy of the complex, the total energy of the complex.
  • systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the computer-implemented methods described herein.
  • non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform any of the computer-implemented methods described herein.
  • FIG. 1 provides a block diagram of an example prediction system, in accordance with some implementations of disclosed methods and systems.
  • FIG. 2 provides a non-limiting example of a block diagram for a system for generating conformer ensembles for a peptide-protein complex, in accordance with some implementations of the disclosed methods and systems.
  • FIG. 3A provides a schematic illustration of using a protein language model to convert a peptide sequence into a peptide embedded representation, in accordance with some implementations of the disclosed methods and systems.
  • FIG. 3B provides a schematic illustration of using a protein language model to convert a protein sequence (e.g., an MHC sequence) into a protein embedded representation, in accordance with some implementations of the disclosed methods and systems.
  • FIG.4 provides a schematic illustration of using a trained machine learning model (e.g., a Structural Constraint Prediction Model) to output pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a pMHC complex) based on input embedded representations of peptide and protein (e.g., MHC) sequences, in accordance with some implementations of the disclosed methods and systems.
  • a trained machine learning model e.g., a Structural Constraint Prediction Model
  • FIG.5 provides a schematic illustration of using pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a peptide – MHC complex) output by a trained machine learning model (e.g., a Structural Constraint Prediction Model) to generate conformer ensembles for the peptide-protein complex (e.g., a pMHC complex), in accordance with some implementations of the disclosed methods and systems.
  • FIG. 6 provides a non-limiting schematic illustration of a process for training a machine learning model to predict pairwise distance data and dihedral angles for a peptide- protein complex, in accordance with one implementation of the disclosed methods.
  • FIG.7 provides a non-limiting schematic illustration of a process of using a trained machine learning model to predict pairwise distance data and dihedral angles for a peptide- protein complex which can then be used to: (i) select a template structure from a protein structure database, and (ii) provide constraints on allowable conformations when generating an ensemble of conformers for the peptide-protein complex, in accordance with one implementation of disclosed methods.
  • FIG.8 provides a non-limiting example that illustrates how predicted distributions of pairwise distance data and dihedral angle data for a pMHC complex can be converted to structural constraints, and how pMHC conformational ensembles can then be generated based on such structural constraints.
  • FIG. 9 provides a non-limiting example of a process flowchart for generating a plurality of compatible structures for a peptide-protein complex, in accordance with one implementation of the disclosed methods. [0048] FIG.
  • FIG. 10 provides a non-limiting example of a process flowchart for training a machine learning model to predict distributions of pairwise distance data and dihedral angle data for a peptide-protein complex, in accordance with one implementation of the disclosed methods.
  • FIG. 11A provides a non-limiting example of data for the root mean square deviation (RMSD) of predicted ⁇ and ⁇ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where the ⁇ and ⁇ dihedral angles were treated independently during training and inference.
  • RMSD root mean square deviation
  • FIG. 11B provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of predicted ⁇ and ⁇ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where the ⁇ and ⁇ dihedral angles were treated either independently or jointly during training and inference.
  • RMSD root mean square deviation
  • FIG. 11C provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of predicted ⁇ and ⁇ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where one data set (red symbols) was generated using a prior protein structure prediction model (AlphaFold – FineTune), and the other data set (green symbols) was generated using one implementation of the methods disclosed herein.
  • FIG. 11D provides a non-limiting example of preliminary data for predicted average dihedral distance in a set of benchmark pMHC complexes comprising peptides of different length.
  • FIG.11E provides a non-limiting example of preliminary data for the recovery of crystal structure contacts for a set of benchmark pMHC complexes using the disclosed structural prediction methods.
  • sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 17 of 108
  • Fig.12A provides a non-limiting example of data for the recovery of distances for test sets of pMHC complexes using the disclosed structural prediction methods.
  • Fig.12B provides a non-limiting example of data for the recovery of dihedral angles for test sets of pMHC complexes using the disclosed structural prediction methods.
  • Fig.12A provides a non-limiting example of data for the recovery of distances for test sets of pMHC complexes using the disclosed structural prediction methods.
  • Fig.12B provides a non-limiting example of data for the recovery of dihedral angles for test sets of pMHC complexes using the disclosed structural prediction methods.
  • Fig.12A provides
  • Fig. 12C provides a non-limiting example of data for the recovery of dihedral angles using the disclosed structural prediction methods, for a representative pMHC 9-mer example (1I1Y).
  • Fig. 13 provides a non-limiting example of data for the number of overlapped models in conformer ensembles predicted using methods of the present disclosure or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations.
  • Fig. 14A provides a non-limiting example of evaluation metrics comparing conformer ensembles obtained using the disclosed structural prediction methods or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations in a first test set.
  • Fig. 14A provides a non-limiting example of evaluation metrics comparing conformer ensembles obtained using the disclosed structural prediction methods or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations in a first test set.
  • FIG. 14B provides a non-limiting example of evaluation metrics comparing conformer ensembles obtained using the disclosed structural prediction methods or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations in a second test set.
  • Fig.15A provides a non-limiting example of the ability of the disclosed structural prediction methods to capture dynamic behaviors of peptides involved in TCR recognition.
  • Fig. 15B provides data for the same peptide as in Fig. 15B, illustrating that molecular dynamics (MD) simulations are unable to overcome the energy barriers associated with sidechain flipping in the TCR bund conformation of the peptide.
  • FIG. 16A provides a non-limiting example of control data for pMHC template structure selection.
  • FIG. 16B provides a non-limiting example of control data for pMHC template structure refinement.
  • FIG. 17A provides a non-limiting example of molecular dynamics data for a sampling of peptide conformations within a pMHC complex. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 18 of 108
  • FIG. 17B provides a non-limiting example of a graphical representation of the pMHC complex of FIG.17A.
  • FIG. 17C provides a non-limiting example of a plot of peptide backbone fluctuations as a function of residue number for the pMHC complex of FIG.17A.
  • FIG. 17A provides a non-limiting example of molecular dynamics data for a sampling of peptide conformations within a pMHC complex. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 18 of 108
  • FIG. 18A provides a non-limiting example of molecular dynamics data for a sampling of peptide conformations within a pMHC complex.
  • FIG. 18B provides a non-limiting example of a graphical representation of the pMHC complex of FIG.18A.
  • FIG. 18C provides a non-limiting example of a plot of peptide backbone fluctuations as a function of residue number for the pMHC complex of FIG.18A.
  • FIG.19 provides a non-limiting example of a block diagram of a computer system, in accordance with some implementations of the methods and systems disclosed herein. [0071] FIG.
  • FIG. 20 provides a non-limiting example of a block diagram of an artificial intelligence (AI) architecture included as part of the example computing system of FIG.4, in accordance with some implementations of the methods and systems disclosed herein.
  • AI artificial intelligence
  • similar components and/or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label.
  • Machine learning-based methods for modeling conformational states of peptide- protein complexes that account for the dynamic nature of three-dimensional peptide and protein structures are described.
  • the disclosed methods can be used to generate an ensemble of allowed peptide-protein conformers (e.g., an ensemble of allowed three-dimensional structures for a peptide-protein complex that can exist at any given point in time – also referred to herein as “compatible structures” for the peptide-protein complex, which can impact the function of the peptide-protein complex at any given point in time).
  • the ensemble of allowed peptide-protein conformers can be represented as an allowed distribution of structural parameters for the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 19 of 108 peptide-protein complex.
  • the ensemble of allowed peptide-protein conformers can be used to more accurately predict features and functional properties of the peptide-protein complex (e.g., complementary interface features, complementary surface fingerprint features, and/or binding of the peptide-protein complex to other proteins or protein complexes).
  • the disclosed methods for generating a plurality of compatible structures for a peptide-protein complex can comprise a first step of inputting peptide sequence data (e.g.
  • protein sequence data e.g. an amino acid sequence or pseudo-sequence, or an embedding thereof
  • protein which can be a full length protein or a truncated version thereof including at least the portion of the protein interacting with the peptide
  • a trained machine learning model e.g., a structural constraint prediction model or a machine learning model comprising a structural constraint prediction model and one or more embedding models configured to obtain a protein embedded representation and a peptide embedded representation, also referred to herein as “protein embedding” and “peptide embedding”.
  • the structural constraint prediction model can be configured (e.g.
  • the predicted structural constraint data can then be used to identify an initial structure for the peptide-protein complex and generate the plurality of compatible structures for the peptide-protein complex based on the initial structure, and the predicted values or distributions of pairwise distances and dihedral angles.
  • the method can further comprise predicting binding of the peptide-protein complex to another protein or protein complex based on the plurality of sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 20 of 108 compatible structures for the peptide-protein complex.
  • the improvements in processing performance and prediction accuracy enabled by the disclosed methods and systems are based on the use of embedded representations of peptide and protein sequences as input, machine-learning based prediction of structural constraints for the peptide-protein complex (e.g., predicted pairwise distance data or posterior distributions for pairwise distance data for amino acid residues in the peptide and the protein, and predicted dihedral angle data or posterior distributions for dihedral angle data for the peptide in the peptide-protein complex), and the use of the predicted structural constraints to select an initial template structure for the peptide-protein complex and generate a plurality of compatible structures that captures the dynamic behavior of the peptide-protein complex in solution.
  • structural constraints for the peptide-protein complex e.g., predicted pairwise distance data or posterior distributions for pairwise distance data for amino acid residues in the peptide and the protein, and predicted dihedral angle data or posterior distributions for dihedral angle data for the peptide in the peptide-protein complex
  • FIG. 1 provides a block diagram of an example prediction system, in accordance with some embodiments.
  • Prediction system 100 can be used, for example, to predict structural properties of peptide-protein complexes (e.g., pMHC complexes), to determine a plurality of allowable conformations (e.g., a conformer ensemble) for a peptide-protein complex (e.g., a pMHC complex), and/or to determine a predicted pMHC interaction with one or more other proteins in an immunoprotein complex (IPC) related to the immunological activity of peptides (such as e.g. interaction between a pMHC and a T cell receptor molecule (TCR) or a part thereof) and, in particular, mutated peptides (e.g., neoantigen peptides).
  • IPC immunoprotein complex
  • the prediction system sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 21 of 108 100 includes computing platform 102, data store 104, and display system 106.
  • Computing platform 102 may take various forms.
  • the computing platform 102 includes a single computer (or computer system) or multiple computers in communication with each other.
  • the computing platform 102 can be a cloud computing platform.
  • Data store 104 and display system 106 are each in communication with computing platform 102.
  • one or more of: data store 104 or display system 106 can be considered part of, or otherwise integrated with, computing platform 102.
  • computing platform 102, data store 104, and display system 106 can be separate components in communication with each other, but in other examples, some combination of these components can be integrated together. Communication between the different components can be implemented using any number of wired communications links, wireless communications links, optical communications links, or a combination thereof.
  • the prediction system 100 includes a sequence analyzer 108, which can be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 can be implemented in the computing platform 102. The sequence analyzer 108 receives sequence data 110 for processing.
  • sequence data 110 can be sent as input into the sequence analyzer 108, retrieved from the data store 104 or some other type of storage (e.g., cloud storage), accessed from cloud storage, or obtained in some other manner.
  • the sequence data 110 can be retrieved from the data store 104 in response to receiving user input entered by a user via an input device (not shown in FIG.1).
  • the sequence data 110 can be generated from processing of a set of samples 112.
  • the set of samples 112 may take the form of one or more biological samples from one or more subjects (e.g., a diseased sample, a healthy sample, or a combination thereof).
  • the set of samples 112 may include a sample obtained from a tumor of a subject (e.g.
  • the tumor can be a manifestation of, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myelogenous leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, T cell sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 22 of 108 lymphocytic leukemia, non-small cell lung cancer, small-cell lung cancer, or a combination thereof.
  • a sample in the set of samples 112 may include, for example, various IPC molecules, various peptides, nucleic acids coding for said IPC molecules and/or peptides, or a combination thereof.
  • the peptides may include one or more mutated peptides (e.g., neoantigens).
  • the IPC molecules may include, for example, various MHC molecules, various TCR molecules, or a combination thereof.
  • the set of samples 112 includes IPC 114 (e.g., MHC Class I molecule, MHC Class II molecule, various TCR molecules, etc.) and/or nucleic acid sequences coding for IPC 114, from which the sequence of IPC 114 can be inferred.
  • the set of samples can include at least one protein 123 (i.e., the source protein) or a nucleic acid sequence coding for protein 123 or a part thereof, from which the sequence of the protein 123 (referred to herein as protein sequence 160) or part thereof (such as amino acid chain 116 or peptide 118) can be inferred.
  • An amino acid chain 116 can be identified from the at least one protein 123 and can be a chain of amino acids that includes a peptide 118 (minimal epitope).
  • the amino acid chain 116 can optionally additionally include an N-flank 120, and/or a C-flank 122, referring respectively to a sequence located at the N-terminus side of the peptide 118 and the C-terminus side of the peptide 118.
  • the peptide 118 can be considered a mutated peptide when it includes one or more variants (e.g., one or more sequence variations) when compared to a corresponding reference sequence.
  • the protein 123 can be a source protein for the amino acid chain 116, which can be generated through proteolysis, which is the process by which proteins (e.g., the protein 123) are broken down into smaller polypeptides or amino acids.
  • the protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, where specific enzymes called proteases cut the peptide bonds between amino acids in the protein 123.
  • the set of samples 112 can be processed to generate the sequence data 110. In some embodiments, multiple samples in the set of samples 112 can be processed at different times.
  • the prediction system 100 includes a sample analyzer that can be used in processing the set of samples 112 to generate the sequence data 110.
  • the sequence data 110 includes, for example, at least one amino acid sequence 129 and at least one IPC sequence 124 (e.g., one IPC sequence 124 corresponding to IPC 114).
  • the amino acid sequence 129 may comprise one or more of: a peptide sequence 126 (e.g., one peptide sequence 126 corresponding sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 23 of 108 to peptide 118), an amino-terminal flanking (N-flank) sequence 128 (e.g., one N-flank sequence 128 corresponding to N-flank 120), or a carboxy-terminal flanking (C-flank) sequence 130 (e.g., one C-flank sequence 130 corresponding to C-flank 122).
  • N-flank sequence 128 e.g., one N-flank sequence 128 corresponding to N-flank 120
  • IPC sequence 124 can be, for example, an MHC sequence 135 that characterizes at least a portion of the MHC.
  • IPC sequence 124 can be, for example, a TCR sequence 131 that characterizes at least a portion of the TCR.
  • IPC sequence 124 may include both an MHC sequence 135 characterizing at least a portion of an MHC molecule and a TCR sequence 131 characterizing at least a portion of a TCR molecule.
  • the sequence data 110 may include IPC sequence 124 in the form of an MHC sequence 135 characterizing at least a portion of an MHC molecule, as well as a separate TCR sequence 131 characterizing at least a portion of a TCR.
  • Protein sequence 160 characterizes at least a portion of the protein 123.
  • the protein sequence 160 can be identified by performing a reverse lookup in a database (e.g., the UniProt database) based on the mutated peptide data (e.g., IPC sequence 124) obtained from the sample.
  • Protein sequence 160 may be used to provide context and/or orthogonal information when characterizing a peptide 118, for example to compare a peptide 118 to a wild-type equivalent, to determine expression of the protein 123 in one or more samples or tissues, etc.
  • Availability of the protein sequence 160 is completely optional, and methods of the disclosure can make use of e.g. only peptide sequence 118, peptide sequence 116 or parts thereof.
  • Peptide sequence 126 characterizes at least a portion of the peptide 118.
  • N-flank sequence 128 characterizes at least a portion of the N-flank 120.
  • the corresponding sequence for N-flank 120 can be trimmed to generate the N-flank sequence 128.
  • N-flank sequence 128 can comprise or consist of the sequence of a predetermined number of residues upstream (i.e. N-terminal) of peptide 118.
  • C- flank sequence 130 characterizes at least a portion of the C-flank 122.
  • Sequence analyzer 108 receives the sequence data 110 as input for processing.
  • the sequence analyzer 108 includes one or more machine-learning models 132 that process the sequence data 110.
  • the sequence analyzer 108 can process the sequence data 110 (e.g., using embedding engine(s) 140) prior to sending the sequence data 110 (or an embedded form thereof) into one or more of the machine learning models 132 for further processing.
  • machine learning models 132 may comprise, for example, one or more of: one or more embedding engines 140, a structural constraint prediction model 141, and a conformer ensemble generation engine 142.
  • Embedding generation engine(s) 140 can be configured to process peptide and/or protein sequence (e.g., MHC sequence) data (e.g. IPC sequence 124 and/or amino acid sequence 129) and output embeddings for the respective sequences.
  • an embedding of a peptide or protein sequence can refer to a vector representation of a peptide or protein sequence that encodes important features of the peptide or protein in a way that facilitates processing by downstream computational models.
  • An embedding of a peptide or protein sequence can be a latent representation of a machine learning model trained to learn a latent space in which input sequences can be represented and from which the sequences can be reconstructed and/or one or more features of the input sequences can be predicted.
  • Structural constraint prediction model 141 can be configured to process peptide and/or protein embeddings for a peptide-protein complex, which are outputted by the embedding generation engine(s) 140, and output a predicted set of structural constraints associated with the peptide- protein complex.
  • the predicted set of structural constraints can include e.g., predicted pairwise distance data (or a distribution thereof) for at least one peptide amino acid residue and at least one protein amino acid residue in the peptide-protein complex and/or predicted dihedral angle data (or a distribution thereof) for at least one peptide amino acid residue in the peptide in the peptide-protein complex.
  • Conformer ensemble generation engine 142 can be configured to generate a plurality of compatible structures for the peptide-protein complex (i.e., a conformer ensemble 148) based on an initial template structure and the set of structural constraints output by the structural constraint prediction model 141.
  • the conformer ensemble generation engine 142 can be configured to take as input an initial template structure and a set of structural constraints output by the structural constraint prediction model 141 for a peptide-protein sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 25 of 108 complex, and produce as output a plurality of compatible structures for the peptide-protein complex.
  • the plurality of compatible structures can form a conformer ensemble, also referred to as conformational ensemble, structural ensemble or simply ensemble.
  • a conformer ensemble for a peptide-protein complex can be a set of conformations that together describe the structure of the peptide-protein complex.
  • the conformer ensemble may be expected to capture a range of conformations that the peptide-protein complex may adopt under normal (e.g. physiological) conditions.
  • the one or more machine-learning models 132 can be used in either a training mode or a prediction mode. In the training mode, one or more of the machine-learning models 132 can be trained using training data 133. Examples of the training data 133 are described in more detail below.
  • the one or more machine-learning models 132 are trained such that they can be used in the prediction mode.
  • One or more of the machine-learning models 132 are used to process the IPC sequence 124 and the amino acid sequence 129.
  • separate processing engines such as e.g. separate embedding engines 140
  • IPC sequences and amino acid sequences can be used for processing IPC sequences and amino acid sequences. This can enable improved predictive performance of the one or more machine-learning models 132.
  • one or more of the machine- learning models 132 process one or more of: a MHC sequence 135, a TCR sequence 131, a protein sequence 160, a peptide sequence 126, an N-flank sequence 128, or a C-flank sequence 130. Examples of implementations for these different processing engines are described in greater detail below.
  • processing engine identify at least one software component and/or a combination of at least one software component and at least one hardware component which are designed/programmed/configured to interact and/or communicate data to other software and/or hardware components including but not limited to other processing engines.
  • the one or more machine-learning models 132 process the sequence data 110 to generate an output that can be used to generate a report 144.
  • the report 144 may include the exact output of any one or more of the machine-learning models 132, a transformed (e.g. further processed) or filtered version of the output of any one or more of the machine learning models 132, or both.
  • the report 144 may include notifications, recommendations, alerts, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 26 of 108 or other information generated by the sequence analyzer 108 based on the output of one or more of the machine-learning models 132.
  • the report 144 can be an output that includes, for example, information about structural constraints for a peptide-protein complex (e.g., a pMHC complex) such as e.g., information output by structural constraint prediction model 141 or information derived therefrom, a conformer ensemble 148 for a peptide-protein complex (e.g., a pMHC complex) such as e.g.
  • a conformer ensemble 148 output by ensemble generation model 142 predicted features of a peptide-protein complex (e.g., a pMHC complex) such as e.g. predicted features derived at least in part from a conformer ensemble 148 output by ensemble generation model 142, which can include e.g., immunological activity of interest with respect to one or more peptides (e.g., one or more mutated peptides).
  • a peptide-protein complex e.g., a pMHC complex
  • predicted features derived at least in part from a conformer ensemble 148 output by ensemble generation model 142 which can include e.g., immunological activity of interest with respect to one or more peptides (e.g., one or more mutated peptides).
  • the report 144 may include information about structural predictions 146 (e.g., predicted distributions of pairwise distances for amino acid residues and/or dihedral angles) for a peptide protein complex (e.g.., a pMHC complex), conformer ensembles 148 (e.g., a set of allowed conformations that are consistent with the structural predictions) for a peptide protein complex (e.g.., a pMHC complex), peptide-protein complex feature predictions 162, an immunological activity prediction relating to the amino acid 116 (e.g., peptide 118, N-flank 120, C-flank 122, etc.) and IPC 114 (e.g., MHC-I, MHC-II, TCR, etc.), or any combination thereof.
  • structural predictions 146 e.g., predicted distributions of pairwise distances for amino acid residues and/or dihedral angles
  • conformer ensembles 148 e.g., a set of allowed conformations that
  • the report 144 may include, for example, interaction information (e.g., an interaction affinity prediction that predicts a binding affinity between a peptide and an MHC, or an interaction prediction that predicts whether an MHC allele or allotype will present a peptide at a cell surface), immunogenicity information (e.g., an immunogenicity prediction that predicts the ability of a peptide to provoke an immune response in the context of an MHC molecule), or both.
  • the interaction information may provide predictions about a selected set of interactions between the amino acid sequence 129 and the IPC sequence 124.
  • the immunogenicity information may provide predictions about the immunogenicity of amino acid sequence 129 (e.g., including the immunogenicity of the peptide 118).
  • the immunogenicity information may provide predictions about the immunogenicity of peptide 118 in the context of the IPC 114, where the IPC is an MHC molecule.
  • the immunogenicity information may provide predictions about the immunogenicity of peptide 118 (or a peptide-protein complex comprising the peptide 118 and an IPC114 where the IPC is an MHC molecule) in relation to an IPC 114, where the IPC in a TCR molecule.
  • sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 27 of 108 [0094]
  • a report 144 can be displayed on a graphical user interface (GUI) 150 on the display system 106.
  • GUI graphical user interface
  • a user may view and/or interact with the report 144 via the graphical user interface 150.
  • the user may use the report 144 to make decisions about the treatment of a subject from which at least one of the set of samples 112 was obtained (or collected).
  • the user may use the report 144 to make decisions about the manufacture of a treatment for a subject from which at least one of the set of samples 112 was obtained (or collected).
  • the user may use the report 144 to select a peptide sequence for use in manufacturing a vaccine (such as e.g. a personalized cancer vaccine).
  • the prediction system 100 sends the report 144 to the remote system 152 (e.g., wirelessly).
  • the remote system 152 can be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, a tablet, a laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 can be a treatment manufacturing system (or machine) or a portion thereof.
  • FIG.2 provides a non-limiting example of a block diagram for a system 200 (e.g., a computer-implemented system) for generating conformer ensembles for a peptide-protein complex.
  • System 200 can be implemented using the Prediction System 100 described in FIG. 1.
  • system 200 may comprise the Sequence Analyzer 108 and one or more of the Machine Learning Models 132 in FIG.1.
  • the input to the system 200 includes protein sequence data (e.g., MHC Sequence 135) for at least one protein and peptide sequence data (e.g., Peptide Sequence 126) for at least one peptide.
  • the input can comprise peptide sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more than 50 peptides, or portions thereof.
  • peptide sequence data can comprise sequence data for peptides that are at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues in length.
  • the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof.
  • the peptides can comprise mutated peptides (e.g., neoantigens expressed in tumor cells), or portions thereof.
  • the normal peptide and/or mutated peptides may have been isolated from a subject (e.g., a patient) or otherwise identified from a sample previously obtained from a subject (such as e.g.
  • the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium.
  • the peptides can comprise viral peptides, or portions, thereof expressed in a virus.
  • the input can comprise protein sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, or more than 1000 proteins, or portions thereof.
  • protein sequence data can comprise sequences or pseudosequences of at least 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or at least 200 amino acid residues in length.
  • the proteins can comprise IPC proteins.
  • the IPC proteins also referred to herein as IPC molecules
  • MHC Sequence 135) and Peptide Sequence 126 can be processed by Embedding Generation Engine(s) 140 to generate protein embedded representation (e.g. MHC embedded representation 206) and Peptide embedded representation 208.
  • Protein embedded representations such as MHC embedded representation 206 and Peptide embedded representation 208 can be represented as data structures in the form of vectors.
  • each of embedded representation can be represented as a vector of size (t, e) where each residue of a protein (e.g., IPC 114, and peptide 118 shown in FIG.1) is represented by a token, t, and the contextual information of each residue is encoded in (e) dimensions.
  • Embedding Generation Engine(s) 140 can add “special tokens” which can be used to denote the beginning or end of the protein sequence.
  • special tokens For an embedded representation instantiated as a vector of size (t, e), t can represent the total number of residue tokens and special tokens.
  • Protein embedded representations can encode important features of a protein or peptide in a way that can be processed by downstream computational models. [0101] A non-exhaustive list of features that can be implicitly encoded in a protein embedded representation (i.e.
  • Evolutionary Information Proteins with similar sequences or structures are often evolutionarily related. Some features can capture evolutionary relationships and similarities, which can be useful for tasks like protein classification or function prediction.
  • Interaction with Other Molecules Information about how proteins interact with other molecules like DNA, RNA, or small molecules (like drugs) can also be encoded. This includes binding sites and interaction domains.
  • Post-Translational Modifications Proteins often undergo modifications after translation (like phosphorylation or glycosylation) which affect their function. These modifications can sometimes be inferred from the sequence and context and thus can be encoded as features.
  • the Embedding Generation Engine(s) 140 includes a Protein Embedding Model 202 for receiving an MHC sequence 135 and generating an MHC embedded representation 206, and a Peptide Embedding Model 204 for receiving a peptide sequence 126 and for generating a peptide embedded representation 208.
  • the Peptide Embedding Model 204 and the Protein Embedding Model 202 are the same model; in some instances, the Peptide Embedding Model 204 and the Protein Embedding Model 202 are different models.
  • Example methods for generating peptide and/or protein sequence embedding are described in, e.g., Ofer et al. (2021), “The Language of Proteins: NLP, Machine Learning & Protein Sequences”, Computational and Structural Biotechnology Journal 19:1750-1758.
  • the Peptide Embedding Model 204 and/or the Protein Embedding Model 202 can comprise large language models (LLMs) or protein language sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 30 of 108 models (PLMs), e.g., deep learning algorithms configured to strings of characters in order to perform predictive or generative tasks such as recognizing, summarizing, translating, predicting, and generating content.
  • LLMs and PLMs (which are special cases of LLMs trained using protein sequence data) are typically trained at least partially in an unsupervised manner (i.e.
  • LLMs and PLMs are frequently (but not exclusively) transformer based deep neural networks.
  • Large language models are described in more detail in, for example, Li et al. (2021), “BioSeq-BLM: A Platform for Analyzing DNA, RNA and Protein Sequences Based on Biological Language Models”, Nucleic Acids Res.49, e129–e129, and Naveed et al. (2023), “A Comprehensive Overview of Large Language Models”, arxiv.org/pdf/2307.06435.pdf.
  • the Peptide Embedding Model 204 and/or the Protein Embedding Model 202 can comprise an Evolutionary Scale Model (ESM), e.g., a transformer- based LLM trained on unlabeled data for over 250 million distinct protein sequences (specifically, using all sequence data from UniParc - see UniProt Consortium, The universal protein resource (UniProt). Nucleic Acids Res.36, D190–D195 (2008)) using masked language modeling.
  • ESM Evolutionary Scale Model
  • ESM was shown to generate a vector representation for an input protein sequence that can encode biochemical properties of amino acids in the sequence, the biological variation and remote homology of the sequence, and secondary structure and tertiary contacts of the sequence (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118).
  • MHC embedded representation 206 and Peptide embedded representation 208 are then used as the input to trained machine learning model (Structural Constraint Prediction Model 141) configured to process the embedded sequence data and output, for example, Pairwise Distance Data 210 (i.e., predicted pairwise distance data for at least one peptide amino acid residue of the peptide and at least one protein amino acid residue of the protein in a peptide-protein complex) and Dihedral Angle Data 212 (i.e., predicted dihedral angle data for at least one peptide amino acid residue in the peptide in the peptide-protein complex).
  • Pairwise Distance Data 210 i.e., predicted pairwise distance data for at least one peptide amino acid residue of the peptide and at least one protein amino acid residue of the protein in a peptide-protein complex
  • Dihedral Angle Data 212 i.e., predicted dihedral angle data for at least one peptide amino acid residue in the peptide in the peptide-protein complex.
  • a posterior distribution sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 31 of 108 can be characterized by, for example, a mean value, and a standard deviation or any other measure of uncertainty / variability (or any other set of statistical parameters for characterizing a distribution, which depend on the type of distribution that is inferred).
  • Structural Constraint Prediction Model 141 can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution).
  • Structural Constraint Prediction Model 141 can comprise a trained neural network, as described in more detail elsewhere herein.
  • the trained neural network is a deep neural network.
  • the trained neural network can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, a trained recurrent neural network, a trained state space model, or other suitable type of neural network or deep learning model.
  • Structural Constraint Prediction Model 141 comprises a multilayer perceptron.
  • the method can further comprise iterating over a plurality of candidate peptide sequences, or portions thereof, for at least one MHC class I or class II protein.
  • the candidate peptide sequences, or portions thereof may be ranked according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex and/or one or more features of one or more peptide-protein complexes comprising the candidate peptide sequences and MHC I or MHC II proteins, predicted using the conformer ensembles 148 obtained for the respective complexes (such as e.g.
  • peptide-MHC-TCR complexes or other complexes comprising a peptide-MHC and antigen recognition molecule.
  • metrics such as scores estimated from one or more sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 32 of 108 energy functions computed for a Conformer Ensemble can be used to determine whether a peptide-MHC complex is likely to be formed, to quantify the stability of a peptide-MHC complex (e.g. by quantifying the energy of the peptide-MHC complex using an energy function), and/or to quantify the binding of a peptide to an MHC molecule (e.g.
  • scores estimated from one or more energy functions computed for a plurality of Conformer Ensembles each comprising a given MHC allele and a candidate peptide with a candidate peptide sequence can be used to rank the candidate peptide sequences, such as e.g.
  • the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof, isolated from a subject (e.g., a patient).
  • the peptides can comprise mutated peptides (e.g., neoantigens expressed in tumor cells), or portions thereof, isolated from a subject (e.g., a patient).
  • the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium.
  • the peptides can comprise viral peptides, or portions, thereof expressed in a virus.
  • the method can further comprise developing a personalized anti- cancer vaccine, an anti-viral vaccine, an anti-bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof.
  • FIG.3A depicts an example process for generating a peptide sequence embedding, in accordance with some embodiments.
  • a protein language model (PLM) 302 (corresponding to peptide embedding model 204 in FIG.2) can receive a peptide sequence and output a peptide sequence embedding 208 (corresponding to peptide embedded representation 208 in FIG.2).
  • PLMs are machine-learning models (e.g., deep-learning models) sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 33 of 108 that can be based on natural language processing methods such as e.g. attention-based models (including but not limited to transformers) and can be trained on unlabeled sets of protein sequences.
  • Protein language models can be trained to understand (i.e. using self-supervised learning, sometimes referred to as “pre-training”) and also optionally to predict one or more properties of peptides and/or proteins (sometimes referred to as fine tuning or task specific training or downstream task training) based on the amino acid sequence forming such a peptide and/or protein.
  • protein language models can infer a range of characteristics from amino acid sequences, including primary, secondary, tertiary, and quaternary structures of peptides and/or proteins as applicable.
  • PLMs can be used e.g. to predict how proteins fold, what domains are present, where active sites are located, and how stable a protein is.
  • PLMs can also forecast protein-protein, protein-peptide, and protein-nucleic acid interactions, post-translational modifications, and the effects of mutations.
  • PLMs can identify localization signals within a cell, understand evolutionary relationships, and predict protein function.
  • the PLM 302 can comprise an Evolutionary Scale Modeling (ESM model) or a variation of the ESM model.
  • ESM model Evolutionary Scale Modeling
  • the PLM 302 can comprise ProteinBERT, UniRep, or other suitable type of PLM (e.g. any protein language model known in the art).
  • the PLM 302 comprises a pretrained protein language model such as a pretrained ESM model.
  • a pretrained PLM can refer to a PLM that has been trained in a self-supervised manner to learn a representation of protein sequences (i.e.
  • the input peptide sequence 126 can include a sequence of amino acid residues and the PLM 302 can be configured to obtain a plurality of embeddings (i.e., vector representations) by obtaining, for each amino acid residue, a corresponding embedding.
  • the model can be further configured to obtain a single embedding 208 by aggregating the plurality of embeddings corresponding to the sequence of amino acid residues (e.g., by performing element-wise averaging).
  • FIG. 3B depicts an example process for generating an MHC sequence (or other protein sequence) embedding, in accordance with some embodiments.
  • a PLM 304 (which may be the same model or a different model than PLM 302 illustrated in FIG.3A, and which corresponds to Protein Embedding Model 202 in FIG.2) can receive an MHC sequence 135 and output an MHC sequence embedding 206 (corresponding to MHC embedded representation 206 in FIG. 2).
  • a PLM can be trained using a large number of proteins and thus can encode useful information and context about the input sequence as represented by the sequence embedding.
  • the PLM 304 can comprise an Evolutionary Scale Modeling (ESM model) or a variation of the ESM model.
  • the PLM 304 can comprise ProteinBERT, UniRep, or the like.
  • the PLM 304 comprises a pretrained protein language model such as a pretrained ESM model.
  • the PLM304 and the PLM 302 both comprise an ESM model, a ProteinBERT model, a UniRep model, or a variation of any of these.
  • the PLM304 and the PLM 302 both comprise the same protein language model.
  • the MHC sequence 135 can include a plurality of amino acids that compose a corresponding allele.
  • the plurality of residues may be consecutive or non-consecutive.
  • the MHC sequence 135 may be a pseudo-sequence.
  • the pseudo-sequence may be, for each MHC allele (also referred to herein as “MHC molecule” or “HLA allele” or “HLA molecule”) a sequence of 34 amino acids at positions typically found within 4 ⁇ from peptides bound to the MHC allele.
  • the PLM 304 can be configured to obtain a plurality of embeddings (i.e., vector representations) by obtaining, for each amino acid, a corresponding embedding.
  • the model can be further configured to obtain a single embedding 206 by aggregating the plurality of embeddings corresponding to the sequence of amino acid residues (e.g., by performing element-wise averaging).
  • FIG.4 provides a schematic illustration of using a trained machine learning model (e.g., Structural Constraint Prediction Model 141) to output pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a pMHC complex) based on input peptide and protein (e.g., MHC) embeddings.
  • a trained machine learning model e.g., Structural Constraint Prediction Model 141
  • MHC protein
  • Peptide embedded representation 208 and MHC embedded representation 206 are provided as input to Structural Constraint Prediction Model 141, which can be configured to infer posterior distributions of Pairwise Distance Data 210 for peptide and protein amino acid sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 35 of 108 residues in the peptide-protein (e.g.
  • the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict one or more parameters characterizing respective distributions (e.g.
  • the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict one or more parameters characterizing respective distributions (e.g. posterior distributions) of one or more dihedral angles for each of one or more peptide amino acid residues in the peptide-protein complex.
  • the Structural Constraint Prediction Model 141 is configured (i.e.
  • the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict both: (i) one or more parameters characterizing respective distributions of distances between respective pairs of peptide and protein amino acid residues in the peptide-protein complex, and (ii) one or more parameters characterizing respective distributions (e.g. posterior distributions) of one or more dihedral angles for each of one or more peptide amino acid residues in the peptide-protein complex.
  • the Structural Constraint Prediction Model 141 may be trained to predict individual values (e.g. samples of distributions of pairwise distances and/or dihedral angles), to which corresponding distributions can be fitted, thereby identifying said distributions.
  • the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter that can be characterized by, for example, a mean value, a standard deviation, or any other set of statistical parameters for characterizing a distribution.
  • the predicted pairwise distance data and/or the predicted dihedral angle data can comprise one or more parameters that characterize each of one or more respective distributions for each pairwise distance and/or dihedral angle for which a prediction is made.
  • references to predicted pairwise distances between amino acid residues sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 36 of 108 can refer to distances between selected heavy atoms in the backbone of the residues.
  • the selected heavy atoms may be the alpha carbons (C ⁇ ) of the respective residues.
  • a “pairwise distance” may be a distance between the alpha carbons of two amino acid residues.
  • the two amino acid residues comprise an amino acid residue in the peptide and an amino acid residue in the protein.
  • References to predicted dihedral angles can refer to prediction of the ⁇ and/or ⁇ dihedral angle of a peptide residue (e.g.
  • Structural Constraint Prediction Model 141 can comprise a trained neural network, as described in more detail elsewhere herein.
  • the trained machine learning model can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network.
  • peptide-protein complex e.g., a peptide – MHC complex
  • a trained machine learning model e.g., Structural Constraint Prediction Model 141
  • Conformer Ensembles 148 for the peptide-protein complex e.g., a pMHC complex
  • the predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 are used to select an initial Template Structure 502 for the pMHC complex from a set of example pMHC experimental structures, such as crystal structures (e.g., crystal structures available in a protein structure database such as the Protein Data Bank (PDB)) based on a maximum likelihood estimate (MLE) for the predicted posterior distributions for pairwise distance and dihedral angles.
  • an initial Template Structure 502 for a peptide-protein complex is selected from a set of example peptide-protein experimental structures each comprising a peptide that has the same length (i.e. same number of amino acids) as the peptide in the peptide- protein (e.g.
  • selecting the initial Template Structure 502 from the set of example peptide-protein experimental structures comprises: (i) for each example peptide-protein experimental structure in the set, determining an average likelihood (which can be in practice calculated as a log-likelihood) of the distances and dihedral angles in the peptide-protein experimental structure using the predicted distributions for pairwise distances and dihedral angles for the peptide-protein complex; and (ii) selecting the example peptide-protein experimental structure that has the highest average likelihood.
  • an initial Template Structure 502 for a peptide-protein complex is selected from a set of example peptide-protein experimental structures each a representative experimental structure for a cluster of peptide- protein experimental structures.
  • an initial Template Structure 502 for a peptide-protein complex comprises a plurality of template structures.
  • each template structure is selected from a set of example peptide-protein experimental structures each a representative experimental structure for a cluster of peptide- protein experimental structures.
  • a plurality of clusters of peptide-protein experimental structures each comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled may be obtained using a clustering algorithm and a predetermined structural similarity metric.
  • a representative experimental structure may be identified for each cluster, thereby obtaining a plurality of representative experimental structures.
  • One or more or each of the representative experimental structures may be used as an initial Template Structure 502.
  • the representative experimental structures may be cluster centroids, such as e.g. structures that are closest to the center of the respective cluster using a predetermined distance metric.
  • the clusters may be obtained by structural clustering of peptide backbones in a set of example peptide-protein experimental structures comprising peptides of a predetermined length (e.g. the length of the peptide for which an initial Template structure is sought).
  • the clusters may be obtained using an agglomerative clustering method and a predetermined distance metric.
  • the initial Template Structure 502 can then be used as input to Ensemble Generation Engine 142.
  • the term “experimental structures” refers to experimentally determined atomic 3D coordinates for a peptide-protein complex. Such structures may be determined using e.g. X-ray crystallography or nuclear magnetic resonance (NMR).
  • crystal structures as used herein refers to atomic 3D coordinates for a peptide-protein complex experimentally determined by x-ray crystallography.
  • the predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 also provide structural constraints on the allowable conformations of the pMHC complex, and are also provided as input to Ensemble Generation Engine 142.
  • a plurality of compatible structures for the pMHC complex (e.g., Conformer Ensemble 148) may then be generated based on the initial Template Structure 502 and the constraints provided by the predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 using Ensemble Generation Engine 142.
  • Compatible structures for the pMHC complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance and dihedral angle data, and (ii) optimize the structure scoring function.
  • Generation of a plurality of compatible structures may comprise, for example, inputting the initial pMHC template structure into a structural modeling software package (e.g., FlexPepDock (Rosetta, www.rosettacommons.org), described in Raveh et al. “Sub-angstrom modeling of complexes between flexible peptides and globular proteins”, Proteins, Vol. 78, Issue 9, July 2010, pp.
  • FlexPepDock Rosetta, www.rosettacommons.org
  • the structural modeling software package can be configured to iteratively sample (e.g., using a Markov Chain – Monte Carlo (MCMC) sampling algorithm) from a plurality of possible structures for the pMHC complex and evaluate each sampled structure using a structure scoring function (e.g., a proxy for a free energy calculation, as discussed in more detail elsewhere herein) to identify the plurality of compatible structures.
  • MCMC Markov Chain – Monte Carlo
  • the predicted pairwise distance data and dihedral angle data can be used as constraints in a peptide docking protocol (also referred to as “refinement protocol” or “refinement procedure”).
  • the predicted pairwise distance data and dihedral angle data can be used in an energy function used in the peptide docking protocol.
  • predicted pairwise distance data can be used as a harmonic constraint in the peptide docking protocol (e.g., it can be introduced as a harmonic function in an energy function using in the peptide docking protocol – see e.g. Alford et al. J. Chem. Theory Comput. 2017, 13, 6, 3031–3048).
  • Predicted dihedral angle data can be used as a circular harmonic constraint in the peptide docking protocol (e.g., it can be introduced as a circular harmonic function in an energy function using in the peptide docking protocol – see e.g. Alford et al. J. Chem. Theory Comput. 2017, 13, 6, 3031–3048).
  • generating a plurality of compatible structures for the peptide-protein complex comprises performing a pre- packing of the template to remove internal clashes in the protein and the peptide.
  • generating a plurality of compatible structures for the peptide-protein complex comprises optimizing the structure of the peptide backbone relative to the receptor protein (e.g., using a MCMC algorithm with Minimization approach), together with on-the-fly side-chain optimization.
  • peptide backbone optimization is repeated a predetermined number of times (N), generating a number (N) of independent optimization trajectories.
  • the peptide docking protocol uses a coarse grained model of the peptide and protein.
  • the peptide backbone and peptide side-chains are optimized iteratively (e.g. using a MCMC algorithm), combined with rigid-body placement of the peptide with respect to the protein.
  • an initial structure model e.g. coarse grained model
  • a template structure i.e. a template peptide-protein complex structure that is experimentally derived
  • the initial structure model e.g. coarse grained model
  • a threading protocol such as e.g.
  • the homology model is subjected to a restricted refinement stage where only the residues of the peptide and the residues of the protein within a predetermined distance with a peptide residue (e.g. 3.5 ⁇ ) are refined (e.g. in the Rosetta force field or corresponding energy function in any other structural modelling environment).
  • the initial structure model is then used for generating the plurality of compatible structures using a structural modeling software package as described above.
  • Conformer Ensemble 148 may be used in downstream processing to, for example, predict features of the pMHC complex (e.g., interface features, surface complementarity features, etc.).
  • FIG. 6 provides a non-limiting schematic illustration of a process for training a machine learning model to predict pairwise distance data and dihedral angles for a pMHC complex.
  • the model was configured to accept MHC embedded representation 608 and peptide embedded representations 606 for an MHC pseudo-sequence with a fixed length of 34 residues, 604, and a peptide sequence of variable length, 602, as input.
  • the task performed by the trained model can be probabilistic regression over the structural properties of the pMHC complex.
  • the model can be sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 40 of 108 trained to estimate the posterior distribution over the pairwise distances between the peptide and MHC amino acid residues and the dihedral angles of the peptide residues.
  • the parameterization of the posterior distribution consists of a factorized Gaussian and the peptide is, e.g., a 9-mer
  • the model outputs the Gaussian mean and standard deviation parameters for the pairwise distances between the MHC and peptide residues (34x9 distances in total) and likewise for the eight phi and psi dihedral angles.
  • the model architecture consisted of a large pretrained model, such as an Evolutionary Scale Model (ESM) (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118; Verkul et al. (2022), “Language Models Generalize Beyond Natural Proteins", bioRxiv 2022-12), and included additional layers dedicated specifically to the probabilistic regression task.
  • ESM Evolutionary Scale Model
  • the training was carried out in two stages, as illustrated in FIG.6. In the first stage, the ESM portion of the model was fine-tuned for peptide-MHC binding classification, 610.
  • a multi-layer perceptron (MLP) with a single (logit) output was appended to the concatenated ESM embeddings of the peptide and MHC.
  • the embeddings of the ESM portion of the model were mean-aggregated across the residue positions and processed by a MLP head (labeled ⁇ Classifier'') with a single (logit) output.
  • the MLP stage was included to take advantage of the wealth of publicly-available pMHC sequences annotated with binary binding labels. For instance, the model can be pre-trained on the labeled dataset of 2 million sequence pairs curated by Chu et al.
  • the classification task can be expected to provide the model, pretrained on general proteins, with an understanding of valid binding interfaces for a wide variety of peptides and MHCs.
  • the binary cross-entropy / log loss function which provides a measure of the difference between predicted binary outcomes and actual binary labels, was used for training.
  • N is the number of data points
  • the y i are the actual labels in the distribution of training data points (q(y))
  • p(y i ) is the probability that data point y i is in the positive class predicted by the model.
  • Supervised training of the regression task requires crystal structure data for pMHC complexes (e.g., the atomic coordinates for each atom in each amino acid residue of the proteins), of which there are only a few hundred examples available in the Protein Data Bank (PDB).
  • PDB Protein Data Bank
  • Each training instance was defined by data for a single MHC residue and a single peptide residue.
  • the model Given the residue pair, the model extracts the slices corresponding to the residues from the ESM embeddings, concatenates them, and passes the resulting tensor through the regressor MLP.
  • the forward call of the model extracts the slices corresponding to a given residue pair from the ESM embeddings (where a “slice” is a peptide residue or MHC residue embedding), concatenates them, and passes the resulting tensor through the new MLP. Every residue pair from the peptide and pseudo-MHC sequence is generated.
  • the output of the model comprises pairwise distance data 614 and dihedral angle data 616, and varies depending on the parameterization used to characterize the predicted posterior distributions.
  • the output of the model can comprise predicted distributions of c-alpha distances for all pairs of amino acid residues (e.g., a number of distributions equal to the number of residues of a pseudosequence of a MHC molecule (typically 34) multiplied by the length of the peptide).
  • the output of the model can comprise joint distributions of phi and psi dihedral angles for a number of residues of the peptide corresponding to the length of the peptide minus 1 (since dihedral angles are between a first residue in Nter and a subsequent residue Cter of the first residue).
  • the model may output the three mean values and a 3 x 3 covariance matrix that govern the trivariate Gaussian, where the diagonal elements of the 3 x 3 covariance matrix comprise the squares of the three standard deviation values, and the off diagonal elements are the covariance values.
  • the parameterization used consists of a factorized (diagonal) Gaussian, the output could simplify to, for example, three sets of mean and standard deviation values.
  • a Gaussian model to characterize the posterior distribution of predicted pairwise distance and dihedral angle data
  • other options include a Gaussian mixture model, (joint or factorized), von Mises distributions sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 42 of 108 for the angles, and a mixture thereof.
  • the loss function used for training the new MLP 612 was the negative log-likelihood (NLL): evaluated at the distance and angle labels, where again, N is the number of data points and p(y i ) is the probability that data point y i has a true label (i.e.
  • the model is trained to output the predictive distribution over the pairwise distances between the peptide and HLA residues as well as the ⁇ , ⁇ dihedral angles of the peptide residues. This can be denoted as: Where L is the peptide length, N is the protein sequence (e.g. pseudosequence) length ⁇ ⁇ the matrix of pairwise distances, are the peptide dihedral angles, and all terms are conditioned on the training set.
  • the predictive distribution is parameterized as a Gaussian mixture with K components (although as explained above other distributions can be used).
  • ⁇ , ⁇ ) denotes the Gaussian density with mean and variance parameters ⁇ , ⁇ , such that ⁇ ( ⁇ ) ⁇ ⁇ , ⁇ ⁇ 2 ( ⁇ ) ⁇ ⁇ R+ are the parameters of the Gaussian component k governing each distance element and likewise for ⁇ R governing each angle element.
  • the ⁇ [0,1] are the mixture weights that sum to 1 across k ⁇ ⁇ 1, ... and likewise for .
  • the value of K (number of components) may be set independently for every predictive distribution, or may be set to be the same for a plurality of (e.g.
  • the value(s) of K may be set as a hyperparameter, e.g. using cross-validation.
  • the model is trained to predict a set of distributional parameters for an input protein sequence embedding and an input peptide sequence embedding, the distributional parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 43 of 108 characterizing the distributions of pairwise distances between pairs of a peptide amino acid residue and a protein amino acid residue ( ⁇ ⁇ ⁇ , ⁇ ), and the joint distribution of dihedral angles for peptide amino acid residues.
  • the structural constraint prediction model (e.g. new MLP 612) is trained using maximum likelihood estimation.
  • the structural constraint prediction model (e.g. new MLP 612) is trained using a loss function defined as: where w is a hyperparameter, L is the length of the peptide, N is the length of the protein sequence (e.g., pseudosequence), and the subscript ⁇ refers to the parameters of the structural constraint prediction model (e.g. new MLP 612).
  • the hyperparameter w may be set to a predetermined value.
  • the structural constraint prediction model e.g. new MLP 612
  • the structural constraint prediction model is trained using a loss function that weighs training losses corresponding to peptides of a first predetermined length for which more training data is available less than training losses corresponding to peptides of a second predetermined length for which less training data is available.
  • a masking step may be implemented so that the missing information is excluded from the training data used to train and/or optimize the model.
  • the joint distribution over phi, psi for an amino acid residue can be decomposed as follows:
  • the joint distribution over phi, psi for an amino acid residue can be decomposed as follows: sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 44 of 108 [0135]
  • the distribution, ⁇ can be optimized using either the left-hand expression or the right-hand expression. Evaluating the left-hand expression requires measurements of both phi and psi.
  • the residue may not be discarded from the training data, and instead the distribution of ⁇ ⁇ , ⁇ ⁇ , ⁇ may be factorized depending on the availability of the labels as p( ⁇ j
  • FIG. 7 provides a non-limiting example that illustrates the process for how predicted distributions of pairwise distance data and dihedral angle data for a pMHC complex can be converted to structural constraints, and how pMHC conformational ensembles can then be generated based on such structural constraints.
  • An MHC sequence or pseudosequence i.e., the ordered set of amino acids of the MHC molecule that typically contacts a peptide
  • amino acid residues that typically contact a bound peptide e.g., 34 amino acid residues (here illustrated as YFAMYGEKV...(SEQ ID NO:1))
  • a peptide sequence such as e.g., a 9-mer peptide sequence (here illustrated as LLFGYPVYV (SEQ ID NO:2)
  • a trained machine learning model e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.
  • pairwise distance data here illustrated as mean (702, left panel) and standard deviation (704, middle panel) values characterizing the distribution of pairwise distance data, for all pairs of peptide and MHC pseudosequence amino acid residues, and to predict dihedral angle data, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 45 of 108 here illustrated as mean and standard deviation values (706, right panel), for peptide dihedral angles for all peptide residues in the complex.
  • these predictions can then be used as constraints within a structural modeling software, such as the FlexPepDock (Rosetta, www.rosettacommons.org) software package, to generate conformational ensembles for the pMHC complex, together with a template structure model selected based on the predicted pairwise distance data and dihedral angle data.
  • the selected template structure (model 802, upper left panel) was based on the predicted pairwise distance and dihedral angle constraints as applied to pMHC crystal structure PDBID: 4NO5, and was used to create the starting model (model 804, upper right panel) for conformer generation using FlexPepDock. Conformer generation was then performed using the iterative process described elsewhere herein.
  • FIG. 9 provides a non-limiting example of a process 900 (e.g., a computer- implemented method) for generating a plurality of compatible structures for a peptide-protein complex.
  • Process 900 can be implemented using the Prediction System 100 described in FIG. 1.
  • Process 200 can be implemented using the Sequence Analyzer 108 and one or more of the Machine Learning Models 132 in FIG.1.
  • peptide sequence data for at least one peptide
  • protein sequence data for at least one protein
  • a trained machine learning model e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1
  • predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in a peptide-protein complex
  • each predicted pairwise distance data comprising one or more predicted distances and/or a one or more parameters characterizing a predicted distribution of distances between a pair of amino acid residues comprising a peptide residue and a protein amino acid residue in the peptide-protein complex); and (ii) predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex (e.g.
  • each predicted dihedral angle data comprising, for a residue of the peptide in the peptide-protein complex: one or more predicted angles for a ⁇ angle, one or more predicted angles for a ⁇ angle, one or more parameters characterizing a predicted distribution of ⁇ angle, one or more parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 46 of 108 characterizing a predicted distribution of ⁇ angle, and/or one or more parameters characterizing a predicted joint distribution of ⁇ and ⁇ angle).
  • the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter (e.g.
  • the predicted pairwise distance data can comprise one or more predicted parameters that together characterize a distribution of pairwise distance between a pair of amino acid residues comprising at least one peptide residue and at least one protein amino acid residue in the peptide-protein complex.
  • the predicted dihedral angle data can comprise one or more parameters that together characterize a distribution of ⁇ angle, a distribution of ⁇ angle and/or a joint distribution of ⁇ and ⁇ angles for at least one amino acid residue in the peptide that forms part of the peptide-protein complex.
  • the trained machine learning model can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution).
  • the predicted pairwise distance data is predicted for all pairs comprising a peptide amino acid residue and a protein pseudosequence residue, where a protein pseudosequence residue is a residue of the protein that is believed to be involved in the interaction with peptides when forming a peptide-protein complex.
  • the predicted dihedral angle data is predicted for all amino acid residues of the peptide apart from the last (C-ter-) amino acid residue.
  • the predicted dihedral angle data is predicted for all amino acid residues of the peptide core (also referred to as minimal epitope, in the context of peptide-MHC complexes).
  • the input can comprise peptide sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more than 50 peptides, or portions thereof.
  • peptide sequence data can comprise sequence data for peptides that are at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues in length.
  • the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof, isolated from a subject (e.g., a patient), derived from nucleic acid sequences isolated from a subject, or obtained from one or more sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 47 of 108 genome, transcriptome and/or proteome reference dataset (e.g. a genome, transcriptome or proteome database).
  • normal peptides e.g., peptides expressed in normal cells
  • a subject e.g., a patient
  • nucleic acid sequences isolated from a subject e.g., a subject
  • proteome reference dataset e.g. a genome, transcriptome or proteome database
  • the peptides can comprise mutated peptides (e.g., neoantigens expressed or likely to be expressed in tumor cells), or portions thereof, isolated from a subject (e.g., a patient) or derived from nucleic acid sequences isolated from a subject.
  • the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium.
  • the peptides can comprise viral peptides, or portions, thereof expressed in a virus.
  • the peptides can comprise peptides from one or more pathogens, such as e.g. bacteria, viruses, parasites (e.g. plasmodium), etc.
  • the peptide sequence data for the at least one peptide can comprise a peptide embedded representation (or peptide embedding, e.g., a machine-friendly vector representation of the peptide sequence) for the at least one peptide generated using a trained peptide embedding machine learning model.
  • the vector representation of the peptide sequence can encode the peptide sequence in terms of the statistical dependencies and patterns present in the peptide sequences used to train the model.
  • the peptide sequence data can comprise a peptide embedded representation from a trained protein language model, such as e.g. peptide embedding model 204.
  • the peptide embedding model 204 may have been trained in a self-supervised manner as part of a protein language model.
  • the peptide embedding model 204 may have been further trained (also referred to as “fine-tuned”) as part of a machine learning model configured to predict a property of a peptide-protein complex comprising the peptide. Further, multiple consecutive fine-tuning steps may have been implemented comprising further training the peptide embedding model 204 as part of a machine learning model configured to predict a different property of the peptide-protein complex comprising the peptide.
  • the properties of the peptide-protein complex may be selected from: functional properties and structural property.
  • the machine learning model configured to predict a property of the peptide-protein complex may comprise a peptide embedding model, a protein embedding model, and a task specific prediction head.
  • the task specific prediction head may be a classification head or a regression head.
  • a predicted functional property may be binding between the peptide and the protein to form the peptide-protein complex.
  • the binding between the peptide and the protein may be predicted as a binary property (binding / non-binding) using a classification head, or as a continuous property (e.g. predicted binding affinity) using a regression head.
  • a predicted functional property may be presentation of the peptide by the protein on the surface of cells, where the protein is a MHC molecule.
  • the presentation of a peptide by a MHC molecule may be predicted as a binary property (e.g.
  • a predicted functional property may be stability of the peptide-protein complex.
  • the stability of a peptide-protein complex may be predicted as a binary (e.g. stable/unstable) or continuous (e.g. measured complex half-life) property.
  • a predicted structural property may be prediction of pairwise distance data and/or dihedral angle data as described herein.
  • a task specific prediction head may be a multilayer perceptron.
  • predicting a binary property can comprise predicting a probability that the peptide-protein complex comprising the peptide belons to a class of a pair of class (e.g. predicting a probability that the peptide binds to the protein).
  • Peptide-MHC binding affinity data, binding data and presentation data is available in publicly available databases and datasets such as e.g. Chu et al. (2022) as mentioned above, and the Immune Epitope Database & Tools (IEDB, www.iedb.org/home_v3.php).
  • the input can comprise protein sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, or more than 1000 proteins, or portions thereof.
  • protein sequence data can comprise sequences or pseudosequences of at least 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or at least 200 amino acid residues in length.
  • the protein sequence data for at least one protein can comprise a protein embedded representation (or protein embedding, e.g., a machine-friendly vector representation of the protein sequence) for the at least one protein generated using a trained protein embedding machine learning model.
  • the vector representation of the protein sequence can encode the protein sequence in terms of the statistical dependencies and patterns present in the protein sequences used to train the model.
  • the protein sequence data can comprise a protein embedded representation from a trained protein language model, such as e.g. protein embedding model 202.
  • the protein embedding model 202 may have been trained in a self-supervised manner as part of a protein language model.
  • the protein embedding model 202 may have been further trained (also referred to as “fine-tuned”) as part of a machine learning model configured to predict a property of a peptide-protein complex comprising the protein, as explained above in relation to the peptide embedding model 204.
  • the trained peptide embedding machine learning model and the trained protein embedding machine learning model can be different models (e.g., Peptide Embedding Model 204 and Protein Embedding Model 202, respectively, as illustrated in FIG. 2).
  • the trained peptide embedding machine learning model and the trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 49 of 108 protein embedding machine learning model can be the same model (e.g., the Embedding Generation Engine 140 illustrated in FIG. 1 and FIG. 2).
  • Example methods for performing peptide and/or protein sequence embedding are described in, e.g., Ofer et al.
  • the trained peptide embedding machine learning model and the trained protein embedding machine learning model can be based on the same model obtained from a protein language model (i.e. a protein/peptide embedding model that has been trained as part of a protein language model in a self-supervised manner).
  • Two instances of the protein/peptide embedding model may have been included in a machine learning model that is trained to predict one or more properties of a peptide-protein complex, in order to obtain the trained protein embedding model and trained peptide embedding model (i.e. further training / fine tuning the pretrained protein embedding model and peptide language model).
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise large language models (LLMs), e.g., deep learning algorithms often based on transformer networks that can recognize, summarize, translate, predict, and generate content using very large data sets (e.g., very large peptide and/or protein sequence data sets).
  • LLMs large language models
  • Large language models are described in more detail in, for example, Li et al. (2021), “BioSeq-BLM: A Platform for Analyzing DNA, RNA and Protein Sequences Based on Biological Language Models”, Nucleic Acids Res.49, e129–e129, and Naveed et al.
  • a protein language model can refer to a large language model that has been trained in a self-supervised manner using data from one or more proteomic databases.
  • a protein language model may be a large language model that has been trained to model the “language” of proteins using large amounts of unlabeled protein sequence data.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise an Evolutionary Scale Model (ESM), e.g., a transformer-based LLM trained on unlabeled data for over 250 million distinct protein sequences that generates a vector representation for an input sequence.
  • ESM Evolutionary Scale Model
  • ESM has been shown to be able to encode biochemical properties of amino acids in the sequence, the biological variation and remote homology of the sequence, and secondary structure and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 50 of 108 tertiary contacts of the sequence (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118).
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise fine-tuned versions of an ESM model, where fine tuning refers to further training of the models as part of a machine learning model trained to predict one or more properties of a peptide-protein complex.
  • the trained machine learning model configured to determine predicted pairwise distance data and dihedral angle data can comprise a trained neural network, as described in more detail elsewhere herein.
  • the trained machine learning model can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network.
  • MLP multilayer perceptron
  • the machine learning model configured to determined predicted pairwise distance data and dihedral angle data can be trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide-protein combinations; training a classifier (or regressor) to predict peptide-protein binding (or other functional properties of a peptide-protein complex) for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and appending a regressor head to at least a portion of the trained classifier (or regressor) and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle
  • the classifier can be a machine learning model comprising a protein embedding model, a peptide embedding model and a classification head.
  • a classification head in this context can be a machine learning model configured to take as inputs a peptide embedded representation from the peptide embedding model, and a protein embedded representation from the protein embedding model, and provide as output a prediction indicative of a class of a plurality of classes that the peptide-protein complex belongs to.
  • the plurality of classes may comprise a class of peptide-protein pairs that bind to each other to form a peptide-protein complex and a class of peptide-protein pairs that do not bind to each other to form a peptide-protein complex.
  • the regressor can be a machine sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 51 of 108 learning model comprising a protein embedding model, a peptide embedding model and a regression head.
  • a regression head in this context can be a machine learning model configured to take as inputs a peptide embedded representation from the peptide embedding model, and a protein embedded representation from the protein embedding model, and provide as output a prediction indicative of binding affinity between the protein and the peptide (e.g. a predicted binding affinity).
  • Appending a regressor head to at least a portion of the trained classifier/regressor and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes can comprise removing the previous classifier/regression head and replacing it with a new regressor head, then training the resulting machine learning model to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes.
  • Crystal structure data may refer to pairwise distances and dihedral angles obtained from experimentally determined structural coordinates of peptide-protein complexes.
  • Pairwise distances and dihedral angles can be calculated from structural coordinates obtained from publicly available structure databases such as e.g. the PDB.
  • the peptide sequence data for the plurality of peptides and/or the protein sequence data for the plurality of proteins used for training can each independently comprise sequence data for at least 2 x 10 4 , 4 x 10 4 , 6 x 10 4 , 8 x 10 4 , 1 x 10 5 , 2 x 10 5 , 3 x 10 5 , 4 x 10 5 , 5 x 10 5 , 6 x 10 5 , 7 x 10 5 , 8 x 10 5 , 9 x 10 5 , 1 x 10 6 , 2 x 10 6 , 3 x 10 6 , 4 x 10 6 , 5 x 10 6 , 6 x 10 6 , 7 x 10 6 , 8 x 10 6 , 9 x 10 6 , 1 x 10 7 , or more than 1 x 10 7 peptide or
  • the peptide-protein binding data for a plurality of peptide-protein combinations used for training can comprise binding data (e.g., binary (yes/no) binding data and/or continuous valued binding affinity data) for at least 2 x 10 4 , 4 x 10 4 , 6 x 10 4 , 8 x 10 4 , 1 x 10 5 , 2 x 10 5 , 3 x 10 5 , 4 x 10 5 , 5 x 10 5 , 6 x 10 5 , 7 x 10 5 , 8 x 10 5 , 9 x 10 5 , 1 x 10 6 , 2 x 10 6 , 3 x 10 6 , 4 x 10 6 , 5 x 10 6 , 6 x 10 6 , 7 x 10 6 , 8 x 10 6 , 9 x 10 6 , 1 x 10 7 , or more than 1 x 10 7 peptide- protein combinations.
  • binding data e.g., binary (yes/no) binding data and/or continuous valued
  • the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide-protein complex can be presented as a separate training instance during training of the regressor head. This provides more training data in cases where the number of peptide-protein crystal structures available is limited.
  • the training data used to train the machine learning model configured to determined predicted pairwise distance data and dihedral angle data can comprise, for each of a plurality of training pairs of amino acid residues comprising a peptide amino acid residue and a protein amino acid residue in a training peptide-protein complex, experimentally determined values of a pairwise distance between the amino acid residues in the pair and one or both dihedral angles associated with the peptide residue in the pair.
  • ground truth distances / dihedral angles
  • label simply “ground truth” or “label”.
  • the training data can therefore comprise ground truth experimentally determined pairwise distances and dihedral angles for each of a plurality of pairs of amino acid residues derived from a plurality of training experimentally determined peptide-protein complexes structures.
  • the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor.
  • Processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex can comprise obtaining pairwise distance data and dihedral angle data for the peptide-protein complex using the trained machine learning model (e.g. a trained machine learning model trained using training data without augmented data), the pairwise distance data comprising a statistical measure of variability in the pairwise distance (e.g.
  • a set of parameters characterizing a predicted distribution of the pairwise distance between a pair of amino acid residues
  • the dihedral angle data comprising a statistical measure of variability in the dihedral angle(s) (e.g. as one of a set of parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 53 of 108 characterizing a predicted distribution of one or both of the dihedral angles of the peptide amino acid residue in a pair), as described herein.
  • Applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex can comprise using the predicted measures of uncertainty to generate an ensemble of conformers using a structural modeling software and a template structure that is an experimentally determined structure for the particular peptide-protein complex.
  • the resulting conformers may then be used to derive an augmented set of “ground truth” pairwise distances and dihedral angles.
  • the augmented pairwise distance data and dihedral angle data can be obtained by applying the methods described herein to generate a conformer ensemble for a peptide-complex for which an experimentally determined structure is available, and using the resulting conformers to derive additional ground truth pairwise distances and dihedral angles for the peptide-protein complex.
  • This additional ground truth pairwise distances and dihedral angles can capture reasonably expected levels of flexibility in the peptide-protein complex that is not captured in experimentally determined crystal structures.
  • the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by selecting one or more peptide-protein pairs that are known to bind to each other to form a complex, and predicting a structure for complexes corresponding to each of said peptide-protein pairs using a structure prediction machine learning algorithm.
  • the machine learning algorithm may be selected from the AlphaFold family of models (e.g. AlphaFold2 or fine-tuned versions thereof, such as e.g. AF-FT which is fine-tuned for peptide-HLA binding predictions, see Motmaen et al. Proc. Nat. Ac.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can be integrated with, and trained simultaneously with, the trained machine learning model.
  • an initial structure also referred to herein as “template” for the peptide-protein complex can be identified based on the predicted pairwise distance data and the predicted dihedral angle data.
  • Identifying the initial structure (also referred to herein as “initial template structure” or simply “template structure”) for the peptide-protein complex based on the predicted pairwise sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 54 of 108 distance data and predicted dihedral angle data can comprise, for example, obtaining a plurality of crystal structures for peptide-protein complexes from a protein structure database; and selecting a crystal structure from the plurality of crystal structures based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted distance data and the predicted dihedral angle data.
  • Similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted distance data and the predicted dihedral angle data can be quantified using a likelihood estimate of the pairwise distance data and dihedral angle data determined for the crystal structure under distributions characterized by the predicted distance data and the predicted dihedral angle data.
  • a crystal structure may be selected from the plurality of crystal structures using a structural similarity metric determined for the crystal structure and a maximum likelihood estimate of pairwise distances and dihedral angles according to the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1).
  • the initial structure is selected based on sequence similarity between the peptide in the peptide-protein complex and the peptides of the same length in the crystal structures for peptide-protein complexes. Sequence similarity may be calculated using any scoring known in the art, such as e.g. BLOSUM62.
  • the plurality of crystal structures obtained may be crystal structures comprising the same protein as the peptide-protein complex. For example, one or more crystal structures may be available for a particular MHC molecule.
  • the plurality of crystal structures obtained may be crystal structures comprising the same protein as the peptide-protein complex, but not necessarily the same peptide.
  • the plurality of crystal structures obtained may be crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled.
  • the plurality of crystal structures may be crystal structures obtained by clustering crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled and identifying a representative structure for each cluster.
  • the initial structure (initial template structure) may comprise a plurality of structures. For peptides up to 10 amino acids, a single template structure may be used. For peptides of 11 or more amino acids, a plurality of template structures may be used.
  • the plurality of template structures may comprise crystal structures obtained by clustering crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled and identifying a representative structure for each cluster.
  • the plurality of crystal structures for peptide-protein complexes obtained from a protein structure database can comprise at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, or more than 300 crystal structures.
  • the protein structure database can comprise, e.g., the Protein Data Bank (PDB) (Research Collaboratory for Structural Bioinformatics; rcsb.org).
  • PDB Protein Data Bank
  • a plurality of compatible structures can be generated for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data.
  • each compatible structure for the peptide-protein complex can conform, for example, with a posterior distribution of the predicted pairwise distance data and with a posterior distribution of the predicted dihedral angle data output by the trained machine learning model.
  • each compatible structure for the peptide-protein complex can conform with the sequence data for the at least one peptide and the at least one protein and with the predicted pairwise distance data and the predicted dihedral angle data.
  • the plurality of compatible structure for the peptide-protein complex are used to obtain a prediction of at least one auxiliary feature.
  • generating the plurality of compatible structures for the peptide- protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data can comprise, for example, inputting the initial structure for the peptide-protein complex into a structural modeling software package (e.g., the Ensemble Generation Engine 142 illustrated in FIG.1 and FIG.2); inputting the peptide sequence data, the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, where compatible structures for the peptide-protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function.
  • a structural modeling software package e.g., the Ensemble Generation Engine 142
  • the initial structure may have already been adapted to include the peptide sequence, such as e.g. using homology modelling and an initial structure comprising a peptide that has a different sequence from the peptide in the peptide-protein complex to be modeled.
  • the iterative sampling can comprise 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1,000 iterations.
  • the plurality of possible structures can comprise 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1,000 possible structures.
  • the plurality of possible structures comprises between 10 and 1000 structures, between 50 and 1000 structures, between 50 and 500 structures, between 100 and 500 structures, between 100 and 400 structures, or about 200 structures.
  • the iterative sampling from the plurality of possible structures for the peptide-protein complex can be performed using a Markov Chain – Monte Carlo (MCMC) sampling algorithm (see, e.g., Ravenzwaaij et al. (2016) “A Simple Introduction to Markov Chain Monte–Carlo Sampling”, Psychon Bull Rev 25:143–154, and Jones et al. (2022), “Markov Chain Monte Carlo in Practice”, Annu. Rev. Stat. Appl. 9:557–578).
  • the Markov Chain – Monte Carlo (MCMC) sampling algorithm can comprise, for example, a Metropolis – Hastings algorithm, a Hamiltonian Monte Carlo algorithm, or a Langevin Monte Carlo algorithm.
  • the iterative sampling from the plurality of possible structures for the peptide-protein complex can be performed using a MCMC algorithm with a Metropolis criterion, e.g. as described in Raveh et al. Proteins, Vol. 78, Issue 9, July 2010, pp.2029-2040
  • the structure scoring function can comprise, e.g., a free energy calculation for the peptide-protein complex (e.g., a binding free energy comprising the free energy difference between the bound complex and unbound peptide and protein, or an overall interaction free energy comprising the free energy difference between the bound complex and isolated peptide and protein).
  • the structural scoring function may be as described in Raveh et al. Proteins, Vol.
  • the structure scoring function can comprise a proxy for a free energy calculation.
  • the structure scoring function can comprise a proxy for a free energy calculation of sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 57 of 108 peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof.
  • the structure scoring function (e.g., a proxy for a free energy calculation) can be based on one or more of ab initio quantum mechanical calculations, density functional theory (DFT) calculations, semi-empirical calculations, molecular mechanics force field calculations, statistical potential calculations, neural potential (or neural force field) calculations that predict the forces acting on individual atoms (e.g., performed using models that recapitulate the results of high-level quantum mechanics calculations at higher speed and/or providing more robust evaluation of bad structures), or a function parameterized by machine learning models trained on structural data.
  • DFT density functional theory
  • the structure scoring function may be a physics-based energy function that determines a total energy of the peptide-protein complex, including one or more of an electrostatic energy, covalent bonding energy, Van der Waals energy, and/or the like. It should be appreciated that different structure scoring functions may be associated with different levels of accuracy and computational complexity. For instance, a more accurate structure scoring function, such as an ab initio quantum mechanics-based structure scoring function, may impose greater computational overhead than a less accurate structure scoring function, such as a molecular mechanics force fields-based structure scoring function.
  • more than one structure scoring functions may be applied as part of identifying compatible structures for the peptide- protein complex, where compatible structures are identified as those structures having, e.g., a lowest structure score (e.g., a lowest total energy).
  • the structural modeling software package can comprise, e.g., AlphaFold (DeepMind, London, UK), Amber (ambermd.org), ESMFold (Meta AI, New York, NY), FlexPepDock (RosettaCommons, rosettacommons.org), HADDOCK (High Ambiguity Driven biomolecular DOCKing; bonvinlab.org), Molecular Operating Environment (MOE; Chemical Computing Group, Montreal, CA), OpenMM (openmm.org), Schrödinger Glide (New York, NY), and AutoDock (autodock.scripps.edu).
  • the structural modeling software package comprises FlexPepDock.
  • a structural modeling software may also be referred to as a “physics-based engine”, because it determines structures that are energetically plausible according to predetermined energetic functions.
  • Any structural modeling algorithm that can obtain a predicted structure using a template sequence, and that sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 58 of 108 can accommodate dihedral angle and/or pairwise distances constraints (e.g. within an energy scoring function) can be used in the context of the present disclosure.
  • the process (or computer-implemented method) illustrated in FIG. 9 can further comprise predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex.
  • the at least one auxiliary feature can comprise, for example, a prediction of binding of the at least one peptide to at least one other protein or protein complex, and/or a prediction of at least one interface feature for the peptide-protein complex (such as e.g. information identifying complementary interface features between the peptide-protein complex and another protein or protein complex, and/or information identifying complementary surface fingerprint features between the peptide-protein complex and another protein or protein complex).
  • the process (or computer-implemented method) illustrated in FIG.9 can further comprise displaying at least a subset of the predicted pairwise distance data and/or the predicted dihedral angle data using any of a variety of techniques for displaying three-dimensional tensor data known to those of skill in the art, e.g., 3D plots, 2D plots of a specified slice of the tensor data, etc.
  • the method can comprise displaying at least a subset of the predicted pairwise distance data determined by the trained machine learning model as, e.g., a two-dimensional heatmap comprising a plot of the pairwise distance between an alpha carbon (C ⁇ ) of a peptide amino acid residue and an alpha carbon (C ⁇ ) of a protein amino acid residue as a function of peptide amino acid residue position and protein amino acid residue position.
  • the plotted pairwise distances can comprise maximum likelihood estimates of the predicted pairwise distances, for example based on predicted parameters of respective distributions of pairwise distances.
  • the plotted pairwise distances can comprise predicted averages of predicted distributions of the pairwise distances.
  • the predicted pairwise distance data can comprise parameters characterizing the distribution of each pairwise distance for which pairwise distance data is predicted. These parameters can be used to obtain a maximum likelihood estimate of each of these pairwise distances.
  • the maximum likelihood estimate can be e.g. a predicted average (i.e. mean). This may be the case for example when parameters of Gaussian distributions are predicted (or any other distributions where the mean of the distribution is the highest likelihood value of the distribution).
  • the plotted pairwise distances can comprise multiple respective two-dimensional plots, each plot sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 59 of 108 displaying the predicted values of a parameter of a predicted distribution for pairwise distances.
  • the plotted pairwise distances can comprise a first two-dimensional plot displaying the means of predicted distributions of pairwise distances, and a second two- dimensional plot displaying the standard deviation or other statistical metric of variability of predicted distributions of pairwise distances.
  • the method can comprise displaying at least a subset of the predicted dihedral angle data determined by the trained machine learning model as, e.g., a two-dimensional scatter plot of phi and psi dihedral angles for at least a subset of the peptide amino acid residues.
  • a two-dimensional scatter plot of phi and psi dihedral angles can comprise an indicator of the means or maximum likelihood estimates of the predicted distributions of phi and psi dihedral angles for each of a plurality of peptide amino acid residues.
  • a two-dimensional scatter plot of phi and psi dihedral angles can comprise an indicator of a statistical metric of variability (e.g. standard deviation) of the predicted distributions of phi and psi dihedral angles for each of a plurality of peptide amino acid residues.
  • a scatter plot of phi angles and a scatter plot of psi dihedral angles may be plotted.
  • Each scatter plot can comprise an indicator of the means or maximum likelihood estimates of the predicted distributions of phi or psi dihedral angles for each of a plurality of peptide amino acid residues, and/or an indicator of a statistical metric of variability for the predicted distributions of phi or psi dihedral angles for each of the plurality of peptide amino acid residues.
  • the process (or computer-implemented method) illustrated in FIG.9 can further comprise predicting features of the peptide-protein complex (e.g., interface features, surface complementarity features, etc.) based on the plurality of compatible structures for the peptide-protein complex.
  • predicting binding of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex can comprise, for example, identifying complementary interface features between the peptide- protein complex and another protein or protein complex; identifying complementary surface sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 60 of 108 fingerprint features between the peptide-protein complex and another protein or protein complex; identifying latent representations of the peptide-protein complex and the other protein or protein complex; and providing the complementary interface features, the complementary surface fingerprint features, and the latent representations of the peptide-protein complex and the other protein or protein complex as input to a second trained machine learning model configured to output a prediction for binding of the peptide-protein complex and the other protein or protein complex.
  • a latent representation of a peptide-protein complex may be any representation in a space in which the peptide-protein complex is projected. This can be used e.g. to identify similarity between complexes (e.g. identify complexes that are more similar to each other) in the latent space.
  • predicting binding of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex can comprise, for example, predicting a structure of the plurality of compatible structures in complex with another protein.
  • complementary interface features can comprise, for example, pairwise lists of amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, total hydrogen bonding energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of hydrogen bonding energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, total van der Waals interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of van der Waals interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, or any combination thereof.
  • complementary surface fingerprint features can comprise, for example, steric hinderance-based complementary shape features, complementary electrostatic features, total electrostatic interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of electrostatic interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 61 of 108 total cation-pi interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of cation-pi interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, complementary polarity features, or any combination thereof.
  • the complementary surface fingerprint features between the peptide-protein complex and the other protein or protein complex can be predicted by a third trained machine learning model based on the latent representations of the peptide-protein complex and the other protein or protein complex.
  • the second trained machine learning model and/or the third trained machine learning model can comprise, for example, a trained neural network or deep learning model. Training a Classifier to Predict Pairwise Distance Data and Dihedral Angle Data [0176]
  • FIG.10 provides a non-limiting example of a flowchart for a process 1000 (e.g., a computer-implemented method) for training a machine learning model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.
  • Process 1000 can be implemented using the prediction system 100 described in FIG.1.
  • process 700 can be implemented using the Sequence Analyzer 108 and one or more of the Machine Learning Model(s) 132 in FIG. 1. All the features described above in relation to e.g. FIG. 9 may apply equally to the methods of FIG.10.
  • peptide sequence data for a plurality of peptides can be received (e.g., by one or more processors of the Prediction System 100 described in FIG.1).
  • the peptide sequence data for the plurality of peptides can comprise sequence data for at least 2 x 10 4 , 4 x 10 4 , 6 x 10 4 , 8 x 10 4 , 1 x 10 5 , 2 x 10 5 , 3 x 10 5 , 4 x 10 5 , 5 x 10 5 , 6 x 10 5 , 7 x 10 5 , 8 x 10 5 , 9 x 10 5 , 1 x 10 6 , 2 x 10 6 , 3 x 10 6 , 4 x 10 6 , 5 x 10 6 , 6 x 10 6 , 7 x 10 6 , 8 x 10 6 , 9 x 10 6 , 1 x 10 7 , or more than 1 x 10 7 peptide sequences.
  • the peptide sequence data for the plurality of peptides can comprise peptide embedded representation data (e.g., machine-friendly vector representations of the peptide sequences) for the plurality of peptides that are generated using a trained peptide sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 62 of 108 embedding machine learning model (e.g., the Embedding Engine(s) 140 illustrated in FIG.1), as described elsewhere herein.
  • peptide embedded representation data e.g., machine-friendly vector representations of the peptide sequences
  • the protein sequence data for the plurality of peptides can comprise sequence data for at least 2 x 10 4 , 4 x 10 4 , 6 x 10 4 , 8 x 10 4 , 1 x 10 5 , 2 x 10 5 , 3 x 10 5 , 4 x 10 5 , 5 x 10 5 , 6 x 10 5 , 7 x 10 5 , 8 x 10 5 , 9 x 10 5 , 1 x 10 6 , 2 x 10 6 , 3 x 10 6 , 4 x 10 6 , 5 x 10 6 , 6 x 10 6 , 7 x 10 6 , 8 x 10 6 , 9 x 10 6 , 1 x 10 7 , or more than 1 x 10 7 protein sequences.
  • the protein sequence data for the plurality of proteins can comprise protein embedded representation data (e.g., machine-friendly vector representations of the protein sequences) for the plurality of proteins that are generated using a trained protein embedding machine learning model(e.g., the Embedding Engine(s) 140 illustrated in FIG.1), as described elsewhere herein.
  • the trained protein embedding machine learning model can be the same model as the trained peptide embedding machine learning model.
  • the peptide-protein binding data for a plurality of peptide-protein combinations can comprise binding data (e.g., binary (yes/no) binding data and/or continuous valued binding affinity data) for at least 2 x 10 4 , 4 x 10 4 , 6 x 10 4 , 8 x 10 4 , 1 x 10 5 , 2 x 10 5 , 3 x 10 5 , 4 x 10 5 , 5 x 10 5 , 6 x 10 5 , 7 x 10 5 , 8 x 10 5 , 9 x 10 5 , 1 x 10 6 , 2 x 10 6 , 3 x 10 6 , 4 x 10 6 , 5 x 10 6 , 6 x 10 6 , 7 x 10 6 , 8 x 10 6 , 9 x 10 6 , 1 x 10 7 , or more than 1 x 10 7 peptide-protein combinations.
  • binding data e.g., binary (yes/no) binding data and/or continuous valued binding affinity data
  • a classifier or regressor e.g., the Structural Constraint Prediction Model 141 illustrated in FIG. 1
  • a classifier or regressor can be trained to predict peptide-protein binding for input pairs of peptide sequence data and protein sequence data using training data comprising, for each of a plurality of training peptide-protein complexes: peptide sequence data, protein sequence data, and corresponding ground truth (i.e. measured or otherwise known) peptide-protein binding data.
  • the ground truth peptide-protein binding data can comprise a binary classification label (e.g. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 63 of 108 binding/non-binding) used to train a classifier, and/or a measure of binding affinity which can be used to train a regressor or from which a binary label can be obtained using a predetermined binding affinity threshold (e.g. for training a classifier).
  • a predetermined binding affinity threshold e.g. for training a classifier.
  • the machine learning model can comprise a supervised learning model (i.e., a model trained using labeled sets of training data), an unsupervised learning model (i.e., a model trained using unlabeled sets of training data), a semi-supervised learning model (i.e., a model trained using a combination of labeled and unlabeled training data), a deep learning model (e.g., a neural network comprising many layers of coupled “nodes” that can be trained in a supervised, unsupervised, or semi-supervised manner), or any combination thereof.
  • the machine learning model can comprise one or more machine learning models (such as e.g.
  • a deep learning model may be an artificial neural network (ANN) comprising a plurality of hidden layers.
  • ANN artificial neural network
  • a deep learning model may comprise one or more attention-based layers, one or more fully connected layers, one or more convolutional layers, and combinations thereof.
  • one or more machine learning models can be utilized to implement the disclosed methods.
  • the disclosed methods can be implemented using a continuous learning approach, where the machine learning model can be periodically or continuously updated based on new training data provided by, e.g., a single local operational system, a plurality of local operational systems, or a plurality of geographically-distributed operational systems.
  • NNs neural networks
  • feedforward neural networks also known as multilayer perceptrons
  • recurrent neural networks convolutional neural networks
  • convolutional neural networks deep neural networks
  • deep recurrent neural networks deep convolutional neural networks
  • attention-based neural networks graph neural networks, or any combination thereof.
  • graph neural networks or any combination thereof.
  • Neural networks generally comprise an interconnected group of nodes organized into multiple layers of nodes.
  • the NN architecture can comprise at least an input layer, one or more hidden layers, and an output layer.
  • the NN can comprise any total number of layers (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100), and any number of hidden layers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100), where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to a predicted output value or set of output values.
  • layers e.g., 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100
  • hidden layers e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100
  • Each layer of the neural network comprises a number of nodes (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 75100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or more than 10,000 nodes).
  • a node receives input data (e.g., peptide sequence data, protein sequence data, concatenations thereof, or other types of input data) that comes either directly from one or more input data nodes or from the output of one or more nodes in previous layers, and performs a specific operation, e.g., a summation operation.
  • a connection from an input to a node is associated with a weight (or weighting factor).
  • the node can, for example, sum up the products of all pairs of inputs, X i , and associated weights, W i .
  • the weighted sum is offset with a bias, b.
  • the output of a node can be gated using a threshold or activation function, f, where f can be a linear or non-linear function.
  • the activation function can be, for example, a rectified linear unit (ReLU) activation function or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parameteric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof.
  • ReLU rectified linear unit
  • the weighting factors, bias values, and threshold values, or other computational parameters of the neural network can be “taught” or “learned” in a training phase using one or more sets of training data (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 sets of training data).
  • the parameters can be trained using the input data from a training data set and a gradient descent or backward propagation method so that the output value(s) (e.g., predicted pairwise distance data and/or dihedral angle data) that the ANN computes are consistent with the examples included in the training data set.
  • the adjustable parameters of the model can be obtained using a back propagation neural network training process that may or may not be sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 65 of 108 performed using the same hardware as that used for implementing the trained model and processing input peptide sequence and/or protein sequence data.
  • Accessing the training data set can include, for example, retrieving the training data set from a local or remote storage, loading the training data set, and/or requesting (and receiving) part or all of the training data set from one or more data stores (e.g., a cloud data storage, a server system, or some other data source).
  • the training data set can be randomly parsed, shuffled, and/or divided to train the classifier model.
  • the machine-learning model can be trained using a static or dynamic learning rate.
  • a dynamic learning rate can be produced using, for example, learning-rate annealing (e.g., using stepwise annealing or cosine annealing).
  • Training can be performed using, for example, a classification loss function and/or a regression loss function.
  • a loss function can be based on, for example, mean square error, median square error, mean absolute error, median absolute error, an entropy-based error, a cross entropy error, and/or a binary cross entropy error.
  • Validation data e.g., a separated subset of the training data set used to train the machine- learning model can be used to assess the performance of the machine-learning model as it is being trained. Training can be terminated if and/or when the target performance is obtained, and/or the maximum number of training iterations have been completed.
  • a regressor head can be appended to at least a portion of the trained classifier and the model comprising the regressor head and the portion of the trained classifier can be trained to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes.
  • a machine learning model comprising a regressor head and at least a portion of the trained classifier (e.g.
  • a trained protein embedding model and a trained peptide embedding model can be trained using training data comprising, for each of a plurality of training peptide- protein complexes: protein sequence data, peptide sequence data, and corresponding ground truth structure data (e.g. experimentally determined crystal structure data) or pairwise distance data and dihedral angle data derived therefrom.
  • the portion of the trained classifier can include a protein embedding model and a peptide embedding model as described herein, trained as part of the classifier model further comprising a classifier head, but excluding the classifier head after such training.
  • the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter (e.g. each pairwise distance and each dihedral angle for which a prediction is made) that can be characterized by, for example, a mean value, a standard deviation, and an uncertainty (or any other set of statistical parameters that characterize a distribution).
  • the trained machine learning model can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution).
  • the set of example peptide-protein complexes may comprise at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, or more than 300 peptide-protein complexes.
  • the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide- protein complex can be presented as a separate training instance during training of the regressor head, thereby providing more training data in cases where the number of peptide-protein crystal structures available is limited.
  • the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can be integrated with, and trained simultaneously with, the trained machine learning model.
  • the training data set can be randomly parsed, shuffled, and/or divided to train the regressor head.
  • a loss function used as part of model training can comprise an error term (e.g., mean squared error or median squared error) and/or an entropy term (e.g., cross entropy or binary cross entropy).
  • a static or non-static learning rate can be used during training of the regressor head.
  • learning rate annealing e.g., using stepwise annealing or cosine annealing
  • validation-data assessment can be used to potentially terminate training early (e.g., upon determining that a performance target has been met).
  • HLA human leukocyte antigen
  • TCRs T-cell receptors
  • MD molecular dynamics
  • AF2 Java, J.M et al. PLOS Computational Biology 14(12), 1006342 (2018)
  • AF-FT AlphaFold-FineTune
  • the present inventors developed a method that uses ML to guide a physics-based engine, such as e.g. Rosetta (Leman et al. Nature Methods 17(7), 665–680 (2020)), to model conformational flexibility of peptides at the pHLA- I interface.
  • the ML component of the method is trained on pHLA sequences and structures and predicts structural descriptors such as distances between the HLA and the peptide residues and the peptide backbone dihedral angles.
  • the predicted structural descriptors are applied as constraints in a peptide coking protocol in the physics engine (here exemplified as FlexPepDock (FPD) protocol in Rosetta; Raveh et al.
  • Proteins Structure, Function, and Bioinformatics 78(9), 2029-2040, (2010)) to sample peptide conformations.
  • the inventors tested the method on a benchmark set of pHLA complexes consisting of 32 different alleles spread across peptide lengths of 8-13 amino acids. They demonstrated that the approach (termed “p-flex”) can recover 78% and 67% of the peptide backbones sampled by the molecular dynamics simulations for peptides of lengths 8-10 and 11-13 amino acids, respectively. To the best of their knowledge, this is the first method to capture the flexibility of peptides up to 13 amino acids in length in a matter of minutes without relying on time- intensive MD simulations.
  • An ML model was designed to predict the distribution over structural descriptors, namely the C ⁇ -C ⁇ distances between HLA and peptide residues and the peptide backbone dihedral angles ( ⁇ and ⁇ ).
  • a classifier was trained on sequence data with pHLA binding labels to learn about peptide and HLA associations.
  • the regressor head from the common embedding shared with the classifier was trained on structural descriptors from pHLA complexes in the Protein Data Bank (PDB).
  • PDB Protein Data Bank
  • the regressor outputs parameters governing the distributions of distances between peptide and HLA residues and joint distributions of ⁇ and ⁇ dihedral angles of the peptide residues.
  • the model takes in the peptide and HLA sequences and outputs the parameters governing the predictive distribution over the structural descriptors.
  • the predicted descriptors were used select a template structure from the PDB that is similar to a given pHLA complex, for initializing the FlexPepDock protocol (FPD). They are also passed into FPD during refinement as HARMONIC and CIRCULAR sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 69 of 108 HARMONIC constraints.
  • Training data – 3D structures The 3D structure data for training and validation of the ML method was curated from PDB.
  • a single pHLA pair was selected by identifying the HLA chain and the corresponding peptide chain by measuring the distance between the second residue of the peptide and 65th residue in the HLA (arbitrarily selected the residues based on manual examination of multiple crystal structures in the PDB).
  • a finalized trimmed snapshot of the pHLA structure was saved in the structure dataset.
  • the typed alleles, together with amino acid sequences of chains in the trimmed structures, were saved in a separate metadata file. Structures of pHLAs found in both unbound and bound (to TCR or co-receptors) conformational states were kept, but the bound partners were removed before training.
  • the final dataset contained a total of 651 noisy (referred to as noisy because they include structures with missing densities, bound and unbound pHLAs) structures which were subdivided into train, validation and test (blind) sets.
  • the PDB contains many more pHLA structures with peptide lengths up to 10 (592 complexes) compared to those with lengths 11-13 (51 complexes) amino acids.
  • the inventors augmented the training set with structures predicted by AF-FT (Motmaen et al. PNAS 120(9), 2216697120 (2023)), an extension of AF2 (Jumper et al.
  • Training data - Labeled pHLA-I sequences The curated data set of pHLA-I sequences with binary binding labels was obtained from Chu., et.al. (Nature Machine Intelligence 4(3):300-311). This data set consists of (i) binder pHLA sequences and (ii) non- binder pairs that were generated for every binder pair by choosing random peptides of same lengths from the peptidomes in IEDB for same alleles.
  • ML method to predict structural descriptors The input to the constraint prediction model consisted of a HLA pseudosequence s HLA with a fixed length of 34 residues and a peptide sequence s pep of variable length ranging from 8 to 15, inclusive.
  • the model was trained to output the predictive distribution over the pairwise distances between the peptide and HLA residues as well as the ⁇ , ⁇ dihedral angles of the peptide residues ( ⁇ ⁇ , ⁇ , ⁇ ⁇ , ⁇ as described herein, denoted P(D, ⁇ , ⁇
  • This is referred to as “joint” modelling of the dihedral angles and is what was used unless indicated otherwise. Indeed, the distribution factorizes among the residues, but ⁇ and ⁇ are modeled jointly for a given residue.
  • the predictive distribution was parametrized as a Gaussian mixture with K components as explained above.
  • the component numbers (K) for the Gaussian mixtures can in principle vary between each predictive distribution but in the present examples the inventors set them to be equal for simplicity and tuned K via cross validation.
  • the final output of the model is the full set of distributional parameters ⁇ ⁇ as described above.
  • the ESM2 Transformer model described in Rives et al. (2021), PNAS 118(15)e2016239118 available at https://github.com/facebookresearch/esm was used to initialize embedding models for both peptide sequences and MHC sequences. Two copies of this model were included to encode respectively a peptide sequence and an MHC pseudosequence.
  • the ESM2 embeddings were mean-aggregated across the residue positions and a small MLP head with a single (logit) output was appended.
  • the resulting classifier model was trained to predict peptide-MHC binding using the data in Chu et al. (2022) and a binary cross-entropy / log-loss function.
  • the binary sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 71 of 108 classification task is designed to provide the model, pretrained on general proteins, with an understanding of valid binding interfaces for a wide variety of peptides and HLAs.
  • the classification head was then removed and replaced with a new regressor MLP.
  • each training instance as a single HLA residue and a single peptide residue.
  • the model extracts the slices corresponding to the residues from the ESM embeddings, concatenate them, and pass the resulting tensor through the regressor MLP.
  • the training is carried out via maximum likelihood estimation (MLE) and the negative log of the distribution as the loss function (L ( ⁇ ) as explained above, where w is a tuned hyperparameter that trades off between the distance and dihedral losses and the subscript ⁇ represents the neural network (MLP) weight parameters).
  • the dataset for the regressor MLP training includes structures with missing distance or angle labels.
  • the inventors accounted for missingness during training so that, when a distance label for D ij was missing for an HLA residue i and peptide residue j, or when both ⁇ j and ⁇ j labels were missing for a peptide residue j, the corresponding model prediction did not contribute to the loss.
  • the HLA was truncated (residues 1-180) to increase the speed of the simulation. Hydrogens were stripped and sidechains pKas were predicted using PROPKA at pH 7.4, and protonation states were assigned using PDB2PQR with the Amber FF19SB force field.
  • the protein was solvated in an octahedral box of OPC water with the box extending 10 ⁇ from the protein and neutralizing NaCl ions.
  • the system is then equilibrated for 1 ns with a restraint force constant of 1 kcal mol-1 ⁇ -2 , and a final equilibration is run at 300K with no restraints. All restraints were removed for the production stage.
  • the production simulations were carried out using NPT conditions and were run using the hydrogen mass repartition option of Amber, allowing a time step of 4 fs. Langevin dynamics are used to maintain the temperature at 300K with a collision frequency of 3 ps-1. Periodic boundary conditions are applied in all directions, the particle mesh Ewald (PME) method is used to calculate the electrostatic interactions with a real-space cutoff of 10 ⁇ , and hydrogen lengths were constrained with the SHAKE algorithm.
  • PME particle mesh Ewald
  • Simulations were run for 300ns in duplicate adding up to 600ns.
  • the inventors analyzed MD simulations to gain insights into the convergence of the simulations and to understand the conformational ensembles sampled within the MD dataset.
  • the analyses are performed using the cpptraj, unless otherwise specified Root mean squared deviation (RMSD) values were calculated for the backbone atoms (CA,C,O,N) using the ’rms’ CPPTRAJ command after superimposition of the truncated HLA.
  • the root mean squared fluctuation (RMSF) of the peptide was calculated using the CPPTRAJ ’atomicfluct command.
  • the dihedral angles were calculated for the peptide using the ’dihedral’ command.
  • Coarse grained models were homology modeled using structural templates selected based on either (i) sequence similarity of the target to the template peptide sequence calculated using BLOSUM62 or (ii) the predicted distance and dihedral constraints. For the latter, the selection was restricted to structural templates in the PDB with the peptide lengths matching the target peptide length without missing densities. For each template, the average Gaussian log-likelihood of the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 73 of 108 distances and dihedrals was computed as described above.
  • Rosetta was used to homology model a target sequence of interest, using as input a target peptide sequence, the HLA allele and a template structure of the complex.
  • the aligning the sequences of the target and the template were aligned using ClustalOmega.
  • the alignments together with template coordinates were utilized by the Rosetta’s threading protocol to generate a homology model.
  • the model was further subjected to a restricted refinement stage where only residues of the peptide and the residues in the HLA contacted by the peptide (within 3.5 ⁇ A) are refined in the Rosetta force field.
  • Structural clustering of peptide backbones To cluster peptide backbones, the matrix of pairwise cosine difference of the dihedral angles of the peptides was calculated as described above. Clustering of the backbones was performed using agglomerative clustering. The cluster centers were then identified and the closest samples to these clusters centers were selected as centroids for downstream tasks.
  • the distance between the two dihedral angles is defined to be: (see, e.g., North et al. (2011), “A New Clustering of Antibody CDR Loop Conformations”, J. Mol Biol. 406(2):228-256).
  • the D-score was used to capture an alignment free distance between two peptide backbones.
  • the D-score was computed for ⁇ , ⁇ and ⁇ dihedral angles of L ⁇ 1 residues for a peptide length L as ⁇ ⁇ ⁇ ( ⁇ , ⁇ ) ⁇ 2(1 ⁇ cos ( ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ )) where A and B are peptides of length L.
  • the peptide dihedral angles of the predicted complexes and those of the MD complexes represent samples drawn from two distributions.
  • the D-score above quantifies the difference between these distributions by treating the ⁇ , ⁇ , and ⁇ angles as independent — the angular distances are computed for each position and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 74 of 108 aggregated in a simple sum. Position-wise metrics like the D-score does not account for the correlations in the three dihedral angles across the residue positions, however.
  • the root mean square deviation (RMSD) between two structures (such as e.g.
  • a predicted structure and a “ground truth”) structure is calculated as ⁇ where n is the number of pairs of equivalent atoms in the two aligned / superimposed structures, di is the distance (typically in ⁇ , typically Euclidian distance) between the two atoms in pair i.
  • the RMSD can be calculated for any type and/or subset of atoms in a protein or protein-peptide complex. Unless indicated otherwise, the RMSD values provided herein refer to RMSDs calculated using all backbone heavy atoms (C, N, C ⁇ , and O) of both a peptide and a protein in a peptide-protein complex for which two structures are compared. The side chain contact recovery was estimated between the predicted models and either MD models or crystal structures.
  • the number of contacts between the HLA and peptide chains of the reference structure were computed by first identifying the residues on the HLAs that are within a distance of 15 ⁇ between their C ⁇ and the peptide residues (or identifying nearby HLA residues to reduce to the search space for subsequent step). Any-atom to any-atom contacts were then captured from the nearby HLA residues and the peptide residues if they are within a threshold of 3.1 ⁇ . The fraction of contacts recovered by the models were then estimated by counting the overlapping contacts between the reference structures and the models of interest. [0209] FIG.
  • 11A provides a non-limiting example of data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted ⁇ and ⁇ dihedral angles relative to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis – identified by the PDB identifier for the corresponding structure) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues), where the ⁇ and ⁇ dihedral angles were treated independently during training and inference. Results are shown as distributions of RMSD values for a plurality of conformer structures predicted as described herein.
  • RMSD root mean square deviation
  • FIG.11B provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted ⁇ and ⁇ dihedral angles compared to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues), where the ⁇ and ⁇ dihedral angles were treated either independently or jointly for a given sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 75 of 108 amino acid residue during training and inference.
  • RMSD root mean square deviation
  • FIG.11C provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted ⁇ and ⁇ dihedral angles compared to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length, where one data set (red single dots) was generated using a prior protein structure prediction model (AlphaFold – FineTune (AF-FT)), and the other data set (green distribution of dots) was generated using the methods disclosed herein (with joint prediction of the dihedral angles).
  • RMSD root mean square deviation
  • Results are shown as distributions of RMSD values for a plurality of conformer structures predicted as described herein. These preliminary results indicate that AF-FT predicts the crystal structures accurately but does not reflect the dynamics of the pMHC complex as it exists in solution.
  • FIG. 11D provides a non-limiting example of preliminary data for predicted average dihedral distance plotted on the y-axis for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues). The results are consistent with the expectation of greater flexibility of longer peptides.
  • FIG.11E provides a non-limiting example of preliminary data for the recovery of crystal structure contacts (i.e., fraction of amino acid residues in the crystal structure that have distances of less than 3.1 ⁇ that also have distances of less than 3.1 ⁇ in the predicted structures) for a set of benchmark pMHC complexes (comprising peptides of length 8, 9, 10, 11, and 13 amino acid residues), where predicted structures are obtained using the disclosed structural prediction methods.
  • the distribution of the fraction of contacts recovered over the predicted conformer structures (y- axis) is plotted for the benchmark pMHC complexes (x-axis).
  • the data demonstrates a high degree of correlation between the predicted contacts and the crystal structure contacts.
  • the inventors partitioned the training and test sets based on sequence and structural similarity.
  • the train and test sets were composed of 80% and 20% of the total 3D structure data from PDB respectively of peptide lengths up to and including 11 AA.
  • the train set was further reshuffled into a 5-fold cross-validation sets consisting of approx. 80% and 20% of the 80% structure data into train and validation sets.
  • the test set and the validation sets consisted of bins of data points covering all the alleles and peptide lengths that had more than one co-crystal structure.
  • peptides of lengths greater than 11 AA the inventors trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 76 of 108 on the data from PDB and structural models generated using Alphafold-Finetune. The selection of validation sets followed the similar strategy as described above for training the models with structure data from PDB only. To generate structure-based data splits, the inventors binned the structures by peptide lengths (up to 10 AA). For each peptide length, they performed clustering of the peptide backbones using a D-score radius of 5.9. All the singleton clusters were moved to the test set.
  • each randomly selected member was placed in a validation set selected at random until a count of 60 in the validation set was reached.
  • the sequence-based split probes the model’s ability to recapitulate the structural descriptors for peptides with sequences unseen in the training set.
  • the structure-based split stress-tests the model’s ability to generalize to structural clusters not represented in the training set, particularly for longer (11-13 amino acids) peptides. Consisting of PDB structures that are diverse in sequence and structure spaces as well as distinct from the training set, these test sets ensure reliable model evaluation.
  • Fig.12 shows exemplary results demonstrating the method’s ability to recover dihedral angles and distances.
  • FIG. 12A shows Boxplots highlighting the average of predicted distances (sampled from the aggregated Gaussian distance distribution) for the peptide positions grouped by peptide length.
  • a box shows a distance distribution constructed from the minimum (min from 34 distances of a peptide residue to the HLA residues) predicted mean distances averaged across all the modeled samples (200 FPD models across 5 folds) for individual peptide position for every test instance.
  • X-axis in panels A and B are peptide positions.
  • Fig. 12A shows that the method produced accurate predictions of distances for both data splits for shorter peptides. The model was able to discriminate between anchor and non-anchor residue positions.
  • 12B shows Ramachandran plots showing dihedral recovery of shorter peptides with length of 8, 9, or 10 (shown in ’Short’), or longer peptides with length of 11, 12, 13 (shown in ’Long’).
  • Predicted dihedral density is shown by light grey contour (the allowed region covering 98% residues) or a dark grey contour (the favored region 98% residues).
  • Scatters represent true dihedral angles sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 77 of 108 in the test set. The points are colored and shaped by whether they fall in the allowed region (blue circle), or outside the allowed region (red cross).
  • Fig.12B show consistent complete overlaps between predicted distributions and native dihedral angles for all instances (including OOD cases) in the test sets, with at least 99.95% of the predicted samples falling within the allowed region.
  • Fig. 12C shows position-wise Ramachandran plots for the representative 9-mer example, 1I1Y.
  • the allowed region (light green) covers 99.95%, while the favored region (dark green) covers 98% dihedral angles.
  • the peptide positions are highlighted using black circles on the stick figure shown in wheat.
  • Fig. 12C shows that the predicted dihedral distributions are tightly clustered along the anchor residues (residue positions 2 and 8), whereas more spread out along the middle residues (residue positions 5 and 6), which reflect a typical scenario for 9-mer peptides in the crystal structures. This spread along the middle residues allows for wider sampling space and thus helps in capturing peptide plasticity. Relative to shorter peptides, the distance distributions were more spread out as expected owing to higher flexibility of longer peptides ( Figure 12A).
  • the model was able to discriminate between anchor and non-anchor positions, found typically along the N- and C-terminal regions of the peptide.
  • multiple stabilizing interactions with the HLA groove, along positions 1-4 amino acids were found ( Figure 12A), which are reflected in the crystal structures.
  • the predicted dihedral angles overlapped with natives for all long peptide test cases (Fig.12B), and span the allowed range of ⁇ angles. Such degeneracy along predicted dihedral angles together with distances enables sampling of diverse peptide backbones.
  • the inventors sampled the structural ensembles of the benchmark test sets from both data splits using FPD with the guidance of ML constraints.
  • the conformational ensembles were obtained using on one or many templates, refined in FPD. Templates for a given target complex were selected (1) using sequence similarity (or sequence homology), (2) using structural similarity estimated from the predicted structural descriptors of the target complex relative to the co-crystal structures in the PDB, and/or (3) clustering the available peptide backbones of the target length from the PDB or MD samples and locating the cluster centers. Typically for shorter peptides (up to 10 amino acids), the inventors found in this example dataset that choosing a single template based on sequence or structural similarity performed very well.
  • the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 78 of 108 inventors believe that this is because peptides in the HLA groove are anchored and stretched, allowing for some plasticity along the middle residues. For longer peptides (i.e.11 amino acids and longer), choosing multiple different templates was beneficial in allowing to sample diverse peptide backbones. This is because these peptides are expected to be highly flexible and therefore using multiple starting points can help to better capture their flexibility. As a baseline, the inventors sampled structural ensembles of the benchmark sets by starting with their homologs chosen based on sequence similarity and refining without any guidance (i.e.
  • the inventors demonstrate the success of the method by its ability to recover i) native, and ii) overlapping and non- overlapping peptide backbones sampled from MD simulations. Since baseline uses homologs as templates, better recovery of native co-crystal structures is generally expected due to available mutant peptides that are few edit distances away and preserve the desired peptide backbones. The inventors evaluated the models from the present method using the D-score and distributional distances as metrics. A peptide backbone was considered recovered if there was at least one FPD sample within a D-score of 10 from any of the 300 MD samples.
  • Recovery rates were 99.9%, 83.9% and 69% for 8-, 9- and 10-mers, respectively, from the sequence- based split; 100%, 69.7%, and43.8% for 8-, 9- and 10-mers, respectively, from the structure- based split; and 71.7%, 47.5% and 82.6% for all 11-, 12- and 13-mers in the PDB.
  • the template selection strategies were constraint-based for 8- and 9-mers, sequence-based for 10- mers, and multi-template for 11-, 12- and 13-mers, followed by constrained refinement using the predictions from the machine learning model.
  • Fig.13 shows non-limiting examples of data demonstrating that structural models sampled by the present method are diverse and overlap with MD ensembles on the benchmark sets across all the peptide lengths.
  • Fig. 13 is a plot showing the number of overlapped FPD models with MD samples (y-axis; using a D-score cut-of of 10.0) for all the PDBIDs in the benchmark set (x-axis) across peptide lengths 8 (red), 9 (magenta), 10 (blue), 11 (orange), 12 (red) and 13 (brown) (from left to right, shown by the color bar below x-axis).
  • the yellow and gray bars (left and right bars in each pair for a benchmark complex) represent constrained and unconstrained refinement respectively.
  • the constrained refinement strategy retrieved MD backbones better than the baseline strategy for shorter length peptides.
  • constrained refinement recovers a higher number of MD backbones, especially for 19% of the cases while the baseline is better or comparable for 1% or 80% of the.
  • the FPD samples that do sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 79 of 108 not have backbones overlapping with MD are not necessarily incorrect, they are potentially reasonable backbones if they have lower Rosetta energies indicating overall stability of the complex together with binding scores.
  • Fig. 14A-B provide non-limiting examples of data demonstrating that the conformational ensembles sampled by the methods of the present disclosure approximate benchmark MD distributions.
  • Fig.14A-B show distributions of evaluation metrics of the FPD models sampled using constrained and unconstrained refinement relative to the native crystal structures for all the peptide lengths.
  • the plots show: (Top) Peptide backbone heavy-atom (N, C ⁇ , C, and O) RMSD (in ⁇ ); (Middle) D-score values (0 indicated perfect recovery of the peptide backbones); (Bottom) Side-chain contact recovery (0 indicates no recovery whereas 1 indicates perfect side chain placement relative to the natives) of the FPD models relative to the native crystal structures for the benchmark PDBIDs (along x-axis). Data is shown on Fig.14A for sequence-base splits, and on Fig.14B for structure-based splits. for 8-10 AA and all test PDBs for 11-13 AA.
  • the peptide lengths 8 (red), 9 (magenta), 10 (blue), 11 (orange), 12 (red) and 13 (brown) are shown by the color bar (from left to right in the order as indicated above) below x-axis.
  • the baseline approach is expected to sample natives better, however the data showed that constraints during refinement guided sampling of models closer to the native structures.
  • Fig.14A shows this result for 25% of 8-mers (4F7T), 30% of 9-mers (1QRN, 2C7U, 2VLL, 2X4O, 3BVN, 3D18, 3I6G, 3I6L, 3KPM, 4MNQ, 4N8V, 4O2E, 4QRS, 5ISZ, 5MER, 5TEZ, 5VVP, 5WMO, 6RPA, 6Z9W, 7KGP, 7KGS, 7LGD, 7NME, 7R80, 7ZUC), and 25% of 10-mers (1HHH, 5C09, 6MPP, 7S8Q, 7S8S, 8DVG) in test sets from sequence-based splits as highlighted by either peptide backbone heavy-atom RMSD, D-score or side-chain contact recovery ( Figure 14A).
  • FIGS.16A-B provide non-limiting examples of preliminary control data for pMHC template structure selection and refinement, respectively.
  • FIG. 16A provides a plot of the median simulated RMSD of pairwise distances between peptide backbone heavy atoms C, N, C ⁇ , and O in the selected starting template structure (for peptides of 8, 9, 10, 11, and 13 amino acid residues in length) versus the corresponding RMSD for pairwise distances obtained from crystal structure data.
  • FIG. 16A provides a plot of the median simulated RMSD of pairwise distances between peptide backbone heavy atoms C, N, C ⁇ , and O in the selected starting template structure (for peptides of 8, 9, 10, 11, and 13 amino acid residues in length) versus the corresponding RMSD for pairwise distances obtained from crystal structure data.
  • FIG. 16A provides a plot of the median simulated RMSD of pairwise distances between peptide backbone heavy atoms C, N, C ⁇ , and O in the
  • FIGS. 17A-C provide non-limiting examples of molecular dynamics data (300ns simulation) for a sampling of peptide conformations (plot of ⁇ (x-axis) and ⁇ (y-axis) dihedral angles) within a pMHC complex (FIG.
  • FIG. 17A provides a plot of the minimum and maximum values of RMSF obtained for a set of several hundred 9-mer peptides versus peptide residue number. The data indicates the variation in the degree of peptide flexibility within the pMHC complex as a function of peptide residue position.
  • FIGS. 18A-C provide non-limiting examples of molecular dynamics data (300ns simulation) for a sampling of peptide conformations (plot of ⁇ (x-axis) and ⁇ (y-axis) dihedral angles) within a pMHC complex (FIG.
  • FIG.18C provides a plot of RMSF for 11 different 13-mer sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 81 of 108 peptides versus peptide residue number.
  • Fig.15A-B show non-limiting examples demonstrating that peptide conformations sampled using guidance from predicted constraints as described herein can illuminate states that dictate TCR specificity.
  • the inventors used the HLA-A03:01 molecule in complex with a mutant PIK3CA peptide as a use case.
  • the authors discuss a large conformational change of the side chain of the tryptophan at position 6 (pTrp6) that happens upon binding.
  • Fig.15A is Structural representation of HLA-A03:01 in various states.
  • the pink cartoon (PDB ID: 7L1C) on Fig.15A represents the unbound HLA-A03:01 in complex with a mutant PIK3CA peptide.
  • the purple cartoon (PDB ID: 7L1D) on Fig.15B depicts the TCR bound to the HLA-A03:01 in complex with a mutant PIK3CA peptide.
  • the grey cartoons on Fig.15A illustrate the modeled HLA- A03:01 in complex with a mutant PIK3CA peptide, s featuring multiple conformations, which highlight the model’s ability to sample both the bound and unbound conformations observed in the crystal structures.
  • the pink cartoon (PDB ID: 7L1C) on Fig.15B represents the unbound HLA-A03:01 in complex with a mutant PIK3CA peptide.
  • the purple cartoon (PDB ID: 7L1D) on Fig.15B depicts the TCR bound to the HLA-A03:01 sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 82 of 108 in complex with a mutant PIK3CA peptide.
  • the grey cartoons on Fig. 15B illustrate the modeled HLA-A03:01 in complex with a mutant PIK3CA peptide, show casing the highest RMSD states from the MD 7L1C simulations.
  • the alignment- free D-score metric highlight the overlaps between MD and the predicted conformer ensembles.
  • the ML model uses ESM2 sequence embeddings shared across a classifier (to learn amino acid associations between peptide and HLA) and a regressor (to predict the structural constraints).
  • FIG.19 provides a non-limiting example of a block diagram for a computer system, in accordance with some embodiments.
  • Computer system 1900 can be an example of one implementation for Computing Platform 102 described above in FIG.1.
  • Computer system 1900 can be a host computer connected to a network.
  • Computer system 1900 can be a client computer or a server.
  • computer system 1900 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a phone sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 83 of 108 or tablet.
  • the device can include, for example, one or more of processor 1910, input device 1920, output device 1930, storage 1940, and communication device 1960.
  • Input device 1920 and output device 1930 can generally correspond to those described elsewhere herein, and they can either be connectable or integrated with the computer.
  • Input device 1920 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device.
  • Output device 1930 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker.
  • Storage 1940 can be any suitable device that provides storage, such as an electrical, magnetic, or optical memory including a RAM, cache, hard drive, or removable storage disk.
  • Communication device 1960 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus 1970 or wirelessly.
  • Software 1950 which can be stored in memory / storage 1940 and executed by processor 1910, can include, for example, the programming that embodies the functionality of the present disclosure (e.g., as embodied in the methods described above).
  • Software 1950 can also be stored and/or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions.
  • a computer-readable storage medium can be any medium, such as storage 1940, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device.
  • Software 1950 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions.
  • a transport medium can be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device.
  • the transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation medium.
  • Computer system 1900 may be connected to a network, which can be any suitable type of interconnected communication system.
  • the network can implement any suitable communications protocol and can be secured by any suitable security protocol.
  • the network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, Tl or T3 lines, cable networks, DSL, or telephone lines.
  • Computer system 1900 can implement any operating system suitable for operating on the network.
  • Software 1950 can be written in any suitable programming language, such as C, C++, Java, or Python.
  • FIG. 20 illustrates a diagram 2000 of an example artificial intelligence (AI) architecture 2002 (which can be included as part of the one or more computing device(s) 1900 as discussed above with respect to FIG. 19) that can be utilized to determined one or more predicted peptide-protein binding interactions and/or to generate a plurality of compatible structures for a peptide-protein complex, in accordance with the disclosed embodiments.
  • AI artificial intelligence
  • the AI architecture 2002 can be implemented utilizing, for example, one or more processing devices that may include hardware (e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC), a system-on- chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and/or other processing device(s) that can be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions running/executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof.
  • hardware e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC),
  • the AI architecture 2002 may include machine learning (ML) algorithms and functions 2004, natural language processing (NLP) algorithms and functions 2006, expert systems 2008, computer-based vision algorithms and functions 2010, speech recognition algorithms and functions 2012, planning algorithms and functions 2014, and robotics algorithms and functions 2016.
  • ML machine learning
  • NLP natural language processing
  • the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 85 of 108 ML algorithms and functions 2004 may include any statistics-based algorithms that can be suitable for finding patterns across large amounts of data (e.g., “Big Data” such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data).
  • the ML algorithms and functions 2004 may include deep learning algorithms 2018, supervised learning algorithms 2020, and unsupervised learning algorithms 2022.
  • the deep learning algorithms 2018 may include any artificial neural networks (ANNs) that can be utilized to learn deep levels of representations and abstractions from large amounts of data.
  • ANNs artificial neural networks
  • the deep learning algorithms 2018 may include ANNs, such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth.
  • ANNs such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (
  • the supervised learning algorithms 2020 may include any algorithms that can be utilized to apply, for example, what has been learned in the past to new data using labeled examples for predicting future events. For example, starting from the analysis of a known training data set, the supervised learning algorithms 2020 may produce an inferred function to make predictions about the output values. The supervised learning algorithms 2020 may also compare its output with the correct and intended output and find errors in order to modify the supervised learning algorithms 2020 accordingly.
  • the unsupervised learning algorithms 2022 may include any algorithms that may applied, for example, when the data used to train the unsupervised learning algorithms 2022 are neither classified nor labeled.
  • the unsupervised learning algorithms 2022 may study and analyze how systems may infer a function to describe a hidden structure from unlabeled data.
  • the NLP algorithms and functions 2006 may include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and/or text.
  • the NLP algorithms and functions 2006 may include content extraction algorithms or functions 2024, classification algorithms or functions 2026, machine translation algorithms or functions 2028, question answering (QA) algorithms or sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 86 of 108 functions 2030, and text generation algorithms or functions 2032.
  • the content extraction algorithms or functions 2024 may include a means for extracting text or images from electronic documents (e.g., webpages, text editor documents, and so forth) to be utilized, for example, in other applications.
  • the classification algorithms or functions 2026 may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, na ⁇ ve Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon.
  • a supervised learning model e.g., logistic regression, na ⁇ ve Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth
  • the machine translation algorithms or functions 2028 may include any algorithms or functions that can be suitable for automatically converting source text in one language, for example, into text in another language.
  • the QA algorithms or functions 2030 may include any algorithms or functions that can be suitable for automatically answering questions posed by humans in, for example, a natural language, such as that performed by voice-controlled personal assistant devices.
  • the text generation algorithms or functions 2032 may include any algorithms or functions that can be suitable for automatically generating natural language texts.
  • the expert systems 2008 may include any algorithms or functions that can be suitable for simulating the judgment and behavior of a human or an organization that has expert knowledge and experience in a particular field (e.g., stock trading, medicine, sports statistics, and so forth).
  • the computer-based vision algorithms and functions 2010 may include any algorithms or functions that can be suitable for automatically extracting information from images (e.g., photo images, video images).
  • the computer-based vision algorithms and functions 2010 may include image recognition algorithms 2034 and machine vision algorithms 2036.
  • the image recognition algorithms 2034 may include any algorithms that can be suitable for automatically identifying and/or classifying objects, places, people, and so forth that can be included in, for example, one or more image frames or other displayed data.
  • the machine vision algorithms 2036 may include any algorithms that can be suitable for allowing computers to “see”, or, for example, to rely on image sensors cameras with specialized optics to acquire images for processing, analyzing, and/or measuring various data characteristics for decision making purposes.
  • any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend.
  • references in the appended claims to an apparatus or system or a component of an apparatus or system being sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 88 of 108 adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative.
  • this disclosure describes or illustrates certain embodiments as providing particular advantages, certain embodiments may provide none, some, or all of these advantages.
  • peptide As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably to refer to a polymer of amino acid residues. The terms encompass amino acid chains of any length, including full-length proteins with amino acid residues linked by covalent peptide bonds.
  • a “mutant peptide” or “mutated peptide” can refer to a peptide that is not present in the normal tissue (e.g., in the wild type amino acid sequences of normal tissue) of an individual subject.
  • a mutated peptide comprises at least one mutated amino acid and can be present in a diseased tissue (e.g., collected from a particular subject), but not in a normal tissue (e.g., collected from the particular subject, collected from a different subject, and/or as identified in a database as corresponding to normal tissue).
  • a mutated peptide can include an epitope.
  • An epitope is the portion of a mutated peptide to which an MHC molecule or a TCR binds.
  • the binding between the epitope of the mutated peptide and the MHC molecule or TCR can induce an immune response (as a result of the mutated peptide not being associated with a subject’s “self”).
  • a mutated peptide can include or be a neoantigen.
  • a mutated peptide can arise from, as non-limiting examples: a non-synonymous mutation leading to different amino acids in the protein (e.g., point mutation); a read-through mutation in which a stop codon is modified or deleted, leading to translation of a longer protein with a novel tumor-specific sequence at the C-terminus; a splice site mutation that leads to a unique tumor-specific protein sequence; a chromosomal rearrangement that gives rise to a chimeric protein with a tumor- specific sequence at a junction of two proteins (gene fusion); and/or a frameshift insertion or deletion that leads to a new open reading frame with a tumor-specific protein sequence.
  • a mutated peptide can include a polypeptide (as characterized by a polypeptide sequence) and/or can be encoded by a nucleotide sequence.
  • a “C-flank” of a peptide refers to one or more amino acids upstream of the C-terminus of the peptide, from the parent protein.
  • a C-flank of a peptide includes one, two, three, four, five, or more amino acid residues upstream of the C-terminus of the peptide.
  • an “N-flank” of a peptide refers to one or more amino acids downstream of the N-terminus of the peptide, from the parent protein.
  • an N-flank of a peptide includes one, two, three, four, five, or more amino acid residues downstream of the N-terminus of the peptide.
  • an “epitope” of a peptide can refer to a region of the peptide between the C-flank and N-flank and can be recognized by a TCR.
  • the epitope of the peptide is a part of the peptide that is recognized by a TCR on a T cell and MHC I on an antigen-presenting cell.
  • the epitope can be a peptide to which a TCR binds, such as a peptide to which the TCR binds when the peptide is bound to MHC I on an antigen-presenting cell.
  • a “ligand” is a peptide that is found to be presented by an MHC molecule at the cell surface from elution experiments or found to be bound to MHC in an in vitro assay.
  • a “sequence” refers to an amino acid sequence that includes an ordered set of amino acid identifiers.
  • a “peptide sequence” refers to a sequence that identifies amino acids of at least a portion of a peptide. In some cases, the peptide sequence includes a variant-coding sequence that includes a variant that is not observed in a corresponding reference sequence. [0247] When the peptide includes a mutated peptide, the variant-coding sequence, identifies amino acids of the mutation or variant.
  • the variant-coding sequence does not identify amino acids of a mutation or variant (and in that instance, is the same as the reference sequence).
  • a variant-coding sequence can be determined by collecting a disease and/or tumor sample (e.g., that includes tumor cells) and performing a sequencing analysis to identify one or more sequences corresponding to disease and/or tumor cells in the sample.
  • a sequencing analysis outputs an amino acid sequence.
  • a sequencing analysis outputs a nucleic acid sequence, which can be subsequently processed to transform codons into amino acid identifiers and thus to produce an amino acid sequence.
  • a variant-coding sequence can sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 90 of 108 include a sequence of a neoantigen.
  • a variant-coding sequence can, but need not, include one or more termini (e.g., the C-terminus and/or the N-terminus) of the peptide.
  • a variant-coding sequence can include an epitope of the peptide.
  • a variant-coding sequence can identify amino acids within a peptide having one or more variants (e.g., one or more amino acid distinctions) relative to a corresponding reference sequence. In some instances, a variant-coding sequence includes an ordered set of amino acids.
  • a variant-coding sequence identifies a reference peptide (e.g., by identifying a genetic reference sequence, such as by gene, start position, and/or end position; or by gene, start position, and/or length) and one or more point mutations relative to the reference peptide.
  • a “reference sequence” can refer to a sequence that identifies amino acids within at least part of a non-mutated peptide or wild-type peptide (e.g., wild-type, parental sequence).
  • the non-mutated or wild-type peptide can include no variants or fewer variants than are included in a mutated peptide.
  • the reference sequence can include an amino acid sequence encoded by a genetic sequence within a same gene relative to a gene that includes a corresponding variant-coding sequence.
  • the reference sequence can include an amino acid sequence encoded by a genetic sequence spanning the same start and stop within a gene relative to intra-gene positions associated with a genetic sequence associated with a corresponding variant-coding sequence.
  • the reference sequence can be identified by collecting a non-disease and/or non-tumor sample from one or more subjects (who can, but need not, include a subject from which a disease sample was collected to determine a variant-coding sequence) and performing a sequencing analysis using the sample.
  • a “pseudosequence” of an MHC molecule can refer to an ordered set of amino acids of the MHC molecule that typically contacts a peptide.
  • a pseudosequence can be a subset of the amino acid sequence of a protein that is expected to contact a peptide in a peptide-protein complex.
  • the subset can comprise consecutive or non- consecutive amino acids, but the amino acids in a pseudosequence are in the same order as in the full sequence of the protein.
  • a pseudosequence of a protein can be defined as an order set of amino acids in the sequence of the protein that are within a predetermined distance of a peptide amino acid in a structure of the peptide-protein complex.
  • the structure can be experimentally determined or predicted.
  • the amino acids included in a pseudo-sequence can additionally be filtered to only include polymorphic positions.
  • Pseudosequences for many MHC molecules are known in the art (see e.g. Nielsen et al. NetMHCpan, a Method for sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 91 of 108 Quantitative Predictions of Peptide Binding to Any HLA-A and -B Locus Protein of Known Sequence, PLoS One.2007 Aug 29;2(8):e796).
  • a “representation” of a sequence or “sequence representation” can include a set of values that represent or identify amino acids in the sequence and/or a set of values that represent or identify nucleic acids that encode the sequence.
  • each amino acid can be represented by a binary string and/or vector of values that is distinct from each other binary string and/or vector representing each other amino acid.
  • the sequence representation can be generated using, for example, one-hot encoding or using a BLOcks SUbstitution Matrix (BLOSUM) matrix.
  • BLOSUM BLOcks SUbstitution Matrix
  • a multi-dimensional (e.g., 20- or 21- dimensional) array be initialized (e.g., randomly or pseudo-randomly initialized).
  • the initialized array can include, for each amino acid, a unique vector corresponding to that amino acid.
  • the values can be fixed such that use of such a unique vector can be assumed to represent the corresponding amino acid.
  • “presentation” of a peptide refers to at least part of the peptide being presented on a surface of a cell by virtue of being bound to an MHC molecule in a particular manner. The presented peptide can then be accessible to other cells, such as nearby T cells.
  • a “sa “sa “sa “sa “sa sample” can include tissue (e.g., a biopsy), single cell, multiple cells, fragments of cells, or an aliquot of body fluid.
  • the sample can be obtained from a subject by means such as, for example, without limitation, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, intervention, another type of sample collection means, or a combination thereof.
  • a “subject” encompasses one or more cells, tissue, or an organism.
  • the subject can be a human or non-human, whether in vivo, ex vivo, or in vitro, male or female.
  • binding affinity refers to affinity of binding between an amino acid (e.g., a peptide of a specific antigen) and an IPC (e.g., an MHC molecule and/or MHC allele).
  • the binding affinity can characterize a stability, tendency, and/or strength of the binding between the peptide and an IPC.
  • immunogenicity can refer to the ability to elicit an immune response (e.g., via T cells and/or B cells).
  • a peptide that is “immunogenic” can be one that is capable of eliciting an immune response.
  • MHC refers to the major histocompatibility complex.
  • the human MHC is also called the human leukocyte antigen (HLA) complex.
  • HLA human leukocyte antigen
  • Embodiments disclosed herein can include: 1.
  • a computer-implemented method for generating a plurality of compatible structures for a peptide-protein complex comprising: inputting peptide sequence data for at least one peptide and protein sequence data for at least one protein into a trained machine learning model to determine: predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance data and the predicted dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data.
  • each compatible structure for the peptide-protein complex conforms with the sequence data for the at least one peptide and the at least one protein and with a prediction of at least one auxiliary feature.
  • the at least one auxiliary feature comprises a prediction of binding of the at least one peptide to the at least one protein.
  • the at least one auxiliary feature comprises a prediction of at least one interface feature for the peptide-protein complex. 7.
  • the protein sequence data comprises a protein embedded representation for the at least one protein generated using a trained protein embedding machine learning model.
  • the trained peptide embedding machine learning model and the trained protein embedding machine learning model are the same model.
  • the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise foundational models.
  • the trained machine learning model comprises a trained neural network.
  • the trained machine learning model comprises a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network.
  • MLP multilayer perceptron
  • the trained machine learning model is trained using a training data set that includes incomplete structural data for one or more peptide-protein complexes. 17.
  • the machine learning model is trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide-protein combinations; training a classifier to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 95 of 108 appending a regressor head to at least a portion of the trained classifier and training the regressor head to predict probabilistic distributions of pairwise distance data and dihe
  • the augmented pairwise distance data and dihedral angle data is derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor.
  • identifying the initial structure for the peptide-protein complex based on the predicted pairwise distance data and predicted dihedral angle data comprises: obtaining a plurality of crystal structures for peptide-protein complexes from a protein structure database; and selecting a crystal structure from the plurality of crystal structures based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 96 of 108 a maximum likelihood estimate for the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model.
  • PDB Protein Data Bank
  • generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises: inputting the initial structure for the peptide-protein complex into a structural modeling software package; inputting the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, wherein compatible structures for the peptide- protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function.
  • the structure scoring function comprises a proxy for a free energy calculation of peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof.
  • DFT density functional theory
  • the structural modeling software package comprises FlexPepDock.
  • MHC Major Histocompatibility Complex
  • the computer-implemented method of embodiment 37 or embodiment 38 further comprising iterating over a plurality of candidate neoantigen sequences, or portions thereof, expressed in tumor cells from the cancer patient for at least one MHC class I or class II protein to rank order the neoantigen sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex.
  • pMHC-I peptide-MHC class I
  • pMHC-II peptide-MHC class II
  • the computer-implemented method of embodiment 37 or embodiment 38 further comprising iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in a virus or bacterium for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide- MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex.
  • pMHC-I peptide- MHC class I
  • pMHC-II peptide-MHC class II
  • a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the computer-implemented method of any one of embodiments 1 to 43.
  • a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform the computer-implemented method of any one of embodiments 1 to 43.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Medical Informatics (AREA)
  • Data Mining & Analysis (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Evolutionary Biology (AREA)
  • Theoretical Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biophysics (AREA)
  • Artificial Intelligence (AREA)
  • Public Health (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Bioethics (AREA)
  • Epidemiology (AREA)
  • Software Systems (AREA)
  • Databases & Information Systems (AREA)
  • Chemical & Material Sciences (AREA)
  • Medicinal Chemistry (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Investigating Or Analysing Biological Materials (AREA)

Abstract

The present disclosure relates to methods for generating a plurality of compatible structures for a peptide-protein complex that can be used, for example, to identify surface fingerprint and interface features of the peptide-protein complex. The methods can comprise inputting peptide sequence data and protein sequence data into a trained machine learning model to determine: (i) predicted pairwise distance data for peptide amino acid residues and protein amino acid residue in the peptide-protein complex; and (ii) predicted dihedral angle data for peptide amino acid residues in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance and dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data.

Description

ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 1 of 108 MACHINE LEARNING-BASED METHODS FOR MODELING pMHC CONFORMERS CROSS-REFERENCE TO RELATED APPLICATIONS [0001] This application claims the priority benefit of United States Provisional Patent Application Serial No. 63/567,360, filed March 19, 2024, the contents of which are incorporated herein by reference in their entirety. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING [0002] The content of the electronic sequence listing (146392067640seqlist.xml; Size: 2,756 bytes; and Date of Creation: March 3, 2025) is herein incorporated by reference in its entirety. FIELD [0003] This disclosure relates generally to machine learning-based methods for modeling conformational states of peptide-protein complexes that account for the dynamic nature of three-dimensional peptide and protein structures, and that may be used to predict features of the peptide-protein complexes. In some aspects, the present disclosure relates more specifically to the application of the disclosed methods to generate conformer ensembles for peptide-Major Histocompatibility Complex (pMHC) complexes that may be used in downstream predictions of peptide binding and/or other features of peptide-MHC interactions. BACKGROUND [0004] Binding interactions that drive the formation of peptide-protein complexes underlie a variety of biochemical processes of critical importance to living organisms. Examples include, but are not limited to, receptor-ligand binding-dependent cell signaling pathways and enzyme-substrate binding-dependent catalytic reactions. Peptide-protein binding interactions also underlie the adaptive immune system in humans and other vertebrates. Human leukocyte antigens (HLAs) are expressed as cell surface receptor proteins that bind and present antigenic peptides to T cells in a restricted manner, which allows discrimination between self- and foreign antigens. The HLA complex is a complex of genes on chromosome 6 (Chr 6) that encodes the cell-surface proteins responsible for regulation of the immune system. The human major histocompatibility complex (MHC) is a linked set of genetic loci comprising the HLA sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 2 of 108 genes that play a fundamental role in, e.g., the cell-mediated immune response to infectious disease, the acceptance of transplanted tissues, etc. The MHC complex encodes the α-chains of the MHC class I molecules HLA-A, HLA-B, and HLA-C (alleles) and the α- and β-chains of the MHC class II molecules HLA-DR, HLA-DP, and HLA-DQ (allotypes), all of which are expressed in a co-dominant fashion. [0005] Cytotoxic T lymphocytes (cytotoxic T cells) are activated upon recognition of peptides bound to the MHC class I molecules (see, e.g., Ito et al. (2010), “Regulation of the Induction and Function of Cytotoxic T Lymphocytes by Natural Killer T Cell”, Journal of Biomedicine and Biotechnology 2010:641757). Activated cytotoxic T cells secrete essential cytolytic mediators (e.g., perforin, granzyme, etc.) and induce apoptosis in target cells (tumor cells, viral infected cells, etc.). In addition, activated cytotoxic T cells secrete cytokines such as interferon gamma (IFN-γ) and tumor necrosis factor alpha (TNF-α), which enhance antigen presentation and mediate antipathogenic effects. Studies have shown that various cytokines such as interleukin-(IL2) or IFN-γ-producing CD4 T-cells are required for the generation of effective cytotoxic T cell-based immunity to infection and cancer. [0006] Activation of helper T (Th) cells upon recognition of peptide fragments (e.g., epitopes derived from protein antigens) which are bound to MHC class II (MHC-II) molecules of antigen-presenting cells results in the development of antigen-reactive antibodies. Peptide binding to an MHC molecule at a sufficient affinity is a prerequisite for immunogenicity (i.e., the ability of a peptide to trigger an immune response). [0007] The ability to accurately predict binding of peptides comprising candidate MHC-I binding epitopes (typically 8 – 15 amino acid residues in length) and candidate MHC-II binding epitopes (typically 13 – 25 amino acid residues in length) to MHC-I / MHC-II molecules, and the immunogenicity (i.e. their ability to be recognized by a T cell receptor and trigger an immune response) of such peptide-MHC complexes would be valuable for, e.g., identifying the amino acid residues of a protein that could be safely altered to eliminate (or minimize) the immunogenicity of candidate therapeutic proteins and/or to boost the immunogenicity of candidate peptides for inclusion in vaccines, including candidate peptides derived from patient- specific neoantigens for the development of personalized cancer vaccines. Peptide-MHC binding affinity is primarily determined by the amino acid sequence of the peptide binding core (also referred to as “minimal epitope”, referring to the part of a peptide extending between the “anchor residues” of the peptide that fit into pockets of the binding groove of the MHC sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 3 of 108 molecule in the peptide-MHC complex, typically having a median length of about nine amino acid residues); however, the amino-terminal flanking (N-flank) sequence and/or the carboxy- terminal flanking (C-flank) sequence (i.e. one or more residues on the N-terminal and/or C- terminal side of the binding core) can also affect peptide-MHC binding. [0008] There are tools for predicting the ability of peptides to bind to proteins to form complexes (e.g., that predict peptide binding to MHC molecules to form pMHC complexes). However, existing tools often fail to accurately predict properties of the peptide-protein complex (such as the immunogenicity of a specific pMHC). For example, peptides predicted to bind to an MHC molecule often turn out to be false positives in that the resulting pMHC complex does not effectively prime a cellular immune response (i.e. peptides that are predicted to bind to MHC molecules to form a pMHC complex are not in fact immunogenic). Thus, there remains a need for tools that more accurately predict the properties of corresponding peptide- protein complexes. SUMMARY [0009] While there are tools for predicting the ability of peptides to bind to proteins to form complexes, the present inventors postulated that the limitations of these existing tools (i.e. their frequent failure to accurately predict properties of the peptide-protein complex, such as the immunogenicity of a specific pMHC) is at least in part due to their failing to account for the dynamic properties and range of allowed conformational states of peptide-protein complexes. The present inventors postulated that tools that account for the dynamic nature of three- dimensional peptide and protein structures would more accurately predict the properties of corresponding peptide-protein complexes. [0010] Disclosed herein are machine learning (ML)-based methods for modeling conformational states of peptide-protein complexes that account for the dynamic nature of three-dimensional peptide and protein structures. The disclosed methods can be used to generate an ensemble (i.e., a group) of allowed peptide-protein conformers. In some examples, the allowed peptide-protein conformers are allowed three-dimensional structures for a peptide- protein complex that can exist at any given point in time, and that impact the function of the peptide-protein complex at any given point in time. The ensemble of allowed peptide-protein conformers represent an allowed distribution of structural parameters for the peptide-protein complex and can be used to more accurately predict features and functional properties of the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 4 of 108 peptide-protein complex (e.g., complementary interface features, complementary surface fingerprint features, and/or binding of the peptide-protein complex to other proteins or protein complexes). The improvements in prediction accuracy enabled by the disclosed methods and systems are based on the use of embedded representations of peptide and protein sequences as input, machine-learning based prediction of structural constraints for the peptide-protein complex (e.g., predicted posterior distributions for pairwise distance data for amino acid residues in the peptide and the protein, and predicted posterior distributions for dihedral angle data for the peptide in the peptide-protein complex), and the use of the predicted structural constraints to select an initial template structure for the peptide-protein complex and generate a plurality of compatible structures that captures the dynamic behavior of the peptide-protein complex in solution. As noted above, prior methods have failed to account for the dynamic properties and range of allowed conformational states of peptide-protein complexes in solution. The present inventors have shown through extensive benchmarking against ensembles from molecular dynamics (MD) simulations that the approach is able to recover a majority of the backbones sampled by the molecular dynamics simulations for peptide-HLA complexes involving peptide lengths up to 13 amino acids. The inventors further demonstrated that the approach was able to capture the dynamic nature of the wild-type and mutating peptide-HLA complexes recognized by TCRs. To the best of the inventor’s knowledge, the methods described herein are the first methods able to capture the flexibility of peptides (even those up to 13 amino acids in lengths) in a matter of minutes without relying on time-intensive MD simulations. [0011] In some aspects, for example, the disclosed machine learning-based methods can be used to generate conformer ensembles for pMHC complexes that represent an allowed distribution of structural parameters for a given pMHC complex, and that can be used in downstream predictions of the features and functional properties of the pMHC complex (e.g., predictions of complementary interface features, complementary surface fingerprint features, and/or binding of the pMHC complex to other proteins or protein complexes. More accurate predictions of pMHC features and functional properties can, for example, facilitate the identification of epitopes (e.g., peptide fragments derived from neoantigens associated with a tumor) that, when presented on the surface of antigen-presenting cells, provoke a robust immune response. Improved representation of the dynamic nature of peptide-HLA complexes also improves the ability to capture conformations that can be present when the peptide-MHC complex is bound by a TCR. Dynamic allostery is believed to play an important part in TCR sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 5 of 108 recognition of peptide-HLA complexes and correctly capturing this is important to be able to accurately predict TCR cross-reactivity and specificity. The disclosed methods can thus aid in selecting peptide fragments of neoantigens for inclusion in or development of therapeutic cancer vaccines. In some instances, neoantigens associated with a tumor can be patient- specific, and the disclosed methods can aid in selecting peptide fragments of the patient- specific neoantigens for inclusion in or development of personalized cancer vaccines. The disclosed methods can also aid in designing and/or characterizing any TCR-based therapeutic or other tumor antigen recognizing therapy. For example, the disclosed methods can aid in design and/or characterizing of Bi-specific T-cell engagers (BiTE), T cell receptor-engineered T cell (TCR-T) therapy, chimeric antigen receptor (CAR) T cell therapy (CAR-T), etc. [0012] In other aspects, also disclosed herein are systems configured to perform any of the methods disclosed herein, and computer-readable storage media comprising one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform any of the methods disclosed herein. The systems may comprise one or more processors and one or more computer-readable storage media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the methods disclosed herein. [0013] Disclosed herein are computer-implemented methods for generating a plurality of compatible structures for a peptide-protein complex, the methods comprising: inputting peptide sequence data for at least one peptide and protein sequence data for at least one protein into a trained machine learning model to determine: predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance data and the predicted dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data. A plurality of compatible structures may also be referred to as “a conformer ensemble”. An initial structure may also be referred to as “template structure”. The trained machine learning model may be a machine learning model that has been trained to take as input peptide sequence data for at least one peptide and protein sequence data for at least one protein, and produce as output predicted sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 6 of 108 pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex. The machine learning model may have been trained using structural data for a plurality of training peptide-protein complexes. The structural data may comprise, for each of the plurality of training peptide-protein complexes, values of: pairwise distances between at least one peptide amino acid residue of the peptide and at least one protein amino acid residue of protein in the training peptide-protein complex, and one or more dihedral angles for at least one peptide amino acid residue in the peptide in the training peptide-protein complex. The values may be experimentally determined and/or predicted using a structural prediction algorithm. [0014] In some embodiments, the computer-implemented method further comprises predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex. [0015] In some embodiments, each compatible structure for the peptide-protein complex conforms with one or more posterior distributions of the predicted pairwise distance data and with one or more posterior distributions of the predicted dihedral angle data output by the trained machine learning model. The predicted pairwise distance data may comprise, for each of one or more pairs of at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein, a predicted posterior distribution of pairwise distances between said peptide amino acid residue and said protein amino acid residue in the peptide-protein complex. The predicted dihedral angle data may comprise, for each of one or more peptide amino acid residues in the at least one peptide, a predicted posterior distribution of one or more dihedral angles of the peptide amino acid residue in the peptide-protein complex. Predicted data comprising a predicted posterior distribution may refer to the predicted data comprising predicted values of parameters that characterize said distribution. The posterior distributions may be used as constraints in a peptide docking protocol, i.e. in a physics engine for modeling peptide-protein complexes. [0016] In some embodiments, the trained machine learning model has been trained at least in part to predict at least one auxiliary features of input protein-peptide complexes. In some embodiments, each compatible structure for the peptide-protein complex conforms with the sequence data for the at least one peptide and the at least one protein. In embodiments, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 7 of 108 generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises generating a template structure that conforms with the sequence data for the at least one peptide and the at least one protein using the initial structure and a homology modeling method. [0017] In some embodiments, the at least one auxiliary feature comprises a prediction of binding of the at least one peptide to the at least one protein. In some embodiments, the at least one auxiliary feature comprises a prediction of at least one interface feature for the peptide- protein complex. [0018] In some embodiments, the peptide sequence data comprises a peptide embedded representation for the at least one peptide generated using a trained peptide embedding machine learning model. In some embodiments, the protein sequence data comprises a protein embedded representation for the at least one protein generated using a trained protein embedding machine learning model. The trained peptide embedding machine learning model may take as input a peptide sequence and produce as output a peptide embedded representation. The trained protein embedding machine learning model may take as input a protein sequence and produce as output a protein embedded representation. The trained peptide embedding machine learning model and/or the trained protein embedding machine learning model may be machine learning models that have been trained in a self-supervised manner to learn a peptide/protein embedding from an input sequence. For example, the peptide embedding machine learning model and/or the trained protein embedding machine learning model may have been trained using unlabeled data (e.g. protein and/or peptide sequence data) and tasks such as masked language modelling, causal language modelling, etc. [0019] In some embodiments, the trained peptide embedding machine learning model and the trained protein embedding machine learning model are the same model. In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise foundational models. In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise protein language models. In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise large language models (LLM). In some embodiments, the trained peptide embedding machine learning model and/or the trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 8 of 108 protein embedding machine learning model comprise an Evolutionary Scale Model (ESM). In some embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are integrated with, and re-trained simultaneously with, the trained machine learning model. In embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are fine- tuned (e.g. after self-supervised training) for a peptide-protein complex related prediction task. Such training may be supervised training. In embodiments, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are fine- tuned for peptide-protein binding prediction. Fine tuning of a model (e.g. a foundation model) for a particular task may include incorporating the model as part of a machine learning model trained to perform the particular task, and at least partially retraining the parameters of the model for this task. The trained peptide embedding machine learning model and/or the trained protein embedding machine learning model may be deep learning models, e.g. deep neural networks. Any deep learning architecture known in the art suitable for use in language modeling may be used. [0020] In some embodiments, the trained machine learning model comprises a trained neural network. In some embodiments, the trained machine learning model comprises a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained state space model. [0021] In some embodiments, the trained machine learning model is trained using a training data set that includes incomplete structural data for one or more peptide-protein complexes. In some embodiments, the trained machine learning model is trained using a training data set that includes, for each of a plurality of training peptide-protein complexes, values of pairwise distances between peptide amino acid residues and protein amino acid residues in the training peptide-protein complex, and values of dihedral angles for peptide amino acid residues in the training peptide-protein complex, wherein the training data does not comprise all pairwise distances for one or more training peptide-protein complexes and/or does not comprise both dihedral angles ψ and φ for one or more peptide amino acid residues in one or more training peptide-protein complexes. In other words, such incomplete training peptide- protein complexes structural data may still have been used for training the machine learning model. In embodiments, the trained machine learning model has been trained using training sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 9 of 108 data including values of dihedral angles for peptide amino acid residues in training peptide- protein complexes for only one of the dihedral angles ψ and φ. [0022] In some embodiments, the trained machine learning model is further configured to determine an uncertainty in the predicted pairwise distance data and the predicted dihedral angle data. In embodiments, the trained machine learning model is configured to predict parameters characterizing: a distribution of pairwise distance for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and a joint distribution of dihedral angles ψ and φ for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex. In embodiments, the trained machine learning model is configured to predict parameters characterizing: a distribution of pairwise distance for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; a distribution of dihedral angle ψ for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; and a distribution of dihedral angle φ for the at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex. The parameters characterizing a distribution may comprise a parameter that characterizes the maximum likelihood value of the distribution (e.g. average) and a parameter that characterizes the uncertainty (e.g., spread) of the distribution (e.g., standard deviation or variance). [0023] In some embodiments, the machine learning model is trained or has been trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide- protein combinations; training a classifier to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and appending a regressor head to at least a portion of the trained classifier and training the regressor head to predict pairwise distance data and dihedral angle data (e.g. probabilistic distributions of pairwise distance data and dihedral angle data) for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from experimental (e.g. crystal) structure data for a set of example peptide- protein complexes. In some embodiments, the pairwise distance data and dihedral angle data derived from the experimental (e.g. crystal) structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide- sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 10 of 108 protein complex is presented as a separate training instance during training of the regressor head. In some embodiments, the machine learning model has been trained in at least two steps. In a first step the machine learning model may have been trained to predict peptide-protein binding for an input peptide sequence data and an input protein sequence data, using training data comprising peptide sequence data for a plurality of peptides; protein sequence data for a plurality of proteins; and peptide-protein binding data for a plurality of peptide-protein combinations. Predicting peptide-protein binding may comprise classifying peptide-protein combinations as binding / not binding to form a peptide-protein complex. Such training data may comprise peptide and protein sequence data for a plurality of peptides and proteins that are known to bind to each other to form a peptide-protein complex, and peptide and protein sequence data for a plurality of peptides and proteins that are not known or expected to bind to each other to form a peptide-protein complex. In the first step the machine learning model may comprise a peptide sequence embedding model, a protein sequence embedding model and a classification head. Such a model may be referred to as “classifier”. In a second step, at least a part of the machine learning model (e.g. the peptide sequence embedding model and the protein sequence embedding model) may have been trained to predict the predicted pairwise distance data and the predicted dihedral angle data (e.g. probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes). The training in the second step may be based on pairwise distance data and dihedral angle data derived from experimental (e.g. crystal) structure data for a set of example (i.e. training) peptide-protein complexes. In the first step the machine learning model may comprise the peptide sequence embedding model, the protein sequence embedding model and a regression head. Such a model may be referred to as “regressor”. [0024] In some embodiments, the machine learning model (e.g. regressor) is further trained or has been further trained on augmented pairwise distance data and dihedral angle data. Augmented pairwise distance data and dihedral angle data may refer to data that has not been experimentally determined. Such data may have been obtained for example using a structural modeling method, such as e.g., using a trained structure prediction machine learning model, e.g. a model of the AlphaFold family. Such data may have been obtained by obtaining a predicted crystal structure for a peptide and protein sequence known to bind to each other to form a peptide-protein complex using a trained structure prediction machine learning model. In some embodiments, the augmented pairwise distance data and dihedral angle data is derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 11 of 108 least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide- protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide- protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor. [0025] In some embodiments, identifying the initial structure for the peptide-protein complex based on the predicted pairwise distance data and predicted dihedral angle data comprises: obtaining a plurality of experimental (e.g., crystal) structures for peptide-protein complexes from a protein structure database; and selecting an experimental (e.g. crystal) structure from the plurality of experimental (e.g. crystal) structures based on a sequence and/or structural similarity between the experimental structure and the peptide-protein complex. A structural similarity may be a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model. In embodiments, the experimental structure is selected based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and a maximum likelihood estimate for the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model. In embodiments, selecting an experimental structure based on a structural similarity comprises: (i) for each example peptide-protein experimental structure in the plurality of experimental structures, determining a likelihood (e.g. average likelihood, average log-likelihood) of the distances and dihedral angles in the peptide-protein experimental structure using the predicted distributions for pairwise distances and dihedral angles for the peptide-protein complex output by the trained machine learning model; and (ii) selecting the example peptide-protein experimental structure that has the highest likelihood. In some embodiments, the protein structure database comprises the Protein Data Bank (PDB). In embodiments, the initial structure comprises a plurality of initial structures. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 12 of 108 [0026] In some embodiments, generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises: inputting the initial structure for the peptide- protein complex into a structural modeling software package; inputting the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, wherein compatible structures for the peptide-protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function. In embodiments, the one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data are included in the structure scoring function. [0027] In some embodiments, the iterative sampling from the plurality of possible structures for the peptide-protein complex is performed using a Markov Chain – Monte Carlo (MCMC) sampling algorithm. In some embodiments, the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Metropolis – Hastings algorithm. In some embodiments, the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Hamiltonian Monte Carlo algorithm. In some embodiments, the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Langevin Monte Carlo algorithm. [0028] In some embodiments, the structure scoring function comprises a proxy for a free energy calculation. In some embodiments, the structure scoring function comprises a proxy for a free energy calculation of peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof. In some embodiments, the scoring function is based on one or more of ab initio quantum mechanical calculations, density functional theory (DFT) calculations, semiempirical calculations, molecular mechanics force field calculations, statistical potential calculations, neural potential calculations, or machine learning models trained on structural data. In some embodiments, the structural modeling software package comprises FlexPepDock. [0029] In some embodiments, the computer-implemented method further comprises iterating over a plurality of candidate peptides to identify at least one peptide that has a maximum likelihood for formation of the peptide-protein complex. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 13 of 108 [0030] In some embodiments, the at least one peptide comprises at least a portion of a peptide that is expressed in normal cells in a subject. In some embodiments, the at least one peptide comprises at least a portion of a neoantigen expressed in tumor cells from a cancer patient. In some embodiments, the at least one peptide comprises at least a portion of a peptide that is expressed in a virus or bacterium. [0031] In some embodiments, the at least one protein comprises a Major Histocompatibility Complex (MHC) class I protein. In some embodiments, the at least one protein comprises a Major Histocompatibility Complex (MHC) class II protein. In some embodiments, the computer-implemented method further comprises iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in normal cells from the subject for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. In some embodiments, the computer-implemented method further comprises iterating over a plurality of candidate neoantigen sequences, or portions thereof, expressed in tumor cells from the cancer patient for at least one MHC class I or class II protein to rank order the neoantigen sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC- II) complex. In some embodiments, the computer-implemented method further comprises iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in a virus or bacterium for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. [0032] In some embodiments, the computer-implemented method (or a method associated with the computer-implemented method) further comprises developing a personalized anti- cancer vaccine based on the rank ordering of the neoantigen sequences, or portions thereof. In some embodiments, the computer-implemented method (or a method associated with the computer-implemented method) further comprises developing an anti-viral vaccine, an anti- bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof. In some embodiments, the computer-implemented method (or a method associated with the computer-implemented method) further comprises predicting one or more characteristics of a complex comprising the peptide-protein complex and a further protein. The further protein may be e.g. a T cell receptor, antigen recognition molecule, T cell sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 14 of 108 engaging molecule, etc. The predicted characteristics may include one or more of: the likelihood of formation of the complex, the stability of the complex, the binding energy of the complex, the total energy of the complex. [0033] Also disclosed herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform any of the computer-implemented methods described herein. [0034] Also disclosed herein are non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform any of the computer-implemented methods described herein. [0035] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been disclosed by specific, exemplary implementations and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. INCORPORATION BY REFERENCE [0036] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls. BRIEF DESCRIPTION OF THE DRAWINGS [0037] The present disclosure is described in conjunction with the appended figures: sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 15 of 108 [0038] FIG. 1 provides a block diagram of an example prediction system, in accordance with some implementations of disclosed methods and systems. [0039] FIG. 2 provides a non-limiting example of a block diagram for a system for generating conformer ensembles for a peptide-protein complex, in accordance with some implementations of the disclosed methods and systems. [0040] FIG. 3A provides a schematic illustration of using a protein language model to convert a peptide sequence into a peptide embedded representation, in accordance with some implementations of the disclosed methods and systems. [0041] FIG. 3B provides a schematic illustration of using a protein language model to convert a protein sequence (e.g., an MHC sequence) into a protein embedded representation, in accordance with some implementations of the disclosed methods and systems. [0042] FIG.4 provides a schematic illustration of using a trained machine learning model (e.g., a Structural Constraint Prediction Model) to output pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a pMHC complex) based on input embedded representations of peptide and protein (e.g., MHC) sequences, in accordance with some implementations of the disclosed methods and systems. [0043] FIG.5 provides a schematic illustration of using pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a peptide – MHC complex) output by a trained machine learning model (e.g., a Structural Constraint Prediction Model) to generate conformer ensembles for the peptide-protein complex (e.g., a pMHC complex), in accordance with some implementations of the disclosed methods and systems. [0044] FIG. 6 provides a non-limiting schematic illustration of a process for training a machine learning model to predict pairwise distance data and dihedral angles for a peptide- protein complex, in accordance with one implementation of the disclosed methods. [0045] FIG.7 provides a non-limiting schematic illustration of a process of using a trained machine learning model to predict pairwise distance data and dihedral angles for a peptide- protein complex which can then be used to: (i) select a template structure from a protein structure database, and (ii) provide constraints on allowable conformations when generating an ensemble of conformers for the peptide-protein complex, in accordance with one implementation of disclosed methods. MHC pseudosequence YFAMYGEKV…(SEQ ID NO:1); peptide sequence LLFGYPVYV (SEQ ID NO:2). sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 16 of 108 [0046] FIG.8 provides a non-limiting example that illustrates how predicted distributions of pairwise distance data and dihedral angle data for a pMHC complex can be converted to structural constraints, and how pMHC conformational ensembles can then be generated based on such structural constraints. [0047] FIG. 9 provides a non-limiting example of a process flowchart for generating a plurality of compatible structures for a peptide-protein complex, in accordance with one implementation of the disclosed methods. [0048] FIG. 10 provides a non-limiting example of a process flowchart for training a machine learning model to predict distributions of pairwise distance data and dihedral angle data for a peptide-protein complex, in accordance with one implementation of the disclosed methods. [0049] FIG. 11A provides a non-limiting example of data for the root mean square deviation (RMSD) of predicted Φ and Ψ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where the Φ and Ψ dihedral angles were treated independently during training and inference. [0050] FIG. 11B provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of predicted Φ and Ψ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where the Φ and Ψ dihedral angles were treated either independently or jointly during training and inference. [0051] FIG. 11C provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of predicted Φ and Ψ dihedral angles in a set of benchmark pMHC complexes comprising peptides of different length, where one data set (red symbols) was generated using a prior protein structure prediction model (AlphaFold – FineTune), and the other data set (green symbols) was generated using one implementation of the methods disclosed herein. [0052] FIG. 11D provides a non-limiting example of preliminary data for predicted average dihedral distance in a set of benchmark pMHC complexes comprising peptides of different length. [0053] FIG.11E provides a non-limiting example of preliminary data for the recovery of crystal structure contacts for a set of benchmark pMHC complexes using the disclosed structural prediction methods. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 17 of 108 [0054] Fig.12A provides a non-limiting example of data for the recovery of distances for test sets of pMHC complexes using the disclosed structural prediction methods. [0055] Fig.12B provides a non-limiting example of data for the recovery of dihedral angles for test sets of pMHC complexes using the disclosed structural prediction methods. [0056] Fig. 12C provides a non-limiting example of data for the recovery of dihedral angles using the disclosed structural prediction methods, for a representative pMHC 9-mer example (1I1Y). [0057] Fig. 13 provides a non-limiting example of data for the number of overlapped models in conformer ensembles predicted using methods of the present disclosure or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations. [0058] Fig. 14A provides a non-limiting example of evaluation metrics comparing conformer ensembles obtained using the disclosed structural prediction methods or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations in a first test set. [0059] Fig. 14B provides a non-limiting example of evaluation metrics comparing conformer ensembles obtained using the disclosed structural prediction methods or a baseline method that does not use ML-predicted structural constraints, to corresponding benchmark molecular dynamics simulations in a second test set. [0060] Fig.15A provides a non-limiting example of the ability of the disclosed structural prediction methods to capture dynamic behaviors of peptides involved in TCR recognition. [0061] Fig. 15B provides data for the same peptide as in Fig. 15B, illustrating that molecular dynamics (MD) simulations are unable to overcome the energy barriers associated with sidechain flipping in the TCR bund conformation of the peptide. [0062] FIG. 16A provides a non-limiting example of control data for pMHC template structure selection. [0063] FIG. 16B provides a non-limiting example of control data for pMHC template structure refinement. [0064] FIG. 17A provides a non-limiting example of molecular dynamics data for a sampling of peptide conformations within a pMHC complex. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 18 of 108 [0065] FIG. 17B provides a non-limiting example of a graphical representation of the pMHC complex of FIG.17A. [0066] FIG. 17C provides a non-limiting example of a plot of peptide backbone fluctuations as a function of residue number for the pMHC complex of FIG.17A. [0067] FIG. 18A provides a non-limiting example of molecular dynamics data for a sampling of peptide conformations within a pMHC complex. [0068] FIG. 18B provides a non-limiting example of a graphical representation of the pMHC complex of FIG.18A. [0069] FIG. 18C provides a non-limiting example of a plot of peptide backbone fluctuations as a function of residue number for the pMHC complex of FIG.18A. [0070] FIG.19 provides a non-limiting example of a block diagram of a computer system, in accordance with some implementations of the methods and systems disclosed herein. [0071] FIG. 20 provides a non-limiting example of a block diagram of an artificial intelligence (AI) architecture included as part of the example computing system of FIG.4, in accordance with some implementations of the methods and systems disclosed herein. [0072] In the appended figures, similar components and/or features can have the same reference label. Further, various components of the same type can be distinguished by following the reference label by a dash and a second label that distinguishes among the similar components. If only the first reference label is used in the specification, the description is applicable to any one of the similar components having the same first reference label irrespective of the second reference label. DETAILED DESCRIPTION [0073] Machine learning-based methods for modeling conformational states of peptide- protein complexes that account for the dynamic nature of three-dimensional peptide and protein structures are described. The disclosed methods can be used to generate an ensemble of allowed peptide-protein conformers (e.g., an ensemble of allowed three-dimensional structures for a peptide-protein complex that can exist at any given point in time – also referred to herein as “compatible structures” for the peptide-protein complex, which can impact the function of the peptide-protein complex at any given point in time). The ensemble of allowed peptide-protein conformers can be represented as an allowed distribution of structural parameters for the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 19 of 108 peptide-protein complex. The ensemble of allowed peptide-protein conformers can be used to more accurately predict features and functional properties of the peptide-protein complex (e.g., complementary interface features, complementary surface fingerprint features, and/or binding of the peptide-protein complex to other proteins or protein complexes). [0074] In some instances, for example, the disclosed methods for generating a plurality of compatible structures for a peptide-protein complex can comprise a first step of inputting peptide sequence data (e.g. an amino acid sequence or an embedding thereof) for at least one peptide, and protein sequence data (e.g. an amino acid sequence or pseudo-sequence, or an embedding thereof) for at least one protein (which can be a full length protein or a truncated version thereof including at least the portion of the protein interacting with the peptide) into a trained machine learning model (e.g., a structural constraint prediction model or a machine learning model comprising a structural constraint prediction model and one or more embedding models configured to obtain a protein embedded representation and a peptide embedded representation, also referred to herein as “protein embedding” and “peptide embedding”). The structural constraint prediction model can be configured (e.g. trained) to process embedded representations of peptide and protein sequences to determine: (i) predicted pairwise distance data (or a distribution thereof) for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide- protein complex (i.e. one or more predicted distances for each of one or more pairs of residues comprising a peptide residue and a protein residue or parameters of a predicted distribution of said distances, where distances between residues can be expressed as distances between predetermined atoms of the residues such as Cα of the residues); and (ii) predicted dihedral angle data (or a distribution thereof) for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex (i.e. one or more predicted values of dihedral angle ψ and/or one or more predicted values of dihedral angle φ for each of one or more peptide residues in the peptide-protein complex, or parameters of respective distributions of said dihedral angles). The predicted structural constraint data can then be used to identify an initial structure for the peptide-protein complex and generate the plurality of compatible structures for the peptide-protein complex based on the initial structure, and the predicted values or distributions of pairwise distances and dihedral angles. [0075] In some instances, the method can further comprise predicting binding of the peptide-protein complex to another protein or protein complex based on the plurality of sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 20 of 108 compatible structures for the peptide-protein complex. For example, in some instances, predicting binding of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex comprises: identifying complementary interface features between the peptide-protein complex and the other protein or protein complex; identifying complementary surface fingerprint features between the peptide-protein complex and the other protein or protein complex; identifying latent representations of the peptide- protein complex and the other protein or protein complex; and providing the complementary interface features, the complementary surface fingerprint features, and the latent representations of the peptide-protein complex and the other protein or protein complex as input to a second trained machine learning model configured (e.g. trained) to output a prediction for binding of the peptide-protein complex and the other protein or protein complex using said inputs. [0076] The improvements in processing performance and prediction accuracy enabled by the disclosed methods and systems are based on the use of embedded representations of peptide and protein sequences as input, machine-learning based prediction of structural constraints for the peptide-protein complex (e.g., predicted pairwise distance data or posterior distributions for pairwise distance data for amino acid residues in the peptide and the protein, and predicted dihedral angle data or posterior distributions for dihedral angle data for the peptide in the peptide-protein complex), and the use of the predicted structural constraints to select an initial template structure for the peptide-protein complex and generate a plurality of compatible structures that captures the dynamic behavior of the peptide-protein complex in solution. As noted above, prior methods have failed to account for the dynamic properties and range of allowed conformational states of peptide-protein complexes in solution. ML-based generation of conformer ensembles for peptide-protein complexes [0077] FIG. 1 provides a block diagram of an example prediction system, in accordance with some embodiments. Prediction system 100 can be used, for example, to predict structural properties of peptide-protein complexes (e.g., pMHC complexes), to determine a plurality of allowable conformations (e.g., a conformer ensemble) for a peptide-protein complex (e.g., a pMHC complex), and/or to determine a predicted pMHC interaction with one or more other proteins in an immunoprotein complex (IPC) related to the immunological activity of peptides (such as e.g. interaction between a pMHC and a T cell receptor molecule (TCR) or a part thereof) and, in particular, mutated peptides (e.g., neoantigen peptides). The prediction system sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 21 of 108 100 includes computing platform 102, data store 104, and display system 106. Computing platform 102 may take various forms. In some embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers in communication with each other. In some embodiments, the computing platform 102 can be a cloud computing platform. [0078] Data store 104 and display system 106 are each in communication with computing platform 102. In some examples, one or more of: data store 104 or display system 106 can be considered part of, or otherwise integrated with, computing platform 102. Thus, in some examples, computing platform 102, data store 104, and display system 106 can be separate components in communication with each other, but in other examples, some combination of these components can be integrated together. Communication between the different components can be implemented using any number of wired communications links, wireless communications links, optical communications links, or a combination thereof. [0079] The prediction system 100 includes a sequence analyzer 108, which can be implemented using hardware, software, firmware, or a combination thereof. In some embodiments, the sequence analyzer 108 can be implemented in the computing platform 102. The sequence analyzer 108 receives sequence data 110 for processing. For example, the sequence data 110 can be sent as input into the sequence analyzer 108, retrieved from the data store 104 or some other type of storage (e.g., cloud storage), accessed from cloud storage, or obtained in some other manner. In some cases, the sequence data 110 can be retrieved from the data store 104 in response to receiving user input entered by a user via an input device (not shown in FIG.1). [0080] The sequence data 110 can be generated from processing of a set of samples 112. The set of samples 112 may take the form of one or more biological samples from one or more subjects (e.g., a diseased sample, a healthy sample, or a combination thereof). The set of samples 112 may include a sample obtained from a tumor of a subject (e.g. a biopsy or tumor resection sample), or a sample comprising tumor genetic material from a subject (e.g. a liquid biopsy). The tumor can be a manifestation of, for example, lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myelogenous leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, T cell sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 22 of 108 lymphocytic leukemia, non-small cell lung cancer, small-cell lung cancer, or a combination thereof. [0081] A sample in the set of samples 112 may include, for example, various IPC molecules, various peptides, nucleic acids coding for said IPC molecules and/or peptides, or a combination thereof. When the set of samples 112 includes a diseased sample, the peptides may include one or more mutated peptides (e.g., neoantigens). The IPC molecules may include, for example, various MHC molecules, various TCR molecules, or a combination thereof. [0082] In some embodiments, the set of samples 112 includes IPC 114 (e.g., MHC Class I molecule, MHC Class II molecule, various TCR molecules, etc.) and/or nucleic acid sequences coding for IPC 114, from which the sequence of IPC 114 can be inferred. Further, the set of samples can include at least one protein 123 (i.e., the source protein) or a nucleic acid sequence coding for protein 123 or a part thereof, from which the sequence of the protein 123 (referred to herein as protein sequence 160) or part thereof (such as amino acid chain 116 or peptide 118) can be inferred. An amino acid chain 116 can be identified from the at least one protein 123 and can be a chain of amino acids that includes a peptide 118 (minimal epitope). The amino acid chain 116 can optionally additionally include an N-flank 120, and/or a C-flank 122, referring respectively to a sequence located at the N-terminus side of the peptide 118 and the C-terminus side of the peptide 118. The peptide 118 can be considered a mutated peptide when it includes one or more variants (e.g., one or more sequence variations) when compared to a corresponding reference sequence. The protein 123 can be a source protein for the amino acid chain 116, which can be generated through proteolysis, which is the process by which proteins (e.g., the protein 123) are broken down into smaller polypeptides or amino acids. The protein 123 can be broken down into smaller polypeptides or amino acids by enzymatic cleavage, where specific enzymes called proteases cut the peptide bonds between amino acids in the protein 123. [0083] The set of samples 112 can be processed to generate the sequence data 110. In some embodiments, multiple samples in the set of samples 112 can be processed at different times. In some embodiments, the prediction system 100 includes a sample analyzer that can be used in processing the set of samples 112 to generate the sequence data 110. The sequence data 110 includes, for example, at least one amino acid sequence 129 and at least one IPC sequence 124 (e.g., one IPC sequence 124 corresponding to IPC 114). The amino acid sequence 129 may comprise one or more of: a peptide sequence 126 (e.g., one peptide sequence 126 corresponding sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 23 of 108 to peptide 118), an amino-terminal flanking (N-flank) sequence 128 (e.g., one N-flank sequence 128 corresponding to N-flank 120), or a carboxy-terminal flanking (C-flank) sequence 130 (e.g., one C-flank sequence 130 corresponding to C-flank 122). One or more sub- sequences of amino acid sequence 129 (e.g., peptide sequence 126, N-flank sequence 128, and C-flank sequence 130) can be processed separately or as a single sequence. [0084] When IPC 114 is an MHC, IPC sequence 124 can be, for example, an MHC sequence 135 that characterizes at least a portion of the MHC. When IPC 114 is a TCR, IPC sequence 124 can be, for example, a TCR sequence 131 that characterizes at least a portion of the TCR. In some embodiments, IPC sequence 124 may include both an MHC sequence 135 characterizing at least a portion of an MHC molecule and a TCR sequence 131 characterizing at least a portion of a TCR molecule. In some embodiments, the sequence data 110 may include IPC sequence 124 in the form of an MHC sequence 135 characterizing at least a portion of an MHC molecule, as well as a separate TCR sequence 131 characterizing at least a portion of a TCR. [0085] Protein sequence 160 characterizes at least a portion of the protein 123. In some embodiments, the protein sequence 160 can be identified by performing a reverse lookup in a database (e.g., the UniProt database) based on the mutated peptide data (e.g., IPC sequence 124) obtained from the sample. Protein sequence 160 may be used to provide context and/or orthogonal information when characterizing a peptide 118, for example to compare a peptide 118 to a wild-type equivalent, to determine expression of the protein 123 in one or more samples or tissues, etc. Availability of the protein sequence 160 is completely optional, and methods of the disclosure can make use of e.g. only peptide sequence 118, peptide sequence 116 or parts thereof. [0086] Peptide sequence 126 characterizes at least a portion of the peptide 118. N-flank sequence 128 characterizes at least a portion of the N-flank 120. In some embodiments, when the number of amino acids (or amino acid residues) upstream from the N-terminus of the peptide is large, the corresponding sequence for N-flank 120 can be trimmed to generate the N-flank sequence 128. In other words, N-flank sequence 128 can comprise or consist of the sequence of a predetermined number of residues upstream (i.e. N-terminal) of peptide 118. C- flank sequence 130 characterizes at least a portion of the C-flank 122. In some embodiments, when the number of amino acids (or amino acid residues) downstream from the C-terminus is large, the corresponding sequence for C-flank 122 can be trimmed to generate the C-flank sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 24 of 108 sequence 130. In other words, C-flank sequence 130 can comprise or consist of the sequence of a predetermined number of residues downstream (i.e. C-terminal) of peptide 118. [0087] Sequence analyzer 108 receives the sequence data 110 as input for processing. The sequence analyzer 108 includes one or more machine-learning models 132 that process the sequence data 110. In some embodiments, the sequence analyzer 108 can process the sequence data 110 (e.g., using embedding engine(s) 140) prior to sending the sequence data 110 (or an embedded form thereof) into one or more of the machine learning models 132 for further processing. In some embodiments, machine learning models 132 may comprise, for example, one or more of: one or more embedding engines 140, a structural constraint prediction model 141, and a conformer ensemble generation engine 142. [0088] Embedding generation engine(s) 140 can be configured to process peptide and/or protein sequence (e.g., MHC sequence) data (e.g. IPC sequence 124 and/or amino acid sequence 129) and output embeddings for the respective sequences. As used herein, an embedding of a peptide or protein sequence can refer to a vector representation of a peptide or protein sequence that encodes important features of the peptide or protein in a way that facilitates processing by downstream computational models. An embedding of a peptide or protein sequence can be a latent representation of a machine learning model trained to learn a latent space in which input sequences can be represented and from which the sequences can be reconstructed and/or one or more features of the input sequences can be predicted. Structural constraint prediction model 141 can be configured to process peptide and/or protein embeddings for a peptide-protein complex, which are outputted by the embedding generation engine(s) 140, and output a predicted set of structural constraints associated with the peptide- protein complex. The predicted set of structural constraints can include e.g., predicted pairwise distance data (or a distribution thereof) for at least one peptide amino acid residue and at least one protein amino acid residue in the peptide-protein complex and/or predicted dihedral angle data (or a distribution thereof) for at least one peptide amino acid residue in the peptide in the peptide-protein complex. Conformer ensemble generation engine 142 can be configured to generate a plurality of compatible structures for the peptide-protein complex (i.e., a conformer ensemble 148) based on an initial template structure and the set of structural constraints output by the structural constraint prediction model 141. Thus, the conformer ensemble generation engine 142 can be configured to take as input an initial template structure and a set of structural constraints output by the structural constraint prediction model 141 for a peptide-protein sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 25 of 108 complex, and produce as output a plurality of compatible structures for the peptide-protein complex. The plurality of compatible structures can form a conformer ensemble, also referred to as conformational ensemble, structural ensemble or simply ensemble. A conformer ensemble for a peptide-protein complex can be a set of conformations that together describe the structure of the peptide-protein complex. The conformer ensemble may be expected to capture a range of conformations that the peptide-protein complex may adopt under normal (e.g. physiological) conditions. [0089] The one or more machine-learning models 132 can be used in either a training mode or a prediction mode. In the training mode, one or more of the machine-learning models 132 can be trained using training data 133. Examples of the training data 133 are described in more detail below. The one or more machine-learning models 132 are trained such that they can be used in the prediction mode. [0090] One or more of the machine-learning models 132 are used to process the IPC sequence 124 and the amino acid sequence 129. In some embodiments, separate processing engines (such as e.g. separate embedding engines 140) can be used for processing IPC sequences and amino acid sequences. This can enable improved predictive performance of the one or more machine-learning models 132. In some embodiments, one or more of the machine- learning models 132 process one or more of: a MHC sequence 135, a TCR sequence 131, a protein sequence 160, a peptide sequence 126, an N-flank sequence 128, or a C-flank sequence 130. Examples of implementations for these different processing engines are described in greater detail below. [0091] As used herein, the terms “processing engine”, “engine”, and “model” identify at least one software component and/or a combination of at least one software component and at least one hardware component which are designed/programmed/configured to interact and/or communicate data to other software and/or hardware components including but not limited to other processing engines. [0092] The one or more machine-learning models 132 process the sequence data 110 to generate an output that can be used to generate a report 144. The report 144 may include the exact output of any one or more of the machine-learning models 132, a transformed (e.g. further processed) or filtered version of the output of any one or more of the machine learning models 132, or both. In some cases, the report 144 may include notifications, recommendations, alerts, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 26 of 108 or other information generated by the sequence analyzer 108 based on the output of one or more of the machine-learning models 132. [0093] The report 144 can be an output that includes, for example, information about structural constraints for a peptide-protein complex (e.g., a pMHC complex) such as e.g., information output by structural constraint prediction model 141 or information derived therefrom, a conformer ensemble 148 for a peptide-protein complex (e.g., a pMHC complex) such as e.g. a conformer ensemble 148 output by ensemble generation model 142, predicted features of a peptide-protein complex (e.g., a pMHC complex) such as e.g. predicted features derived at least in part from a conformer ensemble 148 output by ensemble generation model 142, which can include e.g., immunological activity of interest with respect to one or more peptides (e.g., one or more mutated peptides). For example, the report 144 may include information about structural predictions 146 (e.g., predicted distributions of pairwise distances for amino acid residues and/or dihedral angles) for a peptide protein complex (e.g.., a pMHC complex), conformer ensembles 148 (e.g., a set of allowed conformations that are consistent with the structural predictions) for a peptide protein complex (e.g.., a pMHC complex), peptide-protein complex feature predictions 162, an immunological activity prediction relating to the amino acid 116 (e.g., peptide 118, N-flank 120, C-flank 122, etc.) and IPC 114 (e.g., MHC-I, MHC-II, TCR, etc.), or any combination thereof. The report 144 may include, for example, interaction information (e.g., an interaction affinity prediction that predicts a binding affinity between a peptide and an MHC, or an interaction prediction that predicts whether an MHC allele or allotype will present a peptide at a cell surface), immunogenicity information (e.g., an immunogenicity prediction that predicts the ability of a peptide to provoke an immune response in the context of an MHC molecule), or both. The interaction information may provide predictions about a selected set of interactions between the amino acid sequence 129 and the IPC sequence 124. The immunogenicity information may provide predictions about the immunogenicity of amino acid sequence 129 (e.g., including the immunogenicity of the peptide 118). The immunogenicity information may provide predictions about the immunogenicity of peptide 118 in the context of the IPC 114, where the IPC is an MHC molecule. The immunogenicity information may provide predictions about the immunogenicity of peptide 118 (or a peptide-protein complex comprising the peptide 118 and an IPC114 where the IPC is an MHC molecule) in relation to an IPC 114, where the IPC in a TCR molecule. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 27 of 108 [0094] In some embodiments, a report 144 can be displayed on a graphical user interface (GUI) 150 on the display system 106. A user may view and/or interact with the report 144 via the graphical user interface 150. In some embodiments, the user may use the report 144 to make decisions about the treatment of a subject from which at least one of the set of samples 112 was obtained (or collected). In some embodiments, the user may use the report 144 to make decisions about the manufacture of a treatment for a subject from which at least one of the set of samples 112 was obtained (or collected). For example, the user may use the report 144 to select a peptide sequence for use in manufacturing a vaccine (such as e.g. a personalized cancer vaccine). [0095] In some embodiments, the prediction system 100 sends the report 144 to the remote system 152 (e.g., wirelessly). The remote system 152 can be a cloud computing platform, cloud storage, another computer system, a user device (e.g., a smartphone, a tablet, a laptop, etc.), or some other type of platform. In some embodiments, the remote system 152 can be a treatment manufacturing system (or machine) or a portion thereof. [0096] FIG.2 provides a non-limiting example of a block diagram for a system 200 (e.g., a computer-implemented system) for generating conformer ensembles for a peptide-protein complex. System 200 can be implemented using the Prediction System 100 described in FIG. 1. For example, system 200 may comprise the Sequence Analyzer 108 and one or more of the Machine Learning Models 132 in FIG.1. [0097] The input to the system 200 includes protein sequence data (e.g., MHC Sequence 135) for at least one protein and peptide sequence data (e.g., Peptide Sequence 126) for at least one peptide. [0098] In some instances, the input can comprise peptide sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more than 50 peptides, or portions thereof. In some instances, peptide sequence data can comprise sequence data for peptides that are at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues in length. In some instances, the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof. In some instances, the peptides can comprise mutated peptides (e.g., neoantigens expressed in tumor cells), or portions thereof. The normal peptide and/or mutated peptides may have been isolated from a subject (e.g., a patient) or otherwise identified from a sample previously obtained from a subject (such as e.g. by analyzing nucleic acid sequence data obtained from a sample from said subject comprising sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 28 of 108 normal and/or tumor derived nucleic acids). In some instances, the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium. In some instances, the peptides can comprise viral peptides, or portions, thereof expressed in a virus. [0099] In some instances, the input can comprise protein sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, or more than 1000 proteins, or portions thereof. In some instances, protein sequence data can comprise sequences or pseudosequences of at least 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or at least 200 amino acid residues in length. In some instances, the proteins can comprise IPC proteins. In some instances, the IPC proteins (also referred to herein as IPC molecules) may include, for example, MHC-I molecules, MHC-II molecules, TCR molecules, or a combination thereof. [0100] Input protein sequence (e.g. MHC Sequence 135) and Peptide Sequence 126 can be processed by Embedding Generation Engine(s) 140 to generate protein embedded representation (e.g. MHC embedded representation 206) and Peptide embedded representation 208. Protein embedded representations such as MHC embedded representation 206 and Peptide embedded representation 208 can be represented as data structures in the form of vectors. For example, each of embedded representation can be represented as a vector of size (t, e) where each residue of a protein (e.g., IPC 114, and peptide 118 shown in FIG.1) is represented by a token, t, and the contextual information of each residue is encoded in (e) dimensions. In some instances, in addition to “residue tokens”, Embedding Generation Engine(s) 140 can add “special tokens” which can be used to denote the beginning or end of the protein sequence. For an embedded representation instantiated as a vector of size (t, e), t can represent the total number of residue tokens and special tokens. Protein embedded representations can encode important features of a protein or peptide in a way that can be processed by downstream computational models. [0101] A non-exhaustive list of features that can be implicitly encoded in a protein embedded representation (i.e. each of these have been shown to be properties that can be captured in a protein embedded representation obtained from a protein/peptide embedding model can capture for an input protein sequence) include: [0102] Amino Acid Sequence: The sequence of amino acids that make up the protein. This sequence determines the protein’s structure and function. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 29 of 108 [0103] Structural Properties: Aspects of the protein’s three-dimensional structure, such as alpha-helices, beta-sheets, and other secondary structures. [0104] Functional Properties: Information about the protein’s biological function, such as enzymatic activity, binding sites, or role in cellular processes. [0105] Evolutionary Information: Proteins with similar sequences or structures are often evolutionarily related. Some features can capture evolutionary relationships and similarities, which can be useful for tasks like protein classification or function prediction. [0106] Interaction with Other Molecules: Information about how proteins interact with other molecules like DNA, RNA, or small molecules (like drugs) can also be encoded. This includes binding sites and interaction domains. [0107] Post-Translational Modifications: Proteins often undergo modifications after translation (like phosphorylation or glycosylation) which affect their function. These modifications can sometimes be inferred from the sequence and context and thus can be encoded as features. [0108] Physical and Chemical Properties: Attributes like hydrophobicity, charge, size, and shape of amino acid residues, which influence how proteins fold and interact with other molecules. [0109] Contextual Information: In some models, the context within which a protein or a part of a protein appears (such as in a particular type of cell or a specific metabolic pathway) can be encoded. [0110] In some instances, the Embedding Generation Engine(s) 140 includes a Protein Embedding Model 202 for receiving an MHC sequence 135 and generating an MHC embedded representation 206, and a Peptide Embedding Model 204 for receiving a peptide sequence 126 and for generating a peptide embedded representation 208. In some instances, the Peptide Embedding Model 204 and the Protein Embedding Model 202 are the same model; in some instances, the Peptide Embedding Model 204 and the Protein Embedding Model 202 are different models. Example methods for generating peptide and/or protein sequence embedding are described in, e.g., Ofer et al. (2021), “The Language of Proteins: NLP, Machine Learning & Protein Sequences”, Computational and Structural Biotechnology Journal 19:1750-1758. [0111] In some instances, the Peptide Embedding Model 204 and/or the Protein Embedding Model 202 can comprise large language models (LLMs) or protein language sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 30 of 108 models (PLMs), e.g., deep learning algorithms configured to strings of characters in order to perform predictive or generative tasks such as recognizing, summarizing, translating, predicting, and generating content. LLMs and PLMs (which are special cases of LLMs trained using protein sequence data) are typically trained at least partially in an unsupervised manner (i.e. trained using unlabeled data and tasks such as masked language modelling, causal language modelling, etc.) using very large data sets (e.g., very large peptide and/or protein sequence data sets). LLMs and PLMs are frequently (but not exclusively) transformer based deep neural networks. Large language models are described in more detail in, for example, Li et al. (2021), “BioSeq-BLM: A Platform for Analyzing DNA, RNA and Protein Sequences Based on Biological Language Models”, Nucleic Acids Res.49, e129–e129, and Naveed et al. (2023), “A Comprehensive Overview of Large Language Models”, arxiv.org/pdf/2307.06435.pdf. [0112] In some instances, the Peptide Embedding Model 204 and/or the Protein Embedding Model 202 can comprise an Evolutionary Scale Model (ESM), e.g., a transformer- based LLM trained on unlabeled data for over 250 million distinct protein sequences (specifically, using all sequence data from UniParc - see UniProt Consortium, The universal protein resource (UniProt). Nucleic Acids Res.36, D190–D195 (2008)) using masked language modeling. ESM was shown to generate a vector representation for an input protein sequence that can encode biochemical properties of amino acids in the sequence, the biological variation and remote homology of the sequence, and secondary structure and tertiary contacts of the sequence (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118). [0113] MHC embedded representation 206 and Peptide embedded representation 208 are then used as the input to trained machine learning model (Structural Constraint Prediction Model 141) configured to process the embedded sequence data and output, for example, Pairwise Distance Data 210 (i.e., predicted pairwise distance data for at least one peptide amino acid residue of the peptide and at least one protein amino acid residue of the protein in a peptide-protein complex) and Dihedral Angle Data 212 (i.e., predicted dihedral angle data for at least one peptide amino acid residue in the peptide in the peptide-protein complex). [0114] In some instances, the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter. A posterior distribution sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 31 of 108 can be characterized by, for example, a mean value, and a standard deviation or any other measure of uncertainty / variability (or any other set of statistical parameters for characterizing a distribution, which depend on the type of distribution that is inferred). In some instances, Structural Constraint Prediction Model 141 can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution). [0115] In some instances, Structural Constraint Prediction Model 141 can comprise a trained neural network, as described in more detail elsewhere herein. In embodiments, the trained neural network is a deep neural network. In some instances, the trained neural network can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, a trained recurrent neural network, a trained state space model, or other suitable type of neural network or deep learning model. In embodiments, Structural Constraint Prediction Model 141 comprises a multilayer perceptron. [0116] Pairwise Distance Data 210 and Dihedral Angle Data 212 can then be used as input to an Ensemble Generation Engine 142 (along with an initial template structure for the pMHC complex) that is configured to use the Pairwise Distance Data 210 and the Dihedral Angle Data 212 as constraints and, starting with the initial template structure, generate a plurality of compatible structures (i.e., Conformer Ensemble 148) for the pMHC complex, where compatible structures for the pMHC complex are structures that: (i) are compatible with one or more constraints imposed by the predicted Pairwise Distance Data 210 and the predicted Dihedral Angle Data 212, and (ii) optimize a structure scoring function, as will be discussed in more detail below. [0117] In some instances, the method can further comprise iterating over a plurality of candidate peptide sequences, or portions thereof, for at least one MHC class I or class II protein. In some embodiments, the candidate peptide sequences, or portions thereof, may be ranked according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex and/or one or more features of one or more peptide-protein complexes comprising the candidate peptide sequences and MHC I or MHC II proteins, predicted using the conformer ensembles 148 obtained for the respective complexes (such as e.g. peptide-MHC-TCR complexes, or other complexes comprising a peptide-MHC and antigen recognition molecule). For example, metrics such as scores estimated from one or more sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 32 of 108 energy functions computed for a Conformer Ensemble can be used to determine whether a peptide-MHC complex is likely to be formed, to quantify the stability of a peptide-MHC complex (e.g. by quantifying the energy of the peptide-MHC complex using an energy function), and/or to quantify the binding of a peptide to an MHC molecule (e.g. by quantifying the energy of a peptide-MHC interface). Further, such metrics can be used to compare a plurality of peptides in relation to their likelihood of forming a peptide-MHC complex with a given set of one or more MHC molecules, their stability in complex with said MHC molecules, and/or their binding affinity with said MHC molecules. Thus, scores estimated from one or more energy functions computed for a plurality of Conformer Ensembles each comprising a given MHC allele and a candidate peptide with a candidate peptide sequence can be used to rank the candidate peptide sequences, such as e.g. to obtain a ranking indicative of one or more of: the likelihood of these candidate peptides forming a complex with the given MHC allele, the stability of complexes comprising the candidate peptides and the given MHC allele, and the binding affinity of the candidate peptides with the given MHC allele. Any energy function known in the art in may be used, including but not limited to e.g. the Rosetta energy function described in Cornell et al. (J Am Chem Soc.1996;118(9):2309–2309). In some instances, the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof, isolated from a subject (e.g., a patient). In some instances, the peptides can comprise mutated peptides (e.g., neoantigens expressed in tumor cells), or portions thereof, isolated from a subject (e.g., a patient). In some instances, the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium. In some instances, the peptides can comprise viral peptides, or portions, thereof expressed in a virus. [0118] In some instances, the method can further comprise developing a personalized anti- cancer vaccine, an anti-viral vaccine, an anti-bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof. For example, the method can comprise selecting one or more peptides of the candidate peptides based on the ranked order, for inclusion in a vaccine. [0119] FIG.3A depicts an example process for generating a peptide sequence embedding, in accordance with some embodiments. With reference to FIG.3A, a protein language model (PLM) 302 (corresponding to peptide embedding model 204 in FIG.2) can receive a peptide sequence and output a peptide sequence embedding 208 (corresponding to peptide embedded representation 208 in FIG.2). PLMs are machine-learning models (e.g., deep-learning models) sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 33 of 108 that can be based on natural language processing methods such as e.g. attention-based models (including but not limited to transformers) and can be trained on unlabeled sets of protein sequences. Protein language models can be trained to understand (i.e. using self-supervised learning, sometimes referred to as “pre-training”) and also optionally to predict one or more properties of peptides and/or proteins (sometimes referred to as fine tuning or task specific training or downstream task training) based on the amino acid sequence forming such a peptide and/or protein. In some instances, protein language models can infer a range of characteristics from amino acid sequences, including primary, secondary, tertiary, and quaternary structures of peptides and/or proteins as applicable. In the case of proteins, PLMs can be used e.g. to predict how proteins fold, what domains are present, where active sites are located, and how stable a protein is. In some cases, PLMs can also forecast protein-protein, protein-peptide, and protein-nucleic acid interactions, post-translational modifications, and the effects of mutations. In some cases, PLMs can identify localization signals within a cell, understand evolutionary relationships, and predict protein function. Likewise, PLMs can provide insights into the dynamic behavior of proteins and identify potential drug binding sites, valuable for drug discovery and understanding the molecular basis of diseases. In some embodiments, the PLM 302 can comprise an Evolutionary Scale Modeling (ESM model) or a variation of the ESM model. In some embodiments, the PLM 302 can comprise ProteinBERT, UniRep, or other suitable type of PLM (e.g. any protein language model known in the art). [0120] In some embodiments, the PLM 302 comprises a pretrained protein language model such as a pretrained ESM model. A pretrained PLM can refer to a PLM that has been trained in a self-supervised manner to learn a representation of protein sequences (i.e. learn a latent space from which protein sequences can be reconstructed, for example by training the model o reconstruct ground truth sequences from corrupted sequences), and has not been further trained or fine-tuned for a specific prediction task. In some embodiments, the input peptide sequence 126 can include a sequence of amino acid residues and the PLM 302 can be configured to obtain a plurality of embeddings (i.e., vector representations) by obtaining, for each amino acid residue, a corresponding embedding. The model can be further configured to obtain a single embedding 208 by aggregating the plurality of embeddings corresponding to the sequence of amino acid residues (e.g., by performing element-wise averaging). [0121] FIG. 3B depicts an example process for generating an MHC sequence (or other protein sequence) embedding, in accordance with some embodiments. With reference to FIG. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 34 of 108 3B, a PLM 304 (which may be the same model or a different model than PLM 302 illustrated in FIG.3A, and which corresponds to Protein Embedding Model 202 in FIG.2) can receive an MHC sequence 135 and output an MHC sequence embedding 206 (corresponding to MHC embedded representation 206 in FIG. 2). As discussed herein, a PLM can be trained using a large number of proteins and thus can encode useful information and context about the input sequence as represented by the sequence embedding. In some embodiments, the PLM 304 can comprise an Evolutionary Scale Modeling (ESM model) or a variation of the ESM model. In some embodiments, the PLM 304 can comprise ProteinBERT, UniRep, or the like. In some embodiments, the PLM 304 comprises a pretrained protein language model such as a pretrained ESM model. In embodiments, the PLM304 and the PLM 302 both comprise an ESM model, a ProteinBERT model, a UniRep model, or a variation of any of these. In embodiments, the PLM304 and the PLM 302 both comprise the same protein language model. This ensures that the embeddings for the peptide and protein come from the same model, enhancing the ability to learn associations between the peptide and protein. In some embodiments, the MHC sequence 135 can include a plurality of amino acids that compose a corresponding allele. The plurality of residues may be consecutive or non-consecutive. The MHC sequence 135 may be a pseudo-sequence. The pseudo-sequence may be, for each MHC allele (also referred to herein as “MHC molecule” or “HLA allele” or “HLA molecule”) a sequence of 34 amino acids at positions typically found within 4Å from peptides bound to the MHC allele. Pseudo-sequences of many MHC molecules are known in the art and described in e.g. Nielsen et al. PLoS One. 2007 Aug 29;2(8):e796. The PLM 304 can be configured to obtain a plurality of embeddings (i.e., vector representations) by obtaining, for each amino acid, a corresponding embedding. The model can be further configured to obtain a single embedding 206 by aggregating the plurality of embeddings corresponding to the sequence of amino acid residues (e.g., by performing element-wise averaging). [0122] FIG.4 provides a schematic illustration of using a trained machine learning model (e.g., Structural Constraint Prediction Model 141) to output pairwise distance data and dihedral angle data for a peptide-protein complex (e.g., a pMHC complex) based on input peptide and protein (e.g., MHC) embeddings. [0123] Peptide embedded representation 208 and MHC embedded representation 206 are provided as input to Structural Constraint Prediction Model 141, which can be configured to infer posterior distributions of Pairwise Distance Data 210 for peptide and protein amino acid sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 35 of 108 residues in the peptide-protein (e.g. pMHC) complex (e.g., distributions of predicted pairwise Cα – Cα distance data for at least one peptide amino acid residue and at least one protein amino acid residue in the peptide-protein complex), and Dihedral Angle Data 212 for the peptide amino acid residues in the pMHC complex (e.g., distributions of predicted dihedral angle data for at least one peptide amino acid residue in the peptide-protein complex). In embodiments, the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict one or more parameters characterizing respective distributions (e.g. posterior distributions) of distances between respective pairs of peptide and protein amino acid residues in the peptide- protein complex, such as e.g. at least one distribution of pairwise Cα – Cα distances between at least one peptide amino acid residue and at least one protein amino acid residue in the peptide- protein complex). In embodiments, the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict one or more parameters characterizing respective distributions (e.g. posterior distributions) of one or more dihedral angles for each of one or more peptide amino acid residues in the peptide-protein complex. In embodiments, the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict one or more parameters characterizing respective joint distributions (e.g. posterior distributions) of two dihedral angles (e.g., φ, ψ) for each of one or more peptide amino acid residues in the peptide- protein complex. In embodiments, the Structural Constraint Prediction Model 141 is configured (i.e. trained) to predict both: (i) one or more parameters characterizing respective distributions of distances between respective pairs of peptide and protein amino acid residues in the peptide-protein complex, and (ii) one or more parameters characterizing respective distributions (e.g. posterior distributions) of one or more dihedral angles for each of one or more peptide amino acid residues in the peptide-protein complex. In embodiments, instead of predicting distributions, the Structural Constraint Prediction Model 141 may be trained to predict individual values (e.g. samples of distributions of pairwise distances and/or dihedral angles), to which corresponding distributions can be fitted, thereby identifying said distributions. The predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter that can be characterized by, for example, a mean value, a standard deviation, or any other set of statistical parameters for characterizing a distribution. In other words, the predicted pairwise distance data and/or the predicted dihedral angle data can comprise one or more parameters that characterize each of one or more respective distributions for each pairwise distance and/or dihedral angle for which a prediction is made. References to predicted pairwise distances between amino acid residues sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 36 of 108 can refer to distances between selected heavy atoms in the backbone of the residues. The selected heavy atoms may be the alpha carbons (Cα) of the respective residues. Thus, a “pairwise distance” may be a distance between the alpha carbons of two amino acid residues. The two amino acid residues comprise an amino acid residue in the peptide and an amino acid residue in the protein. References to predicted dihedral angles can refer to prediction of the φ and/or ψ dihedral angle of a peptide residue (e.g. the peptide amino acid residue of a pair of residues comprising a peptide amino acid residue and a protein amino acid residue). [0124] In some instances, Structural Constraint Prediction Model 141 can comprise a trained neural network, as described in more detail elsewhere herein. In some instances, the trained machine learning model can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network. [0125] FIG. 5 provides a schematic illustration of using Pairwise Distance Data 210 and Dihedral Angle Data 212 for a peptide-protein complex (e.g., a peptide – MHC complex) output by a trained machine learning model (e.g., Structural Constraint Prediction Model 141) to generate Conformer Ensembles 148 for the peptide-protein complex (e.g., a pMHC complex). [0126] The predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 are used to select an initial Template Structure 502 for the pMHC complex from a set of example pMHC experimental structures, such as crystal structures (e.g., crystal structures available in a protein structure database such as the Protein Data Bank (PDB)) based on a maximum likelihood estimate (MLE) for the predicted posterior distributions for pairwise distance and dihedral angles. In embodiments, an initial Template Structure 502 for a peptide-protein complex is selected from a set of example peptide-protein experimental structures each comprising a peptide that has the same length (i.e. same number of amino acids) as the peptide in the peptide- protein (e.g. peptide-MHC) complex. In embodiments, selecting the initial Template Structure 502 from the set of example peptide-protein experimental structures comprises: (i) for each example peptide-protein experimental structure in the set, determining an average likelihood (which can be in practice calculated as a log-likelihood) of the distances and dihedral angles in the peptide-protein experimental structure using the predicted distributions for pairwise distances and dihedral angles for the peptide-protein complex; and (ii) selecting the example peptide-protein experimental structure that has the highest average likelihood. This results in selection of the template that is closest to the peptide-protein complex according to the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 37 of 108 predicted structural constraints. An average likelihood of the distances and dihedral angles in a peptide-protein experimental structure can be calculated as ^^^^^^^^^^^^^^^^(^^^^; ^^^^,^^^^2) = − log(^^^^) − + ^^^^^^^^^^^^^^^^^^^^ where x is a distance or dihedral angle in an experimental structu 2 2 re, and ^^^^,^^^^ are the predicted parameters of the distributions of distances and dihedral angles predicted by the model, here illustrated as Gaussian parameters. In embodiments, an initial Template Structure 502 for a peptide-protein complex is selected from a set of example peptide-protein experimental structures each a representative experimental structure for a cluster of peptide- protein experimental structures. In embodiments, an initial Template Structure 502 for a peptide-protein complex comprises a plurality of template structures. In some such embodiments, each template structure is selected from a set of example peptide-protein experimental structures each a representative experimental structure for a cluster of peptide- protein experimental structures. For example, a plurality of clusters of peptide-protein experimental structures each comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled may be obtained using a clustering algorithm and a predetermined structural similarity metric. A representative experimental structure may be identified for each cluster, thereby obtaining a plurality of representative experimental structures. One or more or each of the representative experimental structures may be used as an initial Template Structure 502. The representative experimental structures may be cluster centroids, such as e.g. structures that are closest to the center of the respective cluster using a predetermined distance metric. The clusters may be obtained by structural clustering of peptide backbones in a set of example peptide-protein experimental structures comprising peptides of a predetermined length (e.g. the length of the peptide for which an initial Template structure is sought). The clusters may be obtained using an agglomerative clustering method and a predetermined distance metric. The predetermined distance metric may be the cosine difference of the dihedral angles of the peptide. This can be calculated e.g. as ^^^^^^^^^^^^^^^^(^^^^,^^^^) = for two peptides A and B, where L is the length of the peptides. The initial Template Structure 502 can then be used as input to Ensemble Generation Engine 142. The term “experimental structures” refers to experimentally determined atomic 3D coordinates for a peptide-protein complex. Such structures may be determined using e.g. X-ray crystallography or nuclear magnetic resonance (NMR). The term “crystal structures” as used herein refers to atomic 3D coordinates for a peptide-protein complex experimentally determined by x-ray crystallography. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 38 of 108 [0127] The predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 also provide structural constraints on the allowable conformations of the pMHC complex, and are also provided as input to Ensemble Generation Engine 142. [0128] A plurality of compatible structures for the pMHC complex (e.g., Conformer Ensemble 148) may then be generated based on the initial Template Structure 502 and the constraints provided by the predicted Pairwise Distance Data 210 and Dihedral Angle Data 212 using Ensemble Generation Engine 142. Compatible structures for the pMHC complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance and dihedral angle data, and (ii) optimize the structure scoring function. Generation of a plurality of compatible structures may comprise, for example, inputting the initial pMHC template structure into a structural modeling software package (e.g., FlexPepDock (Rosetta, www.rosettacommons.org), described in Raveh et al. “Sub-angstrom modeling of complexes between flexible peptides and globular proteins”, Proteins, Vol. 78, Issue 9, July 2010, pp. 2029-2040) and further inputting the predicted pairwise distance data and dihedral angle data for the pMHC complex into the structural modeling software package. The structural modeling software package can be configured to iteratively sample (e.g., using a Markov Chain – Monte Carlo (MCMC) sampling algorithm) from a plurality of possible structures for the pMHC complex and evaluate each sampled structure using a structure scoring function (e.g., a proxy for a free energy calculation, as discussed in more detail elsewhere herein) to identify the plurality of compatible structures. Such a process is known is described in e.g. Raveh et al. Proteins, Vol.78, Issue 9, July 2010, pp.2029-2040. The predicted pairwise distance data and dihedral angle data can be used as constraints in a peptide docking protocol (also referred to as “refinement protocol” or “refinement procedure”). The predicted pairwise distance data and dihedral angle data can be used in an energy function used in the peptide docking protocol. For example, predicted pairwise distance data can be used as a harmonic constraint in the peptide docking protocol (e.g., it can be introduced as a harmonic function in an energy function using in the peptide docking protocol – see e.g. Alford et al. J. Chem. Theory Comput. 2017, 13, 6, 3031–3048). Predicted dihedral angle data can be used as a circular harmonic constraint in the peptide docking protocol (e.g., it can be introduced as a circular harmonic function in an energy function using in the peptide docking protocol – see e.g. Alford et al. J. Chem. Theory Comput. 2017, 13, 6, 3031–3048). In embodiments, generating a plurality of compatible structures for the peptide-protein complex comprises performing a pre- packing of the template to remove internal clashes in the protein and the peptide. In sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 39 of 108 embodiments, generating a plurality of compatible structures for the peptide-protein complex comprises optimizing the structure of the peptide backbone relative to the receptor protein (e.g., using a MCMC algorithm with Minimization approach), together with on-the-fly side-chain optimization. In embodiments, peptide backbone optimization is repeated a predetermined number of times (N), generating a number (N) of independent optimization trajectories. In embodiments, the peptide docking protocol uses a coarse grained model of the peptide and protein. In embodiments, the peptide backbone and peptide side-chains are optimized iteratively (e.g. using a MCMC algorithm), combined with rigid-body placement of the peptide with respect to the protein. In embodiments, an initial structure model (e.g. coarse grained model) is homology modelled using a template structure (i.e. a template peptide-protein complex structure that is experimentally derived). In embodiments, the initial structure model (e.g. coarse grained model) is homology modelled using a method comprising: aligning the sequence of the peptide and the peptide in the template structure, and using a threading protocol (such as e.g. as implemented in Rosetta) to generate a homology model using the alignment and the template structure coordinates. In embodiments, the homology model is subjected to a restricted refinement stage where only the residues of the peptide and the residues of the protein within a predetermined distance with a peptide residue (e.g. 3.5Å) are refined (e.g. in the Rosetta force field or corresponding energy function in any other structural modelling environment). In embodiments, the initial structure model is then used for generating the plurality of compatible structures using a structural modeling software package as described above. [0129] In some instances, Conformer Ensemble 148 may be used in downstream processing to, for example, predict features of the pMHC complex (e.g., interface features, surface complementarity features, etc.). Training a Classifier to Predict Peptide – MHC Binding [0130] FIG. 6 provides a non-limiting schematic illustration of a process for training a machine learning model to predict pairwise distance data and dihedral angles for a pMHC complex. As illustrated in the non-limiting example of FIG. 6, the model was configured to accept MHC embedded representation 608 and peptide embedded representations 606 for an MHC pseudo-sequence with a fixed length of 34 residues, 604, and a peptide sequence of variable length, 602, as input. The task performed by the trained model can be probabilistic regression over the structural properties of the pMHC complex. In particular, the model can be sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 40 of 108 trained to estimate the posterior distribution over the pairwise distances between the peptide and MHC amino acid residues and the dihedral angles of the peptide residues. For instance, if the parameterization of the posterior distribution consists of a factorized Gaussian and the peptide is, e.g., a 9-mer, the model outputs the Gaussian mean and standard deviation parameters for the pairwise distances between the MHC and peptide residues (34x9 distances in total) and likewise for the eight phi and psi dihedral angles. [0131] For this example, the model architecture consisted of a large pretrained model, such as an Evolutionary Scale Model (ESM) (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118; Verkul et al. (2022), “Language Models Generalize Beyond Natural Proteins", bioRxiv 2022-12), and included additional layers dedicated specifically to the probabilistic regression task. [0132] The training was carried out in two stages, as illustrated in FIG.6. In the first stage, the ESM portion of the model was fine-tuned for peptide-MHC binding classification, 610. A multi-layer perceptron (MLP) with a single (logit) output was appended to the concatenated ESM embeddings of the peptide and MHC. Specifically, the embeddings of the ESM portion of the model were mean-aggregated across the residue positions and processed by a MLP head (labeled ``Classifier'') with a single (logit) output. The MLP stage was included to take advantage of the wealth of publicly-available pMHC sequences annotated with binary binding labels. For instance, the model can be pre-trained on the labeled dataset of 2 million sequence pairs curated by Chu et al. (2022), “A Transformer-Based Model to Predict Peptide–HLA Class I Binding and Optimize Mutated Peptides for Vaccine Design", Nature Machine Intelligence 4(3):300-311. The classification task can be expected to provide the model, pretrained on general proteins, with an understanding of valid binding interfaces for a wide variety of peptides and MHCs. The binary cross-entropy / log loss function, which provides a measure of the difference between predicted binary outcomes and actual binary labels, was used for training. Here, N is the number of data points, the yi are the actual labels in the distribution of training data points (q(y)) and p(yi) is the probability that data point yi is in the positive class predicted by the model. It quantifies the dissimilarity between the two distributions, and facilitates model training by penalizing inaccurate predictions. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 41 of 108 [0133] In the second stage of training to predict pairwise distances and dihedral angles, the classification multilayer perceptron (MLP) was discarded and a new MLP was introduced, labelled “Regressor” 612. The embedding models 606 and 608 may be frozen at this stage. Although this is not a requirement, this may be advantageous as structure data used to train the model in the second stage is typically less abundant that the training data used in the first stage, and therefore can reduced the likelihood of model losing the generalizability learned in the first stage. Supervised training of the regression task requires crystal structure data for pMHC complexes (e.g., the atomic coordinates for each atom in each amino acid residue of the proteins), of which there are only a few hundred examples available in the Protein Data Bank (PDB). Each training instance was defined by data for a single MHC residue and a single peptide residue. Given the residue pair, the model extracts the slices corresponding to the residues from the ESM embeddings, concatenates them, and passes the resulting tensor through the regressor MLP. In other words, the forward call of the model extracts the slices corresponding to a given residue pair from the ESM embeddings (where a “slice” is a peptide residue or MHC residue embedding), concatenates them, and passes the resulting tensor through the new MLP. Every residue pair from the peptide and pseudo-MHC sequence is generated. The output of the model comprises pairwise distance data 614 and dihedral angle data 616, and varies depending on the parameterization used to characterize the predicted posterior distributions. The output of the model can comprise predicted distributions of c-alpha distances for all pairs of amino acid residues (e.g., a number of distributions equal to the number of residues of a pseudosequence of a MHC molecule (typically 34) multiplied by the length of the peptide). The output of the model can comprise joint distributions of phi and psi dihedral angles for a number of residues of the peptide corresponding to the length of the peptide minus 1 (since dihedral angles are between a first residue in Nter and a subsequent residue Cter of the first residue). For example, if the parameterization used consists of a joint Gaussian, then for a given residue pair the model may output the three mean values and a 3 x 3 covariance matrix that govern the trivariate Gaussian, where the diagonal elements of the 3 x 3 covariance matrix comprise the squares of the three standard deviation values, and the off diagonal elements are the covariance values. If the parameterization used consists of a factorized (diagonal) Gaussian, the output could simplify to, for example, three sets of mean and standard deviation values. Note that one is not restricted to using a Gaussian model to characterize the posterior distribution of predicted pairwise distance and dihedral angle data; other options include a Gaussian mixture model, (joint or factorized), von Mises distributions sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 42 of 108 for the angles, and a mixture thereof. The loss function used for training the new MLP 612 was the negative log-likelihood (NLL): evaluated at the distance and angle labels, where again, N is the number of data points and p(yi) is the probability that data point yi has a true label (i.e. the joint probability of the ground truth distance and dihedral angles in the predicted posterior distributions). The mean and standard deviation values can then be used as constraints in a peptide-protein docking protocol as explained elsewhere herein. In embodiments, the model is trained to output the predictive distribution over the pairwise distances between the peptide and HLA residues as well as the ϕ, ψ dihedral angles of the peptide residues. This can be denoted as: Where L is the peptide length, N is the protein sequence (e.g. pseudosequence) length ^^^^ ∈ the matrix of pairwise distances, are the peptide dihedral angles, and all terms are conditioned on the training set. In embodiments, the predictive distribution is parameterized as a Gaussian mixture with K components (although as explained above other distributions can be used). This can be denoted as: Where N(.|μ,σ) denotes the Gaussian density with mean and variance parameters μ,σ, such that ^^^^ (^^^^) ^^^^^^^^ ,^^^^^^ 2 (^^^^) ^^^^^^ ∈ ℝ+ are the parameters of the Gaussian component k governing each distance element and likewise for ∈ ℝ governing each angle element. The ∈ [0,1] are the mixture weights that sum to 1 across k∈ {1, … and likewise for . The value of K (number of components) may be set independently for every predictive distribution, or may be set to be the same for a plurality of (e.g. all) predictive distributions. The value(s) of K may be set as a hyperparameter, e.g. using cross-validation. In embodiments, the model is trained to predict a set of distributional parameters for an input protein sequence embedding and an input peptide sequence embedding, the distributional parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 43 of 108 characterizing the distributions of pairwise distances between pairs of a peptide amino acid residue and a protein amino acid residue (^^^^�^^^^^^^^^^^^�^^^^^^^^^^^^^^^^, ^^^^^^^^^^^^^^^^^^^^�), and the joint distribution of dihedral angles for peptide amino acid residues. Using the above notations, and in embodiments in which the predictive distribution for the distances and dihedral angles are parameterized as Gaussian mixtures with K components, this can be denoted as ^^^^ ≔ In embodiments, the structural constraint prediction model (e.g. new MLP 612) is trained using maximum likelihood estimation. In embodiments, the structural constraint prediction model (e.g. new MLP 612) is trained using a loss function defined as: where w is a hyperparameter, L is the length of the peptide, N is the length of the protein sequence (e.g., pseudosequence), and the subscript θ refers to the parameters of the structural constraint prediction model (e.g. new MLP 612). The hyperparameter w may be set to a predetermined value. In embodiments, the structural constraint prediction model (e.g. new MLP 612) is trained using a loss function that weighs training losses corresponding to peptides of a first predetermined length for which more training data is available less than training losses corresponding to peptides of a second predetermined length for which less training data is available. [0134] In some instances, for example when the crystal structure data for one or more peptide-protein complexes (e.g., one or more pMHC complexes) is incomplete (e.g., dihedral angle and/or pairwise distance data is missing for one or more amino acid residues), a masking step may be implemented so that the missing information is excluded from the training data used to train and/or optimize the model. In addition, for cases where the phi angle is available but the psi angle is not, the joint distribution over phi, psi for an amino acid residue can be decomposed as follows: Similarly, for cases where the psi angle is available but the phi angle is not, the joint distribution over phi, psi for an amino acid residue can be decomposed as follows: sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 44 of 108 [0135] The distribution, θ, can be optimized using either the left-hand expression or the right-hand expression. Evaluating the left-hand expression requires measurements of both phi and psi. In the right-hand (decomposed) expression, evaluating the first term (conditional distributions), ^^^^�^^^^ ( ^^^^ ) |^^^^ ( ^^^^ ) � or ^^^^�^^^^ ( ^^^^ ) |^^^^ ( ^^^^ ) �, also requires measurements of both phi and psi, but evaluating the second term (marginal distributions), ^^^^(^^^^ ( ^^^^ ) ) or ^^^^(^^^^ ( ^^^^ ) ), only requires measurement of phi or psi, respectively. This allows one to still train the model on amino acid residues for which only a measurement of phi or psi is available. The decomposition described above is straightforward for Gaussian distributions, but for other distributions other algorithms, such as expectation maximization (EM) may be required. In other words, when only one of the φj and ψj labels are observed for a residue in the training data, the residue may not be discarded from the training data, and instead the distribution of ^^^^�^^^^^^^^,^^^^^^^^�^^^^^^^^^^^^^^^^, ^^^^^^^^^^^^^^^^^^^^� may be factorized depending on the availability of the labels as p(φjj)p(ψj) when the label for ψj is available, and p(ψjj)p(φj) when the label for φj is available. Because the marginal term p(ϕj) admits evaluation even when only the ϕj label is observed, and likewise for ψj , the model can still learn from residues missing a single angle label. In embodiments, residues in the training data for which labels for both dihedral angles are observed are assigned (e.g. randomly) to evaluate either of the conditionals p(ϕj | ψj) or p(ψj | ϕj ) and those with only one angle label are assigned to evaluate the marginal terms (p(ϕj) or p(ψj)). Generating a pMHC Conformational Ensemble [0136] FIG. 7 provides a non-limiting example that illustrates the process for how predicted distributions of pairwise distance data and dihedral angle data for a pMHC complex can be converted to structural constraints, and how pMHC conformational ensembles can then be generated based on such structural constraints. [0137] An MHC sequence or pseudosequence (i.e., the ordered set of amino acids of the MHC molecule that typically contacts a peptide), comprising amino acid residues that typically contact a bound peptide, e.g., 34 amino acid residues (here illustrated as YFAMYGEKV…(SEQ ID NO:1)), and a peptide sequence, such as e.g., a 9-mer peptide sequence (here illustrated as LLFGYPVYV (SEQ ID NO:2)) are input into a trained machine learning model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG. 1) to predict pairwise distance data, here illustrated as mean (702, left panel) and standard deviation (704, middle panel) values characterizing the distribution of pairwise distance data, for all pairs of peptide and MHC pseudosequence amino acid residues, and to predict dihedral angle data, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 45 of 108 here illustrated as mean and standard deviation values (706, right panel), for peptide dihedral angles for all peptide residues in the complex. [0138] As illustrated in FIG.8, these predictions can then be used as constraints within a structural modeling software, such as the FlexPepDock (Rosetta, www.rosettacommons.org) software package, to generate conformational ensembles for the pMHC complex, together with a template structure model selected based on the predicted pairwise distance data and dihedral angle data. The selected template structure (model 802, upper left panel) was based on the predicted pairwise distance and dihedral angle constraints as applied to pMHC crystal structure PDBID: 4NO5, and was used to create the starting model (model 804, upper right panel) for conformer generation using FlexPepDock. Conformer generation was then performed using the iterative process described elsewhere herein. An overlay of 200 conformer models is shown in the lower panel, model 806. Using a Structural Constraint Prediction Model to Generate a Conformer Ensemble [0139] FIG. 9 provides a non-limiting example of a process 900 (e.g., a computer- implemented method) for generating a plurality of compatible structures for a peptide-protein complex. Process 900 can be implemented using the Prediction System 100 described in FIG. 1. For example, process 200 can be implemented using the Sequence Analyzer 108 and one or more of the Machine Learning Models 132 in FIG.1. [0140] At step 902 in FIG.9, peptide sequence data (for at least one peptide) and protein sequence data (for at least one protein) can be input into a trained machine learning model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1) to determine: (i) predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in a peptide-protein complex (e.g. each predicted pairwise distance data comprising one or more predicted distances and/or a one or more parameters characterizing a predicted distribution of distances between a pair of amino acid residues comprising a peptide residue and a protein amino acid residue in the peptide-protein complex); and (ii) predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex (e.g. each predicted dihedral angle data comprising, for a residue of the peptide in the peptide-protein complex: one or more predicted angles for a ψ angle, one or more predicted angles for a φ angle, one or more parameters characterizing a predicted distribution of ψ angle, one or more parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 46 of 108 characterizing a predicted distribution of φ angle, and/or one or more parameters characterizing a predicted joint distribution of ψ and φ angle). [0141] In some instances, the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter (e.g. each distance and each dihedral angle – or pair of dihedral angles - for which a prediction is made) that can be characterized by, for example, a mean value, a standard deviation, and an uncertainty (or any other set of statistical parameters for characterizing a distribution). In other words, the predicted pairwise distance data can comprise one or more predicted parameters that together characterize a distribution of pairwise distance between a pair of amino acid residues comprising at least one peptide residue and at least one protein amino acid residue in the peptide-protein complex. Similarly, the predicted dihedral angle data can comprise one or more parameters that together characterize a distribution of ψ angle, a distribution of φ angle and/or a joint distribution of ψ and φ angles for at least one amino acid residue in the peptide that forms part of the peptide-protein complex. In some instances, the trained machine learning model can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution). In embodiments, the predicted pairwise distance data is predicted for all pairs comprising a peptide amino acid residue and a protein pseudosequence residue, where a protein pseudosequence residue is a residue of the protein that is believed to be involved in the interaction with peptides when forming a peptide-protein complex. In embodiments, the predicted dihedral angle data is predicted for all amino acid residues of the peptide apart from the last (C-ter-) amino acid residue. In embodiments, the predicted dihedral angle data is predicted for all amino acid residues of the peptide core (also referred to as minimal epitope, in the context of peptide-MHC complexes). [0142] In some instances, the input can comprise peptide sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or more than 50 peptides, or portions thereof. In some instances, peptide sequence data can comprise sequence data for peptides that are at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues in length. In some instances, the peptides can comprise normal peptides (e.g., peptides expressed in normal cells), or portions thereof, isolated from a subject (e.g., a patient), derived from nucleic acid sequences isolated from a subject, or obtained from one or more sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 47 of 108 genome, transcriptome and/or proteome reference dataset (e.g. a genome, transcriptome or proteome database). In some instances, the peptides can comprise mutated peptides (e.g., neoantigens expressed or likely to be expressed in tumor cells), or portions thereof, isolated from a subject (e.g., a patient) or derived from nucleic acid sequences isolated from a subject. In some instances, the peptides can comprise bacterial peptides, or portions thereof, expressed by a bacterium. In some instances, the peptides can comprise viral peptides, or portions, thereof expressed in a virus. In some instances, the peptides can comprise peptides from one or more pathogens, such as e.g. bacteria, viruses, parasites (e.g. plasmodium), etc. [0143] In some instances, the peptide sequence data for the at least one peptide can comprise a peptide embedded representation (or peptide embedding, e.g., a machine-friendly vector representation of the peptide sequence) for the at least one peptide generated using a trained peptide embedding machine learning model. The vector representation of the peptide sequence can encode the peptide sequence in terms of the statistical dependencies and patterns present in the peptide sequences used to train the model. The peptide sequence data can comprise a peptide embedded representation from a trained protein language model, such as e.g. peptide embedding model 204. The peptide embedding model 204 may have been trained in a self-supervised manner as part of a protein language model. The peptide embedding model 204 may have been further trained (also referred to as “fine-tuned”) as part of a machine learning model configured to predict a property of a peptide-protein complex comprising the peptide. Further, multiple consecutive fine-tuning steps may have been implemented comprising further training the peptide embedding model 204 as part of a machine learning model configured to predict a different property of the peptide-protein complex comprising the peptide. The properties of the peptide-protein complex may be selected from: functional properties and structural property. The machine learning model configured to predict a property of the peptide-protein complex may comprise a peptide embedding model, a protein embedding model, and a task specific prediction head. The task specific prediction head may be a classification head or a regression head. A predicted functional property may be binding between the peptide and the protein to form the peptide-protein complex. The binding between the peptide and the protein may be predicted as a binary property (binding / non-binding) using a classification head, or as a continuous property (e.g. predicted binding affinity) using a regression head. A predicted functional property may be presentation of the peptide by the protein on the surface of cells, where the protein is a MHC molecule. The presentation of a peptide by a MHC molecule may be predicted as a binary property (e.g. presented / not sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 48 of 108 presented), for example using eluted ligand data. A predicted functional property may be stability of the peptide-protein complex. The stability of a peptide-protein complex may be predicted as a binary (e.g. stable/unstable) or continuous (e.g. measured complex half-life) property. A predicted structural property may be prediction of pairwise distance data and/or dihedral angle data as described herein. A task specific prediction head may be a multilayer perceptron. As the skilled person understands, predicting a binary property can comprise predicting a probability that the peptide-protein complex comprising the peptide belons to a class of a pair of class (e.g. predicting a probability that the peptide binds to the protein). Peptide-MHC binding affinity data, binding data and presentation data is available in publicly available databases and datasets such as e.g. Chu et al. (2022) as mentioned above, and the Immune Epitope Database & Tools (IEDB, www.iedb.org/home_v3.php). [0144] In some instances, the input can comprise protein sequence data for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000, or more than 1000 proteins, or portions thereof. In some instances, protein sequence data can comprise sequences or pseudosequences of at least 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or at least 200 amino acid residues in length. In some instances, the protein sequence data for at least one protein can comprise a protein embedded representation (or protein embedding, e.g., a machine-friendly vector representation of the protein sequence) for the at least one protein generated using a trained protein embedding machine learning model. The vector representation of the protein sequence can encode the protein sequence in terms of the statistical dependencies and patterns present in the protein sequences used to train the model. The protein sequence data can comprise a protein embedded representation from a trained protein language model, such as e.g. protein embedding model 202. The protein embedding model 202 may have been trained in a self-supervised manner as part of a protein language model. The protein embedding model 202 may have been further trained (also referred to as “fine-tuned”) as part of a machine learning model configured to predict a property of a peptide-protein complex comprising the protein, as explained above in relation to the peptide embedding model 204. [0145] In some instances, the trained peptide embedding machine learning model and the trained protein embedding machine learning model can be different models (e.g., Peptide Embedding Model 204 and Protein Embedding Model 202, respectively, as illustrated in FIG. 2). In some instances, the trained peptide embedding machine learning model and the trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 49 of 108 protein embedding machine learning model can be the same model (e.g., the Embedding Generation Engine 140 illustrated in FIG. 1 and FIG. 2). Example methods for performing peptide and/or protein sequence embedding are described in, e.g., Ofer et al. (2021), “The Language of Proteins: NLP, Machine Learning & Protein Sequences”, Computational and Structural Biotechnology Journal 19:1750-1758. In embodiments, the trained peptide embedding machine learning model and the trained protein embedding machine learning model can be based on the same model obtained from a protein language model (i.e. a protein/peptide embedding model that has been trained as part of a protein language model in a self-supervised manner). Two instances of the protein/peptide embedding model (which may respectively be referred to as pretrained protein embedding model and pretrained peptide embedding model) may have been included in a machine learning model that is trained to predict one or more properties of a peptide-protein complex, in order to obtain the trained protein embedding model and trained peptide embedding model (i.e. further training / fine tuning the pretrained protein embedding model and peptide language model). [0146] In some instances, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise large language models (LLMs), e.g., deep learning algorithms often based on transformer networks that can recognize, summarize, translate, predict, and generate content using very large data sets (e.g., very large peptide and/or protein sequence data sets). Large language models are described in more detail in, for example, Li et al. (2021), “BioSeq-BLM: A Platform for Analyzing DNA, RNA and Protein Sequences Based on Biological Language Models”, Nucleic Acids Res.49, e129–e129, and Naveed et al. (2023), “A Comprehensive Overview of Large Language Models”, arxiv.org/pdf/2307.06435.pdf. A protein language model can refer to a large language model that has been trained in a self-supervised manner using data from one or more proteomic databases. Thus, a protein language model may be a large language model that has been trained to model the “language” of proteins using large amounts of unlabeled protein sequence data. [0147] In some instances, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise an Evolutionary Scale Model (ESM), e.g., a transformer-based LLM trained on unlabeled data for over 250 million distinct protein sequences that generates a vector representation for an input sequence. ESM has been shown to be able to encode biochemical properties of amino acids in the sequence, the biological variation and remote homology of the sequence, and secondary structure and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 50 of 108 tertiary contacts of the sequence (see, e.g., Rives et al. (2021), “Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences”, PNAS 118(15)e2016239118). The trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can comprise fine-tuned versions of an ESM model, where fine tuning refers to further training of the models as part of a machine learning model trained to predict one or more properties of a peptide-protein complex. [0148] In some instances, the trained machine learning model configured to determine predicted pairwise distance data and dihedral angle data (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1) can comprise a trained neural network, as described in more detail elsewhere herein. In some instances, the trained machine learning model can comprise, for example, a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network. [0149] In some instances, the machine learning model configured to determined predicted pairwise distance data and dihedral angle data (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1) can be trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide-protein combinations; training a classifier (or regressor) to predict peptide-protein binding (or other functional properties of a peptide-protein complex) for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and appending a regressor head to at least a portion of the trained classifier (or regressor) and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes. The classifier can be a machine learning model comprising a protein embedding model, a peptide embedding model and a classification head. A classification head in this context can be a machine learning model configured to take as inputs a peptide embedded representation from the peptide embedding model, and a protein embedded representation from the protein embedding model, and provide as output a prediction indicative of a class of a plurality of classes that the peptide-protein complex belongs to. The plurality of classes may comprise a class of peptide-protein pairs that bind to each other to form a peptide-protein complex and a class of peptide-protein pairs that do not bind to each other to form a peptide-protein complex. The regressor can be a machine sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 51 of 108 learning model comprising a protein embedding model, a peptide embedding model and a regression head. A regression head in this context can be a machine learning model configured to take as inputs a peptide embedded representation from the peptide embedding model, and a protein embedded representation from the protein embedding model, and provide as output a prediction indicative of binding affinity between the protein and the peptide (e.g. a predicted binding affinity). Appending a regressor head to at least a portion of the trained classifier/regressor and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes can comprise removing the previous classifier/regression head and replacing it with a new regressor head, then training the resulting machine learning model to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes. Crystal structure data may refer to pairwise distances and dihedral angles obtained from experimentally determined structural coordinates of peptide-protein complexes. Pairwise distances and dihedral angles can be calculated from structural coordinates obtained from publicly available structure databases such as e.g. the PDB. [0150] In some instances, the peptide sequence data for the plurality of peptides and/or the protein sequence data for the plurality of proteins used for training can each independently comprise sequence data for at least 2 x 104, 4 x 104, 6 x 104, 8 x 104, 1 x 105, 2 x 105, 3 x 105, 4 x 105, 5 x 105, 6 x 105, 7 x 105, 8 x 105, 9 x 105, 1 x 106, 2 x 106, 3 x 106, 4 x 106, 5 x 106, 6 x 106, 7 x 106, 8 x 106, 9 x 106, 1 x 107, or more than 1 x 107 peptide or protein sequences. [0151] In some instances, the peptide-protein binding data for a plurality of peptide-protein combinations used for training can comprise binding data (e.g., binary (yes/no) binding data and/or continuous valued binding affinity data) for at least 2 x 104, 4 x 104, 6 x 104, 8 x 104, 1 x 105, 2 x 105, 3 x 105, 4 x 105, 5 x 105, 6 x 105, 7 x 105, 8 x 105, 9 x 105, 1 x 106, 2 x 106, 3 x 106, 4 x 106, 5 x 106, 6 x 106, 7 x 106, 8 x 106, 9 x 106, 1 x 107, or more than 1 x 107 peptide- protein combinations. [0152] In some instances, the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide-protein complex can be presented as a separate training instance during training of the regressor head. This provides more training data in cases where the number of peptide-protein crystal structures available is limited. In sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 52 of 108 other words, the training data used to train the machine learning model configured to determined predicted pairwise distance data and dihedral angle data (e.g., the Structural Constraint Prediction Model 141) can comprise, for each of a plurality of training pairs of amino acid residues comprising a peptide amino acid residue and a protein amino acid residue in a training peptide-protein complex, experimentally determined values of a pairwise distance between the amino acid residues in the pair and one or both dihedral angles associated with the peptide residue in the pair. These pairwise distances and dihedral angles may be referred to herein as “ground truth” distances / dihedral angles (or simply “ground truth” or “label”). The training data can therefore comprise ground truth experimentally determined pairwise distances and dihedral angles for each of a plurality of pairs of amino acid residues derived from a plurality of training experimentally determined peptide-protein complexes structures. [0153] In some instances, the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor. Processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex can comprise obtaining pairwise distance data and dihedral angle data for the peptide-protein complex using the trained machine learning model (e.g. a trained machine learning model trained using training data without augmented data), the pairwise distance data comprising a statistical measure of variability in the pairwise distance (e.g. as one of a set of parameters characterizing a predicted distribution of the pairwise distance between a pair of amino acid residues) and the dihedral angle data comprising a statistical measure of variability in the dihedral angle(s) (e.g. as one of a set of parameters sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 53 of 108 characterizing a predicted distribution of one or both of the dihedral angles of the peptide amino acid residue in a pair), as described herein. Applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex can comprise using the predicted measures of uncertainty to generate an ensemble of conformers using a structural modeling software and a template structure that is an experimentally determined structure for the particular peptide-protein complex. The resulting conformers may then be used to derive an augmented set of “ground truth” pairwise distances and dihedral angles. In other words, the augmented pairwise distance data and dihedral angle data can be obtained by applying the methods described herein to generate a conformer ensemble for a peptide-complex for which an experimentally determined structure is available, and using the resulting conformers to derive additional ground truth pairwise distances and dihedral angles for the peptide-protein complex. This additional ground truth pairwise distances and dihedral angles can capture reasonably expected levels of flexibility in the peptide-protein complex that is not captured in experimentally determined crystal structures. In embodiments, the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by selecting one or more peptide-protein pairs that are known to bind to each other to form a complex, and predicting a structure for complexes corresponding to each of said peptide-protein pairs using a structure prediction machine learning algorithm. For example, the machine learning algorithm may be selected from the AlphaFold family of models (e.g. AlphaFold2 or fine-tuned versions thereof, such as e.g. AF-FT which is fine-tuned for peptide-HLA binding predictions, see Motmaen et al. Proc. Nat. Ac. Sci.120(9), 2216697120, 2023). [0154] In some instances, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can be integrated with, and trained simultaneously with, the trained machine learning model. [0155] At step 904 in FIG.9, an initial structure (also referred to herein as “template”) for the peptide-protein complex can be identified based on the predicted pairwise distance data and the predicted dihedral angle data. [0156] Identifying the initial structure (also referred to herein as “initial template structure” or simply “template structure”) for the peptide-protein complex based on the predicted pairwise sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 54 of 108 distance data and predicted dihedral angle data can comprise, for example, obtaining a plurality of crystal structures for peptide-protein complexes from a protein structure database; and selecting a crystal structure from the plurality of crystal structures based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted distance data and the predicted dihedral angle data. Similarity between pairwise distance data and dihedral angle data determined for the crystal structure and the predicted distance data and the predicted dihedral angle data can be quantified using a likelihood estimate of the pairwise distance data and dihedral angle data determined for the crystal structure under distributions characterized by the predicted distance data and the predicted dihedral angle data. Alternatively, a crystal structure may be selected from the plurality of crystal structures using a structural similarity metric determined for the crystal structure and a maximum likelihood estimate of pairwise distances and dihedral angles according to the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG.1). In embodiments, the initial structure is selected based on sequence similarity between the peptide in the peptide-protein complex and the peptides of the same length in the crystal structures for peptide-protein complexes. Sequence similarity may be calculated using any scoring known in the art, such as e.g. BLOSUM62. The plurality of crystal structures obtained may be crystal structures comprising the same protein as the peptide-protein complex. For example, one or more crystal structures may be available for a particular MHC molecule. The plurality of crystal structures obtained may be crystal structures comprising the same protein as the peptide-protein complex, but not necessarily the same peptide. The plurality of crystal structures obtained may be crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled. The plurality of crystal structures may be crystal structures obtained by clustering crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled and identifying a representative structure for each cluster. In embodiment, the initial structure (initial template structure) may comprise a plurality of structures. For peptides up to 10 amino acids, a single template structure may be used. For peptides of 11 or more amino acids, a plurality of template structures may be used. The plurality of template structures may comprise crystal structures obtained by clustering crystal structures comprising a peptide of the same length as the peptide in the peptide-protein complex to be modelled and identifying a representative structure for each cluster. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 55 of 108 [0157] In some instances, the plurality of crystal structures for peptide-protein complexes obtained from a protein structure database can comprise at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, or more than 300 crystal structures. [0158] In some instances, the protein structure database can comprise, e.g., the Protein Data Bank (PDB) (Research Collaboratory for Structural Bioinformatics; rcsb.org). [0159] At step 906 in FIG.9, a plurality of compatible structures can be generated for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data. [0160] In some instances, each compatible structure for the peptide-protein complex can conform, for example, with a posterior distribution of the predicted pairwise distance data and with a posterior distribution of the predicted dihedral angle data output by the trained machine learning model. In some instances, each compatible structure for the peptide-protein complex can conform with the sequence data for the at least one peptide and the at least one protein and with the predicted pairwise distance data and the predicted dihedral angle data. In embodiments, the plurality of compatible structure for the peptide-protein complex are used to obtain a prediction of at least one auxiliary feature. [0161] In some instances, generating the plurality of compatible structures for the peptide- protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data can comprise, for example, inputting the initial structure for the peptide-protein complex into a structural modeling software package (e.g., the Ensemble Generation Engine 142 illustrated in FIG.1 and FIG.2); inputting the peptide sequence data, the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, where compatible structures for the peptide-protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function. Inputting the peptide sequence data is optional and instead the initial structure may have already been adapted to include the peptide sequence, such as e.g. using homology modelling and an initial structure comprising a peptide that has a different sequence from the peptide in the peptide-protein complex to be modeled. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 56 of 108 [0162] In some instances, the iterative sampling can comprise 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1,000 iterations. The number of iterations may be set by one or more stopping criteria of e.g. a MCMC algorithm. Thus, the number of iterations may not be set a priori and may differ between instances. [0163] In some instances, the plurality of possible structures can comprise 2, 4, 6, 8, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1,000 possible structures. In embodiments, the plurality of possible structures comprises between 10 and 1000 structures, between 50 and 1000 structures, between 50 and 500 structures, between 100 and 500 structures, between 100 and 400 structures, or about 200 structures. [0164] In some instances, the iterative sampling from the plurality of possible structures for the peptide-protein complex can be performed using a Markov Chain – Monte Carlo (MCMC) sampling algorithm (see, e.g., Ravenzwaaij et al. (2018) “A Simple Introduction to Markov Chain Monte–Carlo Sampling”, Psychon Bull Rev 25:143–154, and Jones et al. (2022), “Markov Chain Monte Carlo in Practice”, Annu. Rev. Stat. Appl. 9:557–578). In some instances, the Markov Chain – Monte Carlo (MCMC) sampling algorithm can comprise, for example, a Metropolis – Hastings algorithm, a Hamiltonian Monte Carlo algorithm, or a Langevin Monte Carlo algorithm. In embodiments, the iterative sampling from the plurality of possible structures for the peptide-protein complex can be performed using a MCMC algorithm with a Metropolis criterion, e.g. as described in Raveh et al. Proteins, Vol. 78, Issue 9, July 2010, pp.2029-2040 [0165] In some instances, the structure scoring function can comprise, e.g., a free energy calculation for the peptide-protein complex (e.g., a binding free energy comprising the free energy difference between the bound complex and unbound peptide and protein, or an overall interaction free energy comprising the free energy difference between the bound complex and isolated peptide and protein). For example, the structural scoring function may be as described in Raveh et al. Proteins, Vol. 78, Issue 9, July 2010, pp. 2029-2040, or the Rosetta scoring function (see Leman et al., Nature Methods 17(7), 665-680 (2020)). In some instances, the structure scoring function can comprise a proxy for a free energy calculation. In some instances, the structure scoring function can comprise a proxy for a free energy calculation of sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 57 of 108 peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof. [0166] In some instances, the structure scoring function (e.g., a proxy for a free energy calculation) can be based on one or more of ab initio quantum mechanical calculations, density functional theory (DFT) calculations, semi-empirical calculations, molecular mechanics force field calculations, statistical potential calculations, neural potential (or neural force field) calculations that predict the forces acting on individual atoms (e.g., performed using models that recapitulate the results of high-level quantum mechanics calculations at higher speed and/or providing more robust evaluation of bad structures), or a function parameterized by machine learning models trained on structural data. For example, in some cases, the structure scoring function may be a physics-based energy function that determines a total energy of the peptide-protein complex, including one or more of an electrostatic energy, covalent bonding energy, Van der Waals energy, and/or the like. It should be appreciated that different structure scoring functions may be associated with different levels of accuracy and computational complexity. For instance, a more accurate structure scoring function, such as an ab initio quantum mechanics-based structure scoring function, may impose greater computational overhead than a less accurate structure scoring function, such as a molecular mechanics force fields-based structure scoring function. Accordingly, in some cases, more than one structure scoring functions may be applied as part of identifying compatible structures for the peptide- protein complex, where compatible structures are identified as those structures having, e.g., a lowest structure score (e.g., a lowest total energy). [0167] In some instances, the structural modeling software package can comprise, e.g., AlphaFold (DeepMind, London, UK), Amber (ambermd.org), ESMFold (Meta AI, New York, NY), FlexPepDock (RosettaCommons, rosettacommons.org), HADDOCK (High Ambiguity Driven biomolecular DOCKing; bonvinlab.org), Molecular Operating Environment (MOE; Chemical Computing Group, Montreal, CA), OpenMM (openmm.org), Schrödinger Glide (New York, NY), and AutoDock (autodock.scripps.edu). In embodiments, the structural modeling software package comprises FlexPepDock. A structural modeling software may also be referred to as a “physics-based engine”, because it determines structures that are energetically plausible according to predetermined energetic functions. Any structural modeling algorithm that can obtain a predicted structure using a template sequence, and that sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 58 of 108 can accommodate dihedral angle and/or pairwise distances constraints (e.g. within an energy scoring function) can be used in the context of the present disclosure. [0168] In some instances, the process (or computer-implemented method) illustrated in FIG. 9 can further comprise predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex. In some instances, the at least one auxiliary feature can comprise, for example, a prediction of binding of the at least one peptide to at least one other protein or protein complex, and/or a prediction of at least one interface feature for the peptide-protein complex (such as e.g. information identifying complementary interface features between the peptide-protein complex and another protein or protein complex, and/or information identifying complementary surface fingerprint features between the peptide-protein complex and another protein or protein complex). [0169] In some instances, the process (or computer-implemented method) illustrated in FIG.9 can further comprise displaying at least a subset of the predicted pairwise distance data and/or the predicted dihedral angle data using any of a variety of techniques for displaying three-dimensional tensor data known to those of skill in the art, e.g., 3D plots, 2D plots of a specified slice of the tensor data, etc. In some instances, the method can comprise displaying at least a subset of the predicted pairwise distance data determined by the trained machine learning model as, e.g., a two-dimensional heatmap comprising a plot of the pairwise distance between an alpha carbon (Cα) of a peptide amino acid residue and an alpha carbon (Cα) of a protein amino acid residue as a function of peptide amino acid residue position and protein amino acid residue position. The plotted pairwise distances can comprise maximum likelihood estimates of the predicted pairwise distances, for example based on predicted parameters of respective distributions of pairwise distances. The plotted pairwise distances can comprise predicted averages of predicted distributions of the pairwise distances. In other words, the predicted pairwise distance data can comprise parameters characterizing the distribution of each pairwise distance for which pairwise distance data is predicted. These parameters can be used to obtain a maximum likelihood estimate of each of these pairwise distances. The maximum likelihood estimate can be e.g. a predicted average (i.e. mean). This may be the case for example when parameters of Gaussian distributions are predicted (or any other distributions where the mean of the distribution is the highest likelihood value of the distribution). The plotted pairwise distances can comprise multiple respective two-dimensional plots, each plot sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 59 of 108 displaying the predicted values of a parameter of a predicted distribution for pairwise distances. For example, the plotted pairwise distances can comprise a first two-dimensional plot displaying the means of predicted distributions of pairwise distances, and a second two- dimensional plot displaying the standard deviation or other statistical metric of variability of predicted distributions of pairwise distances. In some instances, the method can comprise displaying at least a subset of the predicted dihedral angle data determined by the trained machine learning model as, e.g., a two-dimensional scatter plot of phi and psi dihedral angles for at least a subset of the peptide amino acid residues. A two-dimensional scatter plot of phi and psi dihedral angles can comprise an indicator of the means or maximum likelihood estimates of the predicted distributions of phi and psi dihedral angles for each of a plurality of peptide amino acid residues. A two-dimensional scatter plot of phi and psi dihedral angles can comprise an indicator of a statistical metric of variability (e.g. standard deviation) of the predicted distributions of phi and psi dihedral angles for each of a plurality of peptide amino acid residues. Instead or in addition to a two-dimensional scatter plot of phi and psi dihedral angles, one or both of a scatter plot of phi angles and a scatter plot of psi dihedral angles may be plotted. Each scatter plot can comprise an indicator of the means or maximum likelihood estimates of the predicted distributions of phi or psi dihedral angles for each of a plurality of peptide amino acid residues, and/or an indicator of a statistical metric of variability for the predicted distributions of phi or psi dihedral angles for each of the plurality of peptide amino acid residues. [0170] In some instances, the process (or computer-implemented method) depicted in FIG. 9 can further comprise iterating over a plurality of candidate peptides (e.g., at least 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 candidate peptide sequences or portions thereof) to identify at least one peptide that has a maximum likelihood for formation of the peptide-protein complex. [0171] In some instances, the process (or computer-implemented method) illustrated in FIG.9 can further comprise predicting features of the peptide-protein complex (e.g., interface features, surface complementarity features, etc.) based on the plurality of compatible structures for the peptide-protein complex. In some instances, predicting binding of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex can comprise, for example, identifying complementary interface features between the peptide- protein complex and another protein or protein complex; identifying complementary surface sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 60 of 108 fingerprint features between the peptide-protein complex and another protein or protein complex; identifying latent representations of the peptide-protein complex and the other protein or protein complex; and providing the complementary interface features, the complementary surface fingerprint features, and the latent representations of the peptide-protein complex and the other protein or protein complex as input to a second trained machine learning model configured to output a prediction for binding of the peptide-protein complex and the other protein or protein complex. A latent representation of a peptide-protein complex may be any representation in a space in which the peptide-protein complex is projected. This can be used e.g. to identify similarity between complexes (e.g. identify complexes that are more similar to each other) in the latent space. In some instances, predicting binding of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex can comprise, for example, predicting a structure of the plurality of compatible structures in complex with another protein. [0172] In some instances, complementary interface features can comprise, for example, pairwise lists of amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, total hydrogen bonding energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of hydrogen bonding energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, total van der Waals interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of van der Waals interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, or any combination thereof. [0173] In some instances, complementary surface fingerprint features can comprise, for example, steric hinderance-based complementary shape features, complementary electrostatic features, total electrostatic interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of electrostatic interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 61 of 108 total cation-pi interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, pairwise lists of cation-pi interaction energies for amino acid residues in the peptide-protein complex that are in contact with amino acid residues in the other protein or protein complex, complementary polarity features, or any combination thereof. [0174] In some instances, the complementary surface fingerprint features between the peptide-protein complex and the other protein or protein complex can be predicted by a third trained machine learning model based on the latent representations of the peptide-protein complex and the other protein or protein complex. [0175] In some instances, the second trained machine learning model and/or the third trained machine learning model can comprise, for example, a trained neural network or deep learning model. Training a Classifier to Predict Pairwise Distance Data and Dihedral Angle Data [0176] FIG.10 provides a non-limiting example of a flowchart for a process 1000 (e.g., a computer-implemented method) for training a machine learning model (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG. 1) to predict distributions of pairwise distance data and dihedral angle data for a peptide-protein complex. Process 1000 can be implemented using the prediction system 100 described in FIG.1. For example, process 700 can be implemented using the Sequence Analyzer 108 and one or more of the Machine Learning Model(s) 132 in FIG. 1. All the features described above in relation to e.g. FIG. 9 may apply equally to the methods of FIG.10. [0177] At step 1002 in FIG. 10, peptide sequence data for a plurality of peptides can be received (e.g., by one or more processors of the Prediction System 100 described in FIG.1). [0178] In some instances, the peptide sequence data for the plurality of peptides can comprise sequence data for at least 2 x 104, 4 x 104, 6 x 104, 8 x 104, 1 x 105, 2 x 105, 3 x 105, 4 x 105, 5 x 105, 6 x 105, 7 x 105, 8 x 105, 9 x 105, 1 x 106, 2 x 106, 3 x 106, 4 x 106, 5 x 106, 6 x 106, 7 x 106, 8 x 106, 9 x 106, 1 x 107, or more than 1 x 107 peptide sequences. [0179] In some instances, the peptide sequence data for the plurality of peptides can comprise peptide embedded representation data (e.g., machine-friendly vector representations of the peptide sequences) for the plurality of peptides that are generated using a trained peptide sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 62 of 108 embedding machine learning model (e.g., the Embedding Engine(s) 140 illustrated in FIG.1), as described elsewhere herein. [0180] At step 1004 in FIG. 10, protein sequence data for a plurality of proteins can be received. [0181] In some instances, the protein sequence data for the plurality of peptides can comprise sequence data for at least 2 x 104, 4 x 104, 6 x 104, 8 x 104, 1 x 105, 2 x 105, 3 x 105, 4 x 105, 5 x 105, 6 x 105, 7 x 105, 8 x 105, 9 x 105, 1 x 106, 2 x 106, 3 x 106, 4 x 106, 5 x 106, 6 x 106, 7 x 106, 8 x 106, 9 x 106, 1 x 107, or more than 1 x 107 protein sequences. [0182] In some instances, the protein sequence data for the plurality of proteins can comprise protein embedded representation data (e.g., machine-friendly vector representations of the protein sequences) for the plurality of proteins that are generated using a trained protein embedding machine learning model(e.g., the Embedding Engine(s) 140 illustrated in FIG.1), as described elsewhere herein. In some instances, the trained protein embedding machine learning model can be the same model as the trained peptide embedding machine learning model. [0183] At step 1006 in FIG. 10, peptide-protein binding affinity data for a plurality of peptide-protein combinations can be received. [0184] In some instances, the peptide-protein binding data for a plurality of peptide-protein combinations can comprise binding data (e.g., binary (yes/no) binding data and/or continuous valued binding affinity data) for at least 2 x 104, 4 x 104, 6 x 104, 8 x 104, 1 x 105, 2 x 105, 3 x 105, 4 x 105, 5 x 105, 6 x 105, 7 x 105, 8 x 105, 9 x 105, 1 x 106, 2 x 106, 3 x 106, 4 x 106, 5 x 106, 6 x 106, 7 x 106, 8 x 106, 9 x 106, 1 x 107, or more than 1 x 107 peptide-protein combinations. [0185] At step 1008 in FIG. 10, a classifier or regressor (e.g., the Structural Constraint Prediction Model 141 illustrated in FIG. 1) can be trained to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data. In other words, a classifier or regressor can be trained to predict peptide-protein binding data for input pairs of peptide sequence data and protein sequence data using training data comprising, for each of a plurality of training peptide-protein complexes: peptide sequence data, protein sequence data, and corresponding ground truth (i.e. measured or otherwise known) peptide-protein binding data. The ground truth peptide-protein binding data can comprise a binary classification label (e.g. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 63 of 108 binding/non-binding) used to train a classifier, and/or a measure of binding affinity which can be used to train a regressor or from which a binary label can be obtained using a predetermined binding affinity threshold (e.g. for training a classifier). [0186] Any of a variety of machine learning approaches and algorithms (where a machine learning model, as referred to herein, comprises a trained machine learning algorithm) can be used in implementing the disclosed methods. For example, the machine learning model can comprise a supervised learning model (i.e., a model trained using labeled sets of training data), an unsupervised learning model (i.e., a model trained using unlabeled sets of training data), a semi-supervised learning model (i.e., a model trained using a combination of labeled and unlabeled training data), a deep learning model (e.g., a neural network comprising many layers of coupled “nodes” that can be trained in a supervised, unsupervised, or semi-supervised manner), or any combination thereof. For example, the machine learning model can comprise one or more machine learning models (such as e.g. embedding models) that have been pretrained in a self-supervised manner, then further trained in a supervised or semi-supervised manner (e.g. as part of respective machine learning models trained to perform respective prediction tasks using corresponding training data including at least some labelled data). The embedding models, classifier head, and/or regression head may each be deep learning models. Thus, the machine learning model comprising any one or more of these can also be a deep learning model. A deep learning model may be an artificial neural network (ANN) comprising a plurality of hidden layers. A deep learning model may comprise one or more attention-based layers, one or more fully connected layers, one or more convolutional layers, and combinations thereof. In some instances, one or more machine learning models (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 machine learning models), or an ensemble thereof, can be utilized to implement the disclosed methods. In some instances, the disclosed methods can be implemented using a continuous learning approach, where the machine learning model can be periodically or continuously updated based on new training data provided by, e.g., a single local operational system, a plurality of local operational systems, or a plurality of geographically-distributed operational systems. [0187] Examples of machine learning algorithms that can be employed include, but are not limited to, neural networks, feedforward neural networks (also known as multilayer perceptrons), recurrent neural networks, convolutional neural networks, deep neural networks, sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 64 of 108 deep feedforward neural networks, deep recurrent neural networks, deep convolutional neural networks, attention-based neural networks, graph neural networks, or any combination thereof. [0188] Neural networks (NNs) generally comprise an interconnected group of nodes organized into multiple layers of nodes. For example, the NN architecture can comprise at least an input layer, one or more hidden layers, and an output layer. The NN can comprise any total number of layers (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100), and any number of hidden layers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100), where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to a predicted output value or set of output values. Each layer of the neural network comprises a number of nodes (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 75100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or more than 10,000 nodes). A node receives input data (e.g., peptide sequence data, protein sequence data, concatenations thereof, or other types of input data) that comes either directly from one or more input data nodes or from the output of one or more nodes in previous layers, and performs a specific operation, e.g., a summation operation. In some cases, a connection from an input to a node is associated with a weight (or weighting factor). In some cases, the node can, for example, sum up the products of all pairs of inputs, Xi, and associated weights, Wi. In some cases, the weighted sum is offset with a bias, b. In some cases, the output of a node can be gated using a threshold or activation function, f, where f can be a linear or non-linear function. The activation function can be, for example, a rectified linear unit (ReLU) activation function or other function such as a saturating hyperbolic tangent, identity, binary step, logistic, arcTan, softsign, parameteric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, Sinusoid, Sine, Gaussian, or sigmoid function, or any combination thereof. [0189] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network, can be “taught” or “learned” in a training phase using one or more sets of training data (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 sets of training data). For example, the parameters can be trained using the input data from a training data set and a gradient descent or backward propagation method so that the output value(s) (e.g., predicted pairwise distance data and/or dihedral angle data) that the ANN computes are consistent with the examples included in the training data set. The adjustable parameters of the model can be obtained using a back propagation neural network training process that may or may not be sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 65 of 108 performed using the same hardware as that used for implementing the trained model and processing input peptide sequence and/or protein sequence data. [0190] Accessing the training data set can include, for example, retrieving the training data set from a local or remote storage, loading the training data set, and/or requesting (and receiving) part or all of the training data set from one or more data stores (e.g., a cloud data storage, a server system, or some other data source). [0191] The training data set can be randomly parsed, shuffled, and/or divided to train the classifier model. The machine-learning model can be trained using a static or dynamic learning rate. A dynamic learning rate can be produced using, for example, learning-rate annealing (e.g., using stepwise annealing or cosine annealing). Training can be performed using, for example, a classification loss function and/or a regression loss function. A loss function can be based on, for example, mean square error, median square error, mean absolute error, median absolute error, an entropy-based error, a cross entropy error, and/or a binary cross entropy error. Validation data (e.g., a separated subset of the training data set used to train the machine- learning model can be used to assess the performance of the machine-learning model as it is being trained. Training can be terminated if and/or when the target performance is obtained, and/or the maximum number of training iterations have been completed. [0192] At step 1010 in FIG.10, a regressor head can be appended to at least a portion of the trained classifier and the model comprising the regressor head and the portion of the trained classifier can be trained to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes. Thus, a machine learning model comprising a regressor head and at least a portion of the trained classifier (e.g. comprising a trained protein embedding model and a trained peptide embedding model) can be trained using training data comprising, for each of a plurality of training peptide- protein complexes: protein sequence data, peptide sequence data, and corresponding ground truth structure data (e.g. experimentally determined crystal structure data) or pairwise distance data and dihedral angle data derived therefrom. The portion of the trained classifier can include a protein embedding model and a peptide embedding model as described herein, trained as part of the classifier model further comprising a classifier head, but excluding the classifier head after such training. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 66 of 108 [0193] As noted elsewhere herein, the predicted pairwise distance data and predicted dihedral angle data can comprise a predicted posterior distribution for each parameter (e.g. each pairwise distance and each dihedral angle for which a prediction is made) that can be characterized by, for example, a mean value, a standard deviation, and an uncertainty (or any other set of statistical parameters that characterize a distribution). In some instances, the trained machine learning model can be further configured to determine the mean value, the standard deviation, and/or an uncertainty in the predicted pairwise distance data and/or the predicted dihedral angle data (or any other set of statistical parameters for characterizing a distribution). [0194] In some instances, the set of example peptide-protein complexes may comprise at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, or more than 300 peptide-protein complexes. [0195] As noted elsewhere herein, in some instances the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide- protein complex can be presented as a separate training instance during training of the regressor head, thereby providing more training data in cases where the number of peptide-protein crystal structures available is limited. [0196] In some instances, the regressor can be further trained on augmented pairwise distance data and dihedral angle data, where the augmented pairwise distance data and dihedral angle data can be derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 67 of 108 [0197] In some instances, the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model can be integrated with, and trained simultaneously with, the trained machine learning model. [0198] As noted above, the training data set can be randomly parsed, shuffled, and/or divided to train the regressor head. A loss function used as part of model training can comprise an error term (e.g., mean squared error or median squared error) and/or an entropy term (e.g., cross entropy or binary cross entropy). [0199] A static or non-static learning rate can be used during training of the regressor head. For example, learning rate annealing (e.g., using stepwise annealing or cosine annealing) can be used to reduce the learning rate over the course of performing multiple training iterations. In some instances, validation-data assessment can be used to potentially terminate training early (e.g., upon determining that a performance target has been met). Example – Assessment of Predicted Structural Models [0200] The immune surveillance mechanism in humans is driven by the CD8+ T-cells recognizing cytosolic peptides presented by the class I human leukocyte antigen (HLA) molecules through their receptors. Understanding and predicting the specificity of the T-cell receptors (TCRs) to their cognate antigens have been particularly challenging and the process is further confounded by the TCR promiscuity. Recognition of peptide-HLA (pHLA) molecules by TCRs has also been attributed to the peptide conformational flexibility and dynamics. This has been particularly true for complexes that undergo structurally silent mutations where co-crystal structures do not display noticeable conformational changes. Despite these structurally silent mutations, TCRs that recognize such peptides exhibit a high degree of specificity. For example, the wild- type (WT) and mutant (MT) PIK3CA peptides are both presented by HLA- A*03:01 however, only the MT peptide is recognized by TCR4. To study such systems, traditional molecular dynamics (MD) simulations over long timescales can be employed. While MD is suitable for smaller number of instances, they can become a limiting factor where large-scale data generation is needed especially to address TCR-pHLA binding and/or signaling prediction problems. Recent work leverages machine learning (ML) models to predict protein conformations. For example, AlphaFold2 (AF2) (Jumper, J.M et al. PLOS Computational Biology 14(12), 1006342 (2018)) predicts static structures as seen in the PDB. Developed specificall for the pHLA space is AlphaFold-FineTune (AF-FT) (Motmaen, A, et al., PNAS, 120(9), 2216697120 (2023)), which fine-tune AF2 for pHLA binding sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 68 of 108 prediction but does not focus on modeling the conformational flexiility of peptides. Despite considerable progress in ML-facilitated conformational sampling, dedicated ML approaches for generating peptide conformers at the pHLA interface remain underexplored. Evaluation of predicted conformer distributions can also benefit from more detailed analyses, especially in comparison to MD (molecular dynamics) distributions. The present inventors developed a method that uses ML to guide a physics-based engine, such as e.g. Rosetta (Leman et al. Nature Methods 17(7), 665–680 (2020)), to model conformational flexibility of peptides at the pHLA- I interface. The ML component of the method is trained on pHLA sequences and structures and predicts structural descriptors such as distances between the HLA and the peptide residues and the peptide backbone dihedral angles. The predicted structural descriptors are applied as constraints in a peptide coking protocol in the physics engine (here exemplified as FlexPepDock (FPD) protocol in Rosetta; Raveh et al. Proteins: Structure, Function, and Bioinformatics 78(9), 2029-2040, (2010)) to sample peptide conformations. The inventors tested the method on a benchmark set of pHLA complexes consisting of 32 different alleles spread across peptide lengths of 8-13 amino acids. They demonstrated that the approach (termed “p-flex”) can recover 78% and 67% of the peptide backbones sampled by the molecular dynamics simulations for peptides of lengths 8-10 and 11-13 amino acids, respectively. To the best of their knowledge, this is the first method to capture the flexibility of peptides up to 13 amino acids in length in a matter of minutes without relying on time- intensive MD simulations. [0201] A non-limited example of a method of the present disclosure and results thereof is presented here. An ML model was designed to predict the distribution over structural descriptors, namely the Cα-Cα distances between HLA and peptide residues and the peptide backbone dihedral angles (ϕ and ψ). First, a classifier was trained on sequence data with pHLA binding labels to learn about peptide and HLA associations. Then the regressor head from the common embedding shared with the classifier was trained on structural descriptors from pHLA complexes in the Protein Data Bank (PDB). The regressor outputs parameters governing the distributions of distances between peptide and HLA residues and joint distributions of ϕ and ψ dihedral angles of the peptide residues. At inference time, the model takes in the peptide and HLA sequences and outputs the parameters governing the predictive distribution over the structural descriptors. The predicted descriptors were used select a template structure from the PDB that is similar to a given pHLA complex, for initializing the FlexPepDock protocol (FPD). They are also passed into FPD during refinement as HARMONIC and CIRCULAR sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 69 of 108 HARMONIC constraints. Whereas the model learns these descriptors primarily from static co- crystal structures in the training set, the inventors believe that a well- regularized probabilistic model is capable of extracting shared information— about not only the static structure but also about the dynamics—from different complexes with similar structural properties. The example results below to support this claim; our model, trained on static structures, can recover the static structural descriptors of unseen test complexes as well as approximate their dynamical distribution as estimated by MD simulations. [0202] Training data – 3D structures: The 3D structure data for training and validation of the ML method was curated from PDB. All the PDB structures with keywords “HLA” or “HLA”, restricted to a source organism “Homo sapiens” with a resolution of 10Å or lower were downloaded (as of Feb 2024). Every chain in the downloaded PDB structure was trimmed to a length of 180 amino acids. The chains with lengths (i) lower than 15 amino acids were marked as peptides, (ii) in the template neighborhood of 180 residues and typed by sequence alignment to the HLA sequences in IMGT/IPD database were marked as HLAs. If there were multiple assemblies available in a given PDB, a single pHLA pair was selected by identifying the HLA chain and the corresponding peptide chain by measuring the distance between the second residue of the peptide and 65th residue in the HLA (arbitrarily selected the residues based on manual examination of multiple crystal structures in the PDB). A finalized trimmed snapshot of the pHLA structure was saved in the structure dataset. The typed alleles, together with amino acid sequences of chains in the trimmed structures, were saved in a separate metadata file. Structures of pHLAs found in both unbound and bound (to TCR or co-receptors) conformational states were kept, but the bound partners were removed before training. The final dataset contained a total of 651 noisy (referred to as noisy because they include structures with missing densities, bound and unbound pHLAs) structures which were subdivided into train, validation and test (blind) sets. The PDB contains many more pHLA structures with peptide lengths up to 10 (592 complexes) compared to those with lengths 11-13 (51 complexes) amino acids. To overcome the data sparsity for longer length peptides, the inventors augmented the training set with structures predicted by AF-FT (Motmaen et al. PNAS 120(9), 2216697120 (2023)), an extension of AF2 (Jumper et al. PLOS Computational Biology 14(12), 1006342 (2018)) fine-tuned on pHLA structures from the PDB. Specifically, to generate augmented structure data for peptides of lengths 10-14 amino acids, the inventors randomly sampled 200 binders from the pHLA-I sequences obtained from Chu., et.al. (Nature Machine Intelligence 4(3):300-311), and modeled them with Alphafold-finetune using the default weights. Each sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 70 of 108 model was subsequently refined in the Rosetta force field three times using the Relax protocol. The structural descriptors (distances and dihedrals) from a set of 3000 models were then used to train the regressor component of the method. [0203] Training data - Labeled pHLA-I sequences: The curated data set of pHLA-I sequences with binary binding labels was obtained from Chu., et.al. (Nature Machine Intelligence 4(3):300-311). This data set consists of (i) binder pHLA sequences and (ii) non- binder pairs that were generated for every binder pair by choosing random peptides of same lengths from the peptidomes in IEDB for same alleles. In this labeled data set, there are 2.2M paired sequences distributed across 112 unique HLAs and each HLA has approximately 17% of binder peptide samples. [0204] ML method to predict structural descriptors: The input to the constraint prediction model consisted of a HLA pseudosequence sHLA with a fixed length of 34 residues and a peptide sequence spep of variable length ranging from 8 to 15, inclusive. The model was trained to output the predictive distribution over the pairwise distances between the peptide and HLA residues as well as the ϕ, ψ dihedral angles of the peptide residues (^^^^�^^^^,^^^^,^^^^�^^^^^^^^^^^^^^^^, ^^^^^^^^^^^^^^^^^^^^� as described herein, denoted P(D, φ, ψ| spep, sHLA) in this context, where in this case the matrix of pairwise distances ^^^^ ∈ ℝ34^^^^^^^^ + (as HLA pseudosequences of 34 amino acid lengths are used). This is referred to as “joint” modelling of the dihedral angles and is what was used unless indicated otherwise. Indeed, the distribution factorizes among the residues, but ϕ and ψ are modeled jointly for a given residue. The predictive distribution was parametrized as a Gaussian mixture with K components as explained above. The component numbers (K) for the Gaussian mixtures can in principle vary between each predictive distribution but in the present examples the inventors set them to be equal for simplicity and tuned K via cross validation. The final output of the model is the full set of distributional parameters ^^^^ ≔ as described above. The ESM2 Transformer model described in Rives et al. (2021), PNAS 118(15)e2016239118 available at https://github.com/facebookresearch/esm was used to initialize embedding models for both peptide sequences and MHC sequences. Two copies of this model were included to encode respectively a peptide sequence and an MHC pseudosequence. The ESM2 embeddings were mean-aggregated across the residue positions and a small MLP head with a single (logit) output was appended. The resulting classifier model was trained to predict peptide-MHC binding using the data in Chu et al. (2022) and a binary cross-entropy / log-loss function. The binary sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 71 of 108 classification task is designed to provide the model, pretrained on general proteins, with an understanding of valid binding interfaces for a wide variety of peptides and HLAs. The classification head was then removed and replaced with a new regressor MLP. Supervising this regression task requires 3D structures, of which only a few hundred are available in the PDB. The inventors thus defined each training instance as a single HLA residue and a single peptide residue. Given the residue pair, the model extracts the slices corresponding to the residues from the ESM embeddings, concatenate them, and pass the resulting tensor through the regressor MLP. The training is carried out via maximum likelihood estimation (MLE) and the negative log of the distribution as the loss function (ℒ(^^^^) as explained above, where w is a tuned hyperparameter that trades off between the distance and dihedral losses and the subscript θ represents the neural network (MLP) weight parameters). The dataset for the regressor MLP training includes structures with missing distance or angle labels. The inventors accounted for missingness during training so that, when a distance label for Dij was missing for an HLA residue i and peptide residue j, or when both ϕj and ψj labels were missing for a peptide residue j, the corresponding model prediction did not contribute to the loss. When one of the both ϕj and ψj labels is observed, however, that residue was not discarded and the joint distribution in the predictive distribution for the dihedral angles was factorized as p(ϕj | ψj)*p(ψj) or p(ψj | ϕj)*p(ϕj) the depending on the availability of the labels. Because the marginal term p(ϕj) admits evaluation even when only the ϕj label is observed, and likewise for ψj, the model can still learn from residues missing a single angle label. During training, residues with both ϕ, ψ labels were randomly assigned to evaluate either of the conditionals p(ϕj | ψj ) or p(ψj | ϕj ) and those with only one angle label were assigned to evaluate the marginal terms. The best hyperparameter configuration was determined via 5-fold cross validation. The inventors drew 1,000 predictive samples for the structural descriptors. These samples were then consolidated across 5-fold cross-validation to evaluate the performance of the method. [0205] Molecular dynamics simulations: MD simulations were used to assess the results of the model. MD simulations were initiated from the curated data set of PDB structures described above. Simulations were performed on the bound peptide-HLA complex. The HLA was truncated (residues 1-180) to increase the speed of the simulation. Hydrogens were stripped and sidechains pKas were predicted using PROPKA at pH 7.4, and protonation states were assigned using PDB2PQR with the Amber FF19SB force field. The protein was solvated in an octahedral box of OPC water with the box extending 10Å from the protein and neutralizing NaCl ions. The structures were relaxed with 2000 steps of conjugate-gradient sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 72 of 108 energy minimization, followed by a 0.5 ns heating to 300K in NPT, the protein was restrained using a harmonic restraint potential with a force of 10 kcal mol-1 Å-2. An additional 2000 steps of conjugate-gradient energy minimization were performed without any restraints. Following minimization, the systems were equilibrated using NPT where pressure was maintained at 1 atm and the thermostat at 300K for 0.5 ns with a restraint of 10 kcal mol-1 Å-2 applied to the protein. The system is then equilibrated for 1 ns with a restraint force constant of 1 kcal mol-1 Å-2, and a final equilibration is run at 300K with no restraints. All restraints were removed for the production stage. The production simulations were carried out using NPT conditions and were run using the hydrogen mass repartition option of Amber, allowing a time step of 4 fs. Langevin dynamics are used to maintain the temperature at 300K with a collision frequency of 3 ps-1. Periodic boundary conditions are applied in all directions, the particle mesh Ewald (PME) method is used to calculate the electrostatic interactions with a real-space cutoff of 10 Å, and hydrogen lengths were constrained with the SHAKE algorithm. Simulations were run for 300ns in duplicate adding up to 600ns. The inventors analyzed MD simulations to gain insights into the convergence of the simulations and to understand the conformational ensembles sampled within the MD dataset. The analyses are performed using the cpptraj, unless otherwise specified Root mean squared deviation (RMSD) values were calculated for the backbone atoms (CA,C,O,N) using the ’rms’ CPPTRAJ command after superimposition of the truncated HLA. The root mean squared fluctuation (RMSF) of the peptide was calculated using the CPPTRAJ ’atomicfluct command. The dihedral angles were calculated for the peptide using the ’dihedral’ command. [0206] Sampling of peptide backbones using FPD: The conformer ensembles were sampled using the refinement procedure within the FPD protocol in the Rosetta Software Suite. The refinement procedure used a coarse grained model of the interacting partners. Here, the peptide backbone and its side chains were optimized iteratively combined with rigid-body placement of the peptide with respect to the HLA. The inventors modified the sampling temperature to 0.2 (the default value is set to 0.8) which is used in the Metropolis criterion, to allow for sampling of diverse peptide backbones during simulations. Coarse grained models were homology modeled using structural templates selected based on either (i) sequence similarity of the target to the template peptide sequence calculated using BLOSUM62 or (ii) the predicted distance and dihedral constraints. For the latter, the selection was restricted to structural templates in the PDB with the peptide lengths matching the target peptide length without missing densities. For each template, the average Gaussian log-likelihood of the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 73 of 108 distances and dihedrals was computed as described above. Rosetta was used to homology model a target sequence of interest, using as input a target peptide sequence, the HLA allele and a template structure of the complex. The aligning the sequences of the target and the template were aligned using ClustalOmega. Subsequently, the alignments together with template coordinates were utilized by the Rosetta’s threading protocol to generate a homology model. The model was further subjected to a restricted refinement stage where only residues of the peptide and the residues in the HLA contacted by the peptide (within 3.5 ˚A) are refined in the Rosetta force field. [0207] Structural clustering of peptide backbones: To cluster peptide backbones, the matrix of pairwise cosine difference of the dihedral angles of the peptides was calculated as described above. Clustering of the backbones was performed using agglomerative clustering. The cluster centers were then identified and the closest samples to these clusters centers were selected as centroids for downstream tasks. [0208] Assessment of predicted structural models was performed based on, for example, (i) calculation of the average dihedral distance of modeled peptide backbones relative to “ground truth” values based on crystal structure data for pMHC complexes, (ii) calculation of the root mean square deviation (RMSD) of dihedral angle distances relative to “ground truth” values to check for appropriate placement of the peptide with respect to the MHC in the pMHC complex, and (iii) determination of the fraction of contacts recovered in the model relative to the crystal structure contacts. Dihedral distance (or dihedral angle distance) is a metric that quantifies the distance between two dihedral angles. For two dihedral angles, θ1 and θ2, of the same type and at the same residue position of two different structures, the distance between the two dihedral angles is defined to be: (see, e.g., North et al. (2011), “A New Clustering of Antibody CDR Loop Conformations”, J. Mol Biol. 406(2):228-256). The D-score was used to capture an alignment free distance between two peptide backbones. The D-score was computed for ϕ, ψ and ω dihedral angles of L−1 residues for a peptide length L as ^^^^ − ^^^^^^^^^^^^^^^^^^^^(^^^^,^^^^) ^^^^ 2(1 − cos (^^^^^^^^ − ^^^^^^ ^ ^^ ^^^)) where A and B are peptides of length L. The peptide dihedral angles of the predicted complexes and those of the MD complexes represent samples drawn from two distributions. The D-score above quantifies the difference between these distributions by treating the ϕ, ψ, and ω angles as independent — the angular distances are computed for each position and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 74 of 108 aggregated in a simple sum. Position-wise metrics like the D-score does not account for the correlations in the three dihedral angles across the residue positions, however. The root mean square deviation (RMSD) between two structures (such as e.g. a predicted structure and a “ground truth”) structure, is calculated as ^^^^^^^^^^^^^^^^ where n is the number of pairs of equivalent atoms in the two aligned / superimposed structures, di is the distance (typically in Å, typically Euclidian distance) between the two atoms in pair i. The RMSD can be calculated for any type and/or subset of atoms in a protein or protein-peptide complex. Unless indicated otherwise, the RMSD values provided herein refer to RMSDs calculated using all backbone heavy atoms (C, N, Cα, and O) of both a peptide and a protein in a peptide-protein complex for which two structures are compared. The side chain contact recovery was estimated between the predicted models and either MD models or crystal structures. The number of contacts between the HLA and peptide chains of the reference structure (MD model or crystal structure) were computed by first identifying the residues on the HLAs that are within a distance of 15Å between their Cα and the peptide residues (or identifying nearby HLA residues to reduce to the search space for subsequent step). Any-atom to any-atom contacts were then captured from the nearby HLA residues and the peptide residues if they are within a threshold of 3.1Å. The fraction of contacts recovered by the models were then estimated by counting the overlapping contacts between the reference structures and the models of interest. [0209] FIG. 11A provides a non-limiting example of data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted Φ and Ψ dihedral angles relative to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis – identified by the PDB identifier for the corresponding structure) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues), where the Φ and Ψ dihedral angles were treated independently during training and inference. Results are shown as distributions of RMSD values for a plurality of conformer structures predicted as described herein. In this early example, approximately 25% of the predicted structures exhibit an RMSD of greater than 2.5 Å. FIG.11B provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted Φ and Ψ dihedral angles compared to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues), where the Φ and Ψ dihedral angles were treated either independently or jointly for a given sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 75 of 108 amino acid residue during training and inference. Results are shown as distributions of RMSD values for a plurality of conformer structures predicted as described herein. Joint treatment of the dihedral angles resulted in more accurate (reduced RMSD) predictions. FIG.11C provides a non-limiting example of comparison data for the root mean square deviation (RMSD) of peptide-protein complexes with predicted Φ and Ψ dihedral angles compared to corresponding “ground truth” crystal structures (plotted on the y-axis) for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length, where one data set (red single dots) was generated using a prior protein structure prediction model (AlphaFold – FineTune (AF-FT)), and the other data set (green distribution of dots) was generated using the methods disclosed herein (with joint prediction of the dihedral angles). Results are shown as distributions of RMSD values for a plurality of conformer structures predicted as described herein. These preliminary results indicate that AF-FT predicts the crystal structures accurately but does not reflect the dynamics of the pMHC complex as it exists in solution. FIG. 11D provides a non-limiting example of preliminary data for predicted average dihedral distance plotted on the y-axis for a set of benchmark pMHC complexes (plotted on the x-axis) comprising peptides of different length (8, 9, 10, 11, and 13 amino acid residues). The results are consistent with the expectation of greater flexibility of longer peptides. FIG.11E provides a non-limiting example of preliminary data for the recovery of crystal structure contacts (i.e., fraction of amino acid residues in the crystal structure that have distances of less than 3.1 Å that also have distances of less than 3.1 Å in the predicted structures) for a set of benchmark pMHC complexes (comprising peptides of length 8, 9, 10, 11, and 13 amino acid residues), where predicted structures are obtained using the disclosed structural prediction methods. The distribution of the fraction of contacts recovered over the predicted conformer structures (y- axis) is plotted for the benchmark pMHC complexes (x-axis). The data demonstrates a high degree of correlation between the predicted contacts and the crystal structure contacts. [0210] To evaluate the structural descriptors predicted by the method in this example, the inventors partitioned the training and test sets based on sequence and structural similarity. For the sequence-based splits, the train and test sets were composed of 80% and 20% of the total 3D structure data from PDB respectively of peptide lengths up to and including 11 AA. The train set was further reshuffled into a 5-fold cross-validation sets consisting of approx. 80% and 20% of the 80% structure data into train and validation sets. The test set and the validation sets consisted of bins of data points covering all the alleles and peptide lengths that had more than one co-crystal structure. For peptides of lengths greater than 11 AA, the inventors trained sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 76 of 108 on the data from PDB and structural models generated using Alphafold-Finetune. The selection of validation sets followed the similar strategy as described above for training the models with structure data from PDB only. To generate structure-based data splits, the inventors binned the structures by peptide lengths (up to 10 AA). For each peptide length, they performed clustering of the peptide backbones using a D-score radius of 5.9. All the singleton clusters were moved to the test set. From the remaining clusters, each randomly selected member was placed in a validation set selected at random until a count of 60 in the validation set was reached. The sequence-based split probes the model’s ability to recapitulate the structural descriptors for peptides with sequences unseen in the training set. The structure-based split stress-tests the model’s ability to generalize to structural clusters not represented in the training set, particularly for longer (11-13 amino acids) peptides. Consisting of PDB structures that are diverse in sequence and structure spaces as well as distinct from the training set, these test sets ensure reliable model evaluation. Fig.12 shows exemplary results demonstrating the method’s ability to recover dihedral angles and distances. Fig. 12A shows Boxplots highlighting the average of predicted distances (sampled from the aggregated Gaussian distance distribution) for the peptide positions grouped by peptide length. A box shows a distance distribution constructed from the minimum (min from 34 distances of a peptide residue to the HLA residues) predicted mean distances averaged across all the modeled samples (200 FPD models across 5 folds) for individual peptide position for every test instance. X-axis in panels A and B are peptide positions. Fig. 12A shows that the method produced accurate predictions of distances for both data splits for shorter peptides. The model was able to discriminate between anchor and non-anchor residue positions. The predicted distances underscored that the anchors have generally tighter distributions and lower distances — take, for instance, positions 1 and 8 for 8-mers, positions 1, 2 and 9 for 9-mers and positions 1, 2, 9 and 10 for 10-mers (Figure 12A; the distributions showing mean of minimum distances of a peptide residues to any of the HLA residues across sampled structures from FPD from the test set). The non-anchor residues have higher distances. Next, to analyze the coverage of estimated dihedral density, the inventors overlaid native dihedral angles in the test sets with predicted joint densities of ψ as shown by y-axes and ϕ as shown by x-axes (or Ramachandran plot). Fig. 12B shows Ramachandran plots showing dihedral recovery of shorter peptides with length of 8, 9, or 10 (shown in ’Short’), or longer peptides with length of 11, 12, 13 (shown in ’Long’). Predicted dihedral density is shown by light grey contour (the allowed region covering 98% residues) or a dark grey contour (the favored region 98% residues). Scatters represent true dihedral angles sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 77 of 108 in the test set. The points are colored and shaped by whether they fall in the allowed region (blue circle), or outside the allowed region (red cross). The results on Fig.12B show consistent complete overlaps between predicted distributions and native dihedral angles for all instances (including OOD cases) in the test sets, with at least 99.95% of the predicted samples falling within the allowed region. To further highlight that the method had learned the correct descriptors, a representative example, 1I1Y (9-mer) is shown on Fig. 12C. Fig. 12C shows position-wise Ramachandran plots for the representative 9-mer example, 1I1Y. The allowed region (light green) covers 99.95%, while the favored region (dark green) covers 98% dihedral angles. The peptide positions are highlighted using black circles on the stick figure shown in wheat. The plots highlight results from two data splits marked as such using orange (sequence- based) or green (structure-based) circles. Fig. 12C shows that the predicted dihedral distributions are tightly clustered along the anchor residues (residue positions 2 and 8), whereas more spread out along the middle residues (residue positions 5 and 6), which reflect a typical scenario for 9-mer peptides in the crystal structures. This spread along the middle residues allows for wider sampling space and thus helps in capturing peptide plasticity. Relative to shorter peptides, the distance distributions were more spread out as expected owing to higher flexibility of longer peptides (Figure 12A). The model was able to discriminate between anchor and non-anchor positions, found typically along the N- and C-terminal regions of the peptide. Here, multiple stabilizing interactions with the HLA groove, along positions 1-4 amino acids were found (Figure 12A), which are reflected in the crystal structures. The predicted dihedral angles overlapped with natives for all long peptide test cases (Fig.12B), and span the allowed range of ψ angles. Such degeneracy along predicted dihedral angles together with distances enables sampling of diverse peptide backbones. [0211] To assess the performance of the method, the inventors sampled the structural ensembles of the benchmark test sets from both data splits using FPD with the guidance of ML constraints. The conformational ensembles were obtained using on one or many templates, refined in FPD. Templates for a given target complex were selected (1) using sequence similarity (or sequence homology), (2) using structural similarity estimated from the predicted structural descriptors of the target complex relative to the co-crystal structures in the PDB, and/or (3) clustering the available peptide backbones of the target length from the PDB or MD samples and locating the cluster centers. Typically for shorter peptides (up to 10 amino acids), the inventors found in this example dataset that choosing a single template based on sequence or structural similarity performed very well. Without wishing to be bound by theory, the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 78 of 108 inventors believe that this is because peptides in the HLA groove are anchored and stretched, allowing for some plasticity along the middle residues. For longer peptides (i.e.11 amino acids and longer), choosing multiple different templates was beneficial in allowing to sample diverse peptide backbones. This is because these peptides are expected to be highly flexible and therefore using multiple starting points can help to better capture their flexibility. As a baseline, the inventors sampled structural ensembles of the benchmark sets by starting with their homologs chosen based on sequence similarity and refining without any guidance (i.e. without using constraints as predicted by the ML method described herein). The inventors demonstrate the success of the method by its ability to recover i) native, and ii) overlapping and non- overlapping peptide backbones sampled from MD simulations. Since baseline uses homologs as templates, better recovery of native co-crystal structures is generally expected due to available mutant peptides that are few edit distances away and preserve the desired peptide backbones. The inventors evaluated the models from the present method using the D-score and distributional distances as metrics. A peptide backbone was considered recovered if there was at least one FPD sample within a D-score of 10 from any of the 300 MD samples. Recovery rates were 99.9%, 83.9% and 69% for 8-, 9- and 10-mers, respectively, from the sequence- based split; 100%, 69.7%, and43.8% for 8-, 9- and 10-mers, respectively, from the structure- based split; and 71.7%, 47.5% and 82.6% for all 11-, 12- and 13-mers in the PDB. Here, the template selection strategies were constraint-based for 8- and 9-mers, sequence-based for 10- mers, and multi-template for 11-, 12- and 13-mers, followed by constrained refinement using the predictions from the machine learning model. [0212] Fig.13 shows non-limiting examples of data demonstrating that structural models sampled by the present method are diverse and overlap with MD ensembles on the benchmark sets across all the peptide lengths. Fig. 13 is a plot showing the number of overlapped FPD models with MD samples (y-axis; using a D-score cut-of of 10.0) for all the PDBIDs in the benchmark set (x-axis) across peptide lengths 8 (red), 9 (magenta), 10 (blue), 11 (orange), 12 (red) and 13 (brown) (from left to right, shown by the color bar below x-axis). The yellow and gray bars (left and right bars in each pair for a benchmark complex) represent constrained and unconstrained refinement respectively. The data show that in at least 33% and 50% of the benchmark cases, the constrained refinement strategy retrieved MD backbones better than the baseline strategy for shorter length peptides. Similarly for longer length peptides, constrained refinement recovers a higher number of MD backbones, especially for 19% of the cases while the baseline is better or comparable for 1% or 80% of the. Further, the FPD samples that do sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 79 of 108 not have backbones overlapping with MD are not necessarily incorrect, they are potentially reasonable backbones if they have lower Rosetta energies indicating overall stability of the complex together with binding scores. The total and binding energies can additionally be used to filter samples that are potentially problematic, and which have poor readouts of total and binding energies. [0213] Fig. 14A-B provide non-limiting examples of data demonstrating that the conformational ensembles sampled by the methods of the present disclosure approximate benchmark MD distributions. Fig.14A-B show distributions of evaluation metrics of the FPD models sampled using constrained and unconstrained refinement relative to the native crystal structures for all the peptide lengths. Along the y-axis, the plots show: (Top) Peptide backbone heavy-atom (N, Cα, C, and O) RMSD (in Å); (Middle) D-score values (0 indicated perfect recovery of the peptide backbones); (Bottom) Side-chain contact recovery (0 indicates no recovery whereas 1 indicates perfect side chain placement relative to the natives) of the FPD models relative to the native crystal structures for the benchmark PDBIDs (along x-axis). Data is shown on Fig.14A for sequence-base splits, and on Fig.14B for structure-based splits. for 8-10 AA and all test PDBs for 11-13 AA. The peptide lengths 8 (red), 9 (magenta), 10 (blue), 11 (orange), 12 (red) and 13 (brown) are shown by the color bar (from left to right in the order as indicated above) below x-axis. As discussed earlier, the baseline approach is expected to sample natives better, however the data showed that constraints during refinement guided sampling of models closer to the native structures. In particular, Fig.14A shows this result for 25% of 8-mers (4F7T), 30% of 9-mers (1QRN, 2C7U, 2VLL, 2X4O, 3BVN, 3D18, 3I6G, 3I6L, 3KPM, 4MNQ, 4N8V, 4O2E, 4QRS, 5ISZ, 5MER, 5TEZ, 5VVP, 5WMO, 6RPA, 6Z9W, 7KGP, 7KGS, 7LGD, 7NME, 7R80, 7ZUC), and 25% of 10-mers (1HHH, 5C09, 6MPP, 7S8Q, 7S8S, 8DVG) in test sets from sequence-based splits as highlighted by either peptide backbone heavy-atom RMSD, D-score or side-chain contact recovery (Figure 14A). In addition, similar trends can be seen for 50% of 8-mers (3BWA), 30% of 9-mers (1XR9, 2X4S, 3D18, 7L1C), and 48% of 10-mers (2CLR, 3GIV, 3MRN, 3UTT, 4F7M, 5VZ5, 6D29, 6MPP, 7JYU, 7S8S) in test sets from structure-based splits (Figure 14B). Lastly, among longer length peptides from a test set consisting of all available structures in the PDB, 43% of 11-mers (2FZ3, 2HJK, 3MV7, 3MV8, 4PRB, 4PRD, 4PRE, 4PRI, 5D9S, 5WWI, 5WWU, 6D2R, 6VPZ), 80% of 12-mers (3BW9, 4JQX, 5DDH, 5T6Y), and 14% of 13-mers (3VFS, 6AVF) had sampled conformations closer to the natives relative to the unconstrained-refinement setup (Figure 14A). sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 80 of 108 [0214] FIGS.16A-B provide non-limiting examples of preliminary control data for pMHC template structure selection and refinement, respectively. FIG. 16A provides a plot of the median simulated RMSD of pairwise distances between peptide backbone heavy atoms C, N, Cα, and O in the selected starting template structure (for peptides of 8, 9, 10, 11, and 13 amino acid residues in length) versus the corresponding RMSD for pairwise distances obtained from crystal structure data. FIG. 16B provides a plot of the median simulated RMSD of pairwise distances between peptide backbone heavy atoms C, N, Cα, and O in the generated pMHC conformer ensemble (for peptides of 8, 9, 10, 11, and 13 amino acid residues in length) versus the corresponding RMSD for pairwise distances obtained from the crystal structure data, and illustrates the dynamic nature of the pMHC structure in solution. [0215] FIGS. 17A-C provide non-limiting examples of molecular dynamics data (300ns simulation) for a sampling of peptide conformations (plot of Φ (x-axis) and Ψ (y-axis) dihedral angles) within a pMHC complex (FIG. 17A; PDBID: 2BCK; 9-mer peptide), a graphical representation of the pMHC complex (FIG.17B), and a plot of peptide backbone fluctuations (root mean square fluctuation (RMSF)) as a function of residue number for the pMHC complex (residues 1 to 181 correspond to the MHC molecule and subsequent residues correspond to the peptide) (FIG. 17C), respectively. FIG. 17C provides a plot of the minimum and maximum values of RMSF obtained for a set of several hundred 9-mer peptides versus peptide residue number. The data indicates the variation in the degree of peptide flexibility within the pMHC complex as a function of peptide residue position. The "root mean square fluctuation" (RMSF) is a measure of how much individual atoms or groups of atoms in a molecule fluctuate from their average position over time, essentially indicating the flexibility of a specific region within a structure. This is most commonly used in molecular dynamics simulations to analyze protein dynamics where areas with high RMSF values are considered more flexible. [0216] FIGS. 18A-C provide non-limiting examples of molecular dynamics data (300ns simulation) for a sampling of peptide conformations (plot of Φ (x-axis) and Ψ (y-axis) dihedral angles) within a pMHC complex (FIG. 18A; PDBID: 3VFT; 13-mer peptide), a graphical representation of the pMHC complex (FIG.18B), and a plot of peptide backbone fluctuations (root mean square fluctuation (RMSF)) as a function of residue number for the pMHC complex (residues 1 to 181 correspond to the MHC molecule and subsequent residues correspond to the peptide) (FIG.18C), respectively. FIG.18C provides a plot of RMSF for 11 different 13-mer sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 81 of 108 peptides versus peptide residue number. The data indicates the variation in the degree of peptide flexibility within the pMHC complex as a function of peptide residue position. [0217] Fig.15A-B show non-limiting examples demonstrating that peptide conformations sampled using guidance from predicted constraints as described herein can illuminate states that dictate TCR specificity. To explore the capability of the methods described herein to sample relevant protein conformations, the inventors used the HLA-A03:01 molecule in complex with a mutant PIK3CA peptide as a use case. In a recent study (Ma, J., et al. Nature Communications 16(1), 849 (2025)), the authors discuss a large conformational change of the side chain of the tryptophan at position 6 (pTrp6) that happens upon binding. In the unbound state, pTrp6 packs against the HLA-A3 α1 helix (Figure 15A, pink cartoon) but upon binding, the pTrp6 side chain flip around the axis a, and packs against the α2 helix (Figure 15A, purple cartoon). To flip around the backbone there is a large energetic barrier making it difficult for most models to overcome. Dynamic allostery is hypothesized to be common in TCR recognition of peptide/HLAs highlighting the importance of correctly predicting different conformational samples when predicting cross-reactivity and specificity. Herein, the inventors explored the capability of the methods described herein to sample the bound and unbound conformations of the PIK3CA peptide. They simulated the HLA-A03:01 in complex with a mutant PIK3CA peptide (PDBID:7L1D) and the HLA-A03:01 in complex with a mutant PIK3CA peptide with the TCR removed (PDBID:7L1D) for 300ns. The traditional MD simulations, fail to surpass the energy barriers associated with conformational changes such as sidechain flipping (Figure 15B). These limitations emphasize the importance of incorporating constraints during refinement to achieve a more comprehensive sampling of the conformational landscape. When using the guidance of constraints, the methods described herein effectively modeled the dynamic nature of peptide-HLA complexes, capturing both bound and unbound states. Fig.15A is Structural representation of HLA-A03:01 in various states. The pink cartoon (PDB ID: 7L1C) on Fig.15A represents the unbound HLA-A03:01 in complex with a mutant PIK3CA peptide. The purple cartoon (PDB ID: 7L1D) on Fig.15B depicts the TCR bound to the HLA-A03:01 in complex with a mutant PIK3CA peptide. The grey cartoons on Fig.15A illustrate the modeled HLA- A03:01 in complex with a mutant PIK3CA peptide, showcasing multiple conformations, which highlight the model’s ability to sample both the bound and unbound conformations observed in the crystal structures. The pink cartoon (PDB ID: 7L1C) on Fig.15B represents the unbound HLA-A03:01 in complex with a mutant PIK3CA peptide. The purple cartoon (PDB ID: 7L1D) on Fig.15B depicts the TCR bound to the HLA-A03:01 sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 82 of 108 in complex with a mutant PIK3CA peptide. The grey cartoons on Fig. 15B illustrate the modeled HLA-A03:01 in complex with a mutant PIK3CA peptide, show casing the highest RMSD states from the MD 7L1C simulations. Highlighting that the MD simulations are unable to overcome the energy barriers associated with sidechain flipping. [0218] The inherent plasticity of peptides displayed on the HLA molecules drive TCR specificity and cross-reactivity. Nonetheless, modeling peptide conformational flexibility without employing time-intensive MD simulations for such protein-ligand complexes has been particularly challenging. Additionally, newly developed time-efficient methods focus on monomers or small molecules. In these examples, the inventors combined the strengths of machine learning and molecular modeling tools (e.g. Rosetta (FPD)) and sample peptide conformations of lengths 8-13 amino acids presented by several HLA alleles. They show that they can sample peptide backbones that overlap with backbones from MD simulations across majority of the examples in the blind test sets. To use various metrics including the alignment- free D-score metric highlight the overlaps between MD and the predicted conformer ensembles. The inventors hypothesized that they should be able to model peptide plasticity in pHLAs with guidance from constraints that can restrict large sampling spaces. Thus, they employed machine learning model trained on pHLA binding sequences and complexes to predict the structural descriptors, especially joint distribution of distances and dihedral angles that are used as constraints for sampling. Here, the ML model uses ESM2 sequence embeddings shared across a classifier (to learn amino acid associations between peptide and HLA) and a regressor (to predict the structural constraints). In comparison with a baseline model (no constraints), they showed that the constraint-guided model is better in recovering MD samples and additional plausible diverse peptide backbones informing about the plasticity of peptides within the HLA grooves. Example Computer System [0219] FIG.19 provides a non-limiting example of a block diagram for a computer system, in accordance with some embodiments. Computer system 1900 can be an example of one implementation for Computing Platform 102 described above in FIG.1. [0220] Computer system 1900 can be a host computer connected to a network. Computer system 1900 can be a client computer or a server. As shown in FIG.19, computer system 1900 can be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device), such as a phone sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 83 of 108 or tablet. The device can include, for example, one or more of processor 1910, input device 1920, output device 1930, storage 1940, and communication device 1960. Input device 1920 and output device 1930 can generally correspond to those described elsewhere herein, and they can either be connectable or integrated with the computer. [0221] Input device 1920 can be any suitable device that provides input, such as a touch screen, keyboard or keypad, mouse, or voice-recognition device. Output device 1930 can be any suitable device that provides output, such as a touch screen, haptics device, or speaker. [0222] Storage 1940 can be any suitable device that provides storage, such as an electrical, magnetic, or optical memory including a RAM, cache, hard drive, or removable storage disk. Communication device 1960 can include any suitable device capable of transmitting and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus 1970 or wirelessly. [0223] Software 1950, which can be stored in memory / storage 1940 and executed by processor 1910, can include, for example, the programming that embodies the functionality of the present disclosure (e.g., as embodied in the methods described above). [0224] Software 1950 can also be stored and/or transported within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context disclosure, a computer-readable storage medium can be any medium, such as storage 1940, that can contain or store programming for use by or in connection with an instruction execution system, apparatus, or device. [0225] Software 1950 can also be propagated within any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a transport medium can be any medium that can communicate, propagate, or transport programming for use by or in connection with an instruction execution system, apparatus, or device. The transport readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation medium. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 84 of 108 [0226] Computer system 1900 may be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communications protocol and can be secured by any suitable security protocol. The network can comprise network links of any suitable arrangement that can implement the transmission and reception of network signals, such as wireless network connections, Tl or T3 lines, cable networks, DSL, or telephone lines. [0227] Computer system 1900 can implement any operating system suitable for operating on the network. Software 1950 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client/server arrangement or through a web browser as a web-based application or web service, for example. [0228] FIG. 20 illustrates a diagram 2000 of an example artificial intelligence (AI) architecture 2002 (which can be included as part of the one or more computing device(s) 1900 as discussed above with respect to FIG. 19) that can be utilized to determined one or more predicted peptide-protein binding interactions and/or to generate a plurality of compatible structures for a peptide-protein complex, in accordance with the disclosed embodiments. In certain embodiments, the AI architecture 2002 can be implemented utilizing, for example, one or more processing devices that may include hardware (e.g., a general purpose processor, a graphic processing unit (GPU), an application-specific integrated circuit (ASIC), a system-on- chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), a deep learning processor (DLP), a tensor processing unit (TPU), a neuromorphic processing unit (NPU), and/or other processing device(s) that can be suitable for processing various molecular data and making one or more decisions based thereon), software (e.g., instructions running/executing on one or more processing devices), firmware (e.g., microcode), or some combination thereof. [0229] In certain embodiments, as depicted by FIG. 20, the AI architecture 2002 may include machine learning (ML) algorithms and functions 2004, natural language processing (NLP) algorithms and functions 2006, expert systems 2008, computer-based vision algorithms and functions 2010, speech recognition algorithms and functions 2012, planning algorithms and functions 2014, and robotics algorithms and functions 2016. In certain embodiments, the sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 85 of 108 ML algorithms and functions 2004 may include any statistics-based algorithms that can be suitable for finding patterns across large amounts of data (e.g., “Big Data” such as genomics data, proteomics data, metabolomics data, metagenomics data, transcriptomics data, or other omics data). For example, in certain embodiments, the ML algorithms and functions 2004 may include deep learning algorithms 2018, supervised learning algorithms 2020, and unsupervised learning algorithms 2022. [0230] In certain embodiments, the deep learning algorithms 2018 may include any artificial neural networks (ANNs) that can be utilized to learn deep levels of representations and abstractions from large amounts of data. For example, the deep learning algorithms 2018 may include ANNs, such as a perceptron, a multilayer perceptron (MLP), an autoencoder (AE), a convolution neural network (CNN), a recurrent neural network (RNN), long short term memory (LSTM), a grated recurrent unit (GRU), a restricted Boltzmann Machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and deep Q-networks, a neural autoregressive distribution estimation (NADE), an adversarial network (AN), attentional models (AM), a spiking neural network (SNN), deep reinforcement learning, and so forth. [0231] In certain embodiments, the supervised learning algorithms 2020 may include any algorithms that can be utilized to apply, for example, what has been learned in the past to new data using labeled examples for predicting future events. For example, starting from the analysis of a known training data set, the supervised learning algorithms 2020 may produce an inferred function to make predictions about the output values. The supervised learning algorithms 2020 may also compare its output with the correct and intended output and find errors in order to modify the supervised learning algorithms 2020 accordingly. On the other hand, the unsupervised learning algorithms 2022 may include any algorithms that may applied, for example, when the data used to train the unsupervised learning algorithms 2022 are neither classified nor labeled. For example, the unsupervised learning algorithms 2022 may study and analyze how systems may infer a function to describe a hidden structure from unlabeled data. [0232] In certain embodiments, the NLP algorithms and functions 2006 may include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and/or text. For example, the NLP algorithms and functions 2006 may include content extraction algorithms or functions 2024, classification algorithms or functions 2026, machine translation algorithms or functions 2028, question answering (QA) algorithms or sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 86 of 108 functions 2030, and text generation algorithms or functions 2032. In certain embodiments, the content extraction algorithms or functions 2024 may include a means for extracting text or images from electronic documents (e.g., webpages, text editor documents, and so forth) to be utilized, for example, in other applications. [0233] In certain embodiments, the classification algorithms or functions 2026 may include any algorithms that may utilize a supervised learning model (e.g., logistic regression, naïve Bayes, stochastic gradient descent (SGD), k-nearest neighbors, decision trees, random forests, support vector machine (SVM), and so forth) to learn from the data input to the supervised learning model and to make new observations or classifications based thereon. The machine translation algorithms or functions 2028 may include any algorithms or functions that can be suitable for automatically converting source text in one language, for example, into text in another language. The QA algorithms or functions 2030 may include any algorithms or functions that can be suitable for automatically answering questions posed by humans in, for example, a natural language, such as that performed by voice-controlled personal assistant devices. The text generation algorithms or functions 2032 may include any algorithms or functions that can be suitable for automatically generating natural language texts. [0234] In certain embodiments, the expert systems 2008 may include any algorithms or functions that can be suitable for simulating the judgment and behavior of a human or an organization that has expert knowledge and experience in a particular field (e.g., stock trading, medicine, sports statistics, and so forth). The computer-based vision algorithms and functions 2010 may include any algorithms or functions that can be suitable for automatically extracting information from images (e.g., photo images, video images). For example, the computer-based vision algorithms and functions 2010 may include image recognition algorithms 2034 and machine vision algorithms 2036. The image recognition algorithms 2034 may include any algorithms that can be suitable for automatically identifying and/or classifying objects, places, people, and so forth that can be included in, for example, one or more image frames or other displayed data. The machine vision algorithms 2036 may include any algorithms that can be suitable for allowing computers to “see”, or, for example, to rely on image sensors cameras with specialized optics to acquire images for processing, analyzing, and/or measuring various data characteristics for decision making purposes. [0235] Herein, “or” is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A or B” means “A, B, or both,” unless sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 87 of 108 expressly indicated otherwise or indicated otherwise by context. Moreover, “and” is both joint and several, unless expressly indicated otherwise or indicated otherwise by context. Therefore, herein, “A and B” means “A and B, jointly or severally,” unless expressly indicated otherwise or indicated otherwise by context. [0236] Herein, “automatically” and its derivatives means “without human intervention,” unless expressly indicated otherwise or indicated otherwise by context. [0237] The embodiments disclosed herein are only examples, and the scope of this disclosure is not limited to them. Embodiments according to the present disclosure are in particular disclosed in the attached claims directed to a method, a storage medium, a system and a computer program product, wherein any feature mentioned in one claim category, e.g. method, can be claimed in another claim category, e.g. system, as well. The dependencies or references back in the attached claims are chosen for formal reasons only. However, any subject matter resulting from a deliberate reference back to any previous claims (in particular multiple dependencies) can be claimed as well, so that any combination of claims and the features thereof are disclosed and can be claimed regardless of the dependencies chosen in the attached claims. The subject-matter which can be claimed comprises not only the combinations of features as set out in the attached claims but also any other combination of features in the claims, wherein each feature mentioned in the claims can be combined with any other feature or combination of other features in the claims. Furthermore, any of the embodiments and features described or depicted herein can be claimed in a separate claim and/or in any combination with any embodiment or feature described or depicted herein or with any of the features of the attached claims. [0238] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art would comprehend. The scope of this disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although the present disclosure describes and illustrates respective embodiments herein as including particular components, elements, feature, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Furthermore, reference in the appended claims to an apparatus or system or a component of an apparatus or system being sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 88 of 108 adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative. Additionally, although this disclosure describes or illustrates certain embodiments as providing particular advantages, certain embodiments may provide none, some, or all of these advantages. Example Descriptions of Terms [0239] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably to refer to a polymer of amino acid residues. The terms encompass amino acid chains of any length, including full-length proteins with amino acid residues linked by covalent peptide bonds. [0240] As used herein, a “mutant peptide” or “mutated peptide” can refer to a peptide that is not present in the normal tissue (e.g., in the wild type amino acid sequences of normal tissue) of an individual subject. A mutated peptide comprises at least one mutated amino acid and can be present in a diseased tissue (e.g., collected from a particular subject), but not in a normal tissue (e.g., collected from the particular subject, collected from a different subject, and/or as identified in a database as corresponding to normal tissue). A mutated peptide can include an epitope. An epitope is the portion of a mutated peptide to which an MHC molecule or a TCR binds. Thus, the binding between the epitope of the mutated peptide and the MHC molecule or TCR can induce an immune response (as a result of the mutated peptide not being associated with a subject’s “self”). A mutated peptide can include or be a neoantigen. A mutated peptide can arise from, as non-limiting examples: a non-synonymous mutation leading to different amino acids in the protein (e.g., point mutation); a read-through mutation in which a stop codon is modified or deleted, leading to translation of a longer protein with a novel tumor-specific sequence at the C-terminus; a splice site mutation that leads to a unique tumor-specific protein sequence; a chromosomal rearrangement that gives rise to a chimeric protein with a tumor- specific sequence at a junction of two proteins (gene fusion); and/or a frameshift insertion or deletion that leads to a new open reading frame with a tumor-specific protein sequence. A mutated peptide can include a polypeptide (as characterized by a polypeptide sequence) and/or can be encoded by a nucleotide sequence. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 89 of 108 [0241] As used herein, a “C-flank” of a peptide refers to one or more amino acids upstream of the C-terminus of the peptide, from the parent protein. Optionally, a C-flank of a peptide includes one, two, three, four, five, or more amino acid residues upstream of the C-terminus of the peptide. [0242] As used herein, an “N-flank” of a peptide refers to one or more amino acids downstream of the N-terminus of the peptide, from the parent protein. Optionally, an N-flank of a peptide includes one, two, three, four, five, or more amino acid residues downstream of the N-terminus of the peptide. [0243] As used herein, an “epitope” of a peptide can refer to a region of the peptide between the C-flank and N-flank and can be recognized by a TCR. The epitope of the peptide is a part of the peptide that is recognized by a TCR on a T cell and MHC I on an antigen-presenting cell. For example, the epitope can be a peptide to which a TCR binds, such as a peptide to which the TCR binds when the peptide is bound to MHC I on an antigen-presenting cell. [0244] As used herein, a “ligand” is a peptide that is found to be presented by an MHC molecule at the cell surface from elution experiments or found to be bound to MHC in an in vitro assay. [0245] As used herein, a “sequence” refers to an amino acid sequence that includes an ordered set of amino acid identifiers. [0246] As used herein, a “peptide sequence” refers to a sequence that identifies amino acids of at least a portion of a peptide. In some cases, the peptide sequence includes a variant-coding sequence that includes a variant that is not observed in a corresponding reference sequence. [0247] When the peptide includes a mutated peptide, the variant-coding sequence, identifies amino acids of the mutation or variant. However, when the peptide does not include a mutation or variant, the variant-coding sequence does not identify amino acids of a mutation or variant (and in that instance, is the same as the reference sequence). A variant-coding sequence can be determined by collecting a disease and/or tumor sample (e.g., that includes tumor cells) and performing a sequencing analysis to identify one or more sequences corresponding to disease and/or tumor cells in the sample. In some instances, a sequencing analysis outputs an amino acid sequence. In some instances, a sequencing analysis outputs a nucleic acid sequence, which can be subsequently processed to transform codons into amino acid identifiers and thus to produce an amino acid sequence. A variant-coding sequence can sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 90 of 108 include a sequence of a neoantigen. A variant-coding sequence can, but need not, include one or more termini (e.g., the C-terminus and/or the N-terminus) of the peptide. A variant-coding sequence can include an epitope of the peptide. A variant-coding sequence can identify amino acids within a peptide having one or more variants (e.g., one or more amino acid distinctions) relative to a corresponding reference sequence. In some instances, a variant-coding sequence includes an ordered set of amino acids. In some instances, a variant-coding sequence identifies a reference peptide (e.g., by identifying a genetic reference sequence, such as by gene, start position, and/or end position; or by gene, start position, and/or length) and one or more point mutations relative to the reference peptide. [0248] As used herein, a “reference sequence” can refer to a sequence that identifies amino acids within at least part of a non-mutated peptide or wild-type peptide (e.g., wild-type, parental sequence). The non-mutated or wild-type peptide can include no variants or fewer variants than are included in a mutated peptide. The reference sequence can include an amino acid sequence encoded by a genetic sequence within a same gene relative to a gene that includes a corresponding variant-coding sequence. The reference sequence can include an amino acid sequence encoded by a genetic sequence spanning the same start and stop within a gene relative to intra-gene positions associated with a genetic sequence associated with a corresponding variant-coding sequence. The reference sequence can be identified by collecting a non-disease and/or non-tumor sample from one or more subjects (who can, but need not, include a subject from which a disease sample was collected to determine a variant-coding sequence) and performing a sequencing analysis using the sample. [0249] As used herein, a “pseudosequence” of an MHC molecule can refer to an ordered set of amino acids of the MHC molecule that typically contacts a peptide. Thus, a pseudosequence can be a subset of the amino acid sequence of a protein that is expected to contact a peptide in a peptide-protein complex. The subset can comprise consecutive or non- consecutive amino acids, but the amino acids in a pseudosequence are in the same order as in the full sequence of the protein. A pseudosequence of a protein can be defined as an order set of amino acids in the sequence of the protein that are within a predetermined distance of a peptide amino acid in a structure of the peptide-protein complex. The structure can be experimentally determined or predicted. The amino acids included in a pseudo-sequence can additionally be filtered to only include polymorphic positions. Pseudosequences for many MHC molecules are known in the art (see e.g. Nielsen et al. NetMHCpan, a Method for sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 91 of 108 Quantitative Predictions of Peptide Binding to Any HLA-A and -B Locus Protein of Known Sequence, PLoS One.2007 Aug 29;2(8):e796). [0250] As used herein, a “representation” of a sequence or “sequence representation” can include a set of values that represent or identify amino acids in the sequence and/or a set of values that represent or identify nucleic acids that encode the sequence. For example, each amino acid can be represented by a binary string and/or vector of values that is distinct from each other binary string and/or vector representing each other amino acid. The sequence representation can be generated using, for example, one-hot encoding or using a BLOcks SUbstitution Matrix (BLOSUM) matrix. For example, a multi-dimensional (e.g., 20- or 21- dimensional) array be initialized (e.g., randomly or pseudo-randomly initialized). The initialized array can include, for each amino acid, a unique vector corresponding to that amino acid. The values can be fixed such that use of such a unique vector can be assumed to represent the corresponding amino acid. [0251] As used herein, “presentation” of a peptide refers to at least part of the peptide being presented on a surface of a cell by virtue of being bound to an MHC molecule in a particular manner. The presented peptide can then be accessible to other cells, such as nearby T cells. [0252] As used herein, a “sample” can include tissue (e.g., a biopsy), single cell, multiple cells, fragments of cells, or an aliquot of body fluid. The sample can be obtained from a subject by means such as, for example, without limitation, venipuncture, excretion, ejaculation, massage, biopsy, needle aspirate, lavage sample, scraping, surgical incision, intervention, another type of sample collection means, or a combination thereof. [0253] As used herein, a “subject” encompasses one or more cells, tissue, or an organism. The subject can be a human or non-human, whether in vivo, ex vivo, or in vitro, male or female. A subject can be a mammal, such as a human.As used herein, “binding affinity” refers to affinity of binding between an amino acid (e.g., a peptide of a specific antigen) and an IPC (e.g., an MHC molecule and/or MHC allele). The binding affinity can characterize a stability, tendency, and/or strength of the binding between the peptide and an IPC. [0254] As used herein, “immunogenicity” can refer to the ability to elicit an immune response (e.g., via T cells and/or B cells). A peptide that is “immunogenic” can be one that is capable of eliciting an immune response. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 92 of 108 [0255] As used herein, “MHC” refers to the major histocompatibility complex. The human MHC is also called the human leukocyte antigen (HLA) complex. [0256] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims. Example Embodiments [0257] Embodiments disclosed herein can include: 1. A computer-implemented method for generating a plurality of compatible structures for a peptide-protein complex, the method comprising: inputting peptide sequence data for at least one peptide and protein sequence data for at least one protein into a trained machine learning model to determine: predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance data and the predicted dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data. 2. The computer-implemented method of embodiment 1, further comprising predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 93 of 108 3. The computer-implemented method of embodiments 1 or embodiment 2, wherein each compatible structure for the peptide-protein complex conforms with a posterior distribution of the predicted pairwise distance data and with a posterior distribution of the predicted dihedral angle data output by the trained machine learning model. 4. The computer-implemented method of any one of embodiments 1 to 3, wherein each compatible structure for the peptide-protein complex conforms with the sequence data for the at least one peptide and the at least one protein and with a prediction of at least one auxiliary feature. 5. The computer-implemented method of any one of embodiments 2 to 4, wherein the at least one auxiliary feature comprises a prediction of binding of the at least one peptide to the at least one protein. 6. The computer-implemented method of any one of embodiments 2 to 5, wherein the at least one auxiliary feature comprises a prediction of at least one interface feature for the peptide-protein complex. 7. The computer-implemented method of any one of embodiments 1 to 6, wherein the peptide sequence data comprises a peptide embedded representation for the at least one peptide generated using a trained peptide embedding machine learning model. 8. The computer-implemented method of any one of embodiments 1 to 7, wherein the protein sequence data comprises a protein embedded representation for the at least one protein generated using a trained protein embedding machine learning model. 9. The computer-implemented method of embodiment 8, wherein the trained peptide embedding machine learning model and the trained protein embedding machine learning model are the same model. 10. The computer-implemented method of embodiment 8 or embodiment 9, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise foundational models. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 94 of 108 11. The computer-implemented method of any one of embodiments 8 to 10, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise large language models (LLM). 12. The computer-implemented method of any one of embodiments 8 to 11, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise an Evolutionary Scale Model (ESM). 13. The computer-implemented method of any one of embodiments 8 to 12, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are integrated with, and re-trained simultaneously with, the trained machine learning model. 14. The computer-implemented method of any one of embodiments 1 to 13, wherein the trained machine learning model comprises a trained neural network. 15. The computer-implemented method of embodiment 14, wherein the trained machine learning model comprises a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained graph neural network. 16. The method of any one of embodiments 1 to 15, wherein the trained machine learning model is trained using a training data set that includes incomplete structural data for one or more peptide-protein complexes. 17. The computer-implemented method of any one of embodiments 1 to 16, wherein the trained machine learning model is further configured to determine an uncertainty in the predicted pairwise distance data and the predicted dihedral angle data. 18. The computer-implemented method of any one of embodiments 1 to 17, wherein the machine learning model is trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide-protein combinations; training a classifier to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 95 of 108 appending a regressor head to at least a portion of the trained classifier and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes. 19. The computer-implemented method of embodiment 18, wherein the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide-protein complex is presented as a separate training instance during training of the regressor head. 20. The computer-implemented method of embodiment 18 or embodiment 19, wherein the regressor is further trained on augmented pairwise distance data and dihedral angle data. 21. The computer-implemented method of embodiment 20, wherein the augmented pairwise distance data and dihedral angle data is derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor. 22. The computer-implemented method of any one of embodiments 1 to 21, wherein identifying the initial structure for the peptide-protein complex based on the predicted pairwise distance data and predicted dihedral angle data comprises: obtaining a plurality of crystal structures for peptide-protein complexes from a protein structure database; and selecting a crystal structure from the plurality of crystal structures based on a similarity between pairwise distance data and dihedral angle data determined for the crystal structure and sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 96 of 108 a maximum likelihood estimate for the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model. 23. The computer-implemented method of embodiment 22, wherein the protein structure database comprises the Protein Data Bank (PDB). 24. The computer-implemented method of any one of embodiments 1 to 23, wherein generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises: inputting the initial structure for the peptide-protein complex into a structural modeling software package; inputting the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, wherein compatible structures for the peptide- protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function. 25. The computer-implemented method of embodiment 24, wherein the iterative sampling from the plurality of possible structures for the peptide-protein complex is performed using a Markov Chain – Monte Carlo (MCMC) sampling algorithm. 26. The computer-implemented method of embodiment 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Metropolis – Hastings algorithm. 27. The computer-implemented method of embodiment 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Hamiltonian Monte Carlo algorithm. 28. The computer-implemented method of embodiment 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Langevin Monte Carlo algorithm. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 97 of 108 29. The computer-implemented method of any one of embodiments 24 to 28, wherein the structure scoring function comprises a proxy for a free energy calculation. 30. The computer-implemented method of any one of embodiments 24 to 29, wherein the structure scoring function comprises a proxy for a free energy calculation of peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof. 31. The computer-implemented method of any one of embodiments 24 to 30, wherein the scoring function is based on one or more of ab initio quantum mechanical calculations, density functional theory (DFT) calculations, semiempirical calculations, molecular mechanics force field calculations, statistical potential calculations, neural potential calculations, or machine learning models trained on structural data. 32. The computer-implemented method of any one of embodiments 24 to 31, wherein the structural modeling software package comprises FlexPepDock. 33. The computer-implemented method of any one of embodiments 1 to 32, further comprising iterating over a plurality of candidate peptides to identify at least one peptide that has a maximum likelihood for formation of the peptide-protein complex. 34. The computer-implemented method of any one of embodiments 1 to 33, wherein the at least one peptide comprises at least a portion of a peptide that is expressed in normal cells in a subject. 35. The computer-implemented method of any one of embodiments 1 to 34, wherein the at least one peptide comprises at least a portion of a neoantigen expressed in tumor cells from a cancer patient. 36. The computer-implemented method of any one of embodiments 1 to 35, wherein the at least one peptide comprises at least a portion of a peptide that is expressed in a virus or bacterium. 37. The computer-implemented method of any one of embodiments 1 to 36, wherein the at least one protein comprises a Major Histocompatibility Complex (MHC) class I protein. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 98 of 108 38. The computer-implemented method of any one of embodiments 1 to 36, wherein the at least one protein comprises a Major Histocompatibility Complex (MHC) class II protein. 39. The computer-implemented method of embodiment 37 or embodiment 38, further comprising iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in normal cells from the subject for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 40. The computer-implemented method of embodiment 37 or embodiment 38, further comprising iterating over a plurality of candidate neoantigen sequences, or portions thereof, expressed in tumor cells from the cancer patient for at least one MHC class I or class II protein to rank order the neoantigen sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 41. The computer-implemented method of embodiment 37 or embodiment 38, further comprising iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in a virus or bacterium for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide- MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 42. The computer-implemented method of embodiment 40, further comprising developing a personalized anti-cancer vaccine based on the rank ordering of the neoantigen sequences, or portions thereof. 43. The computer-implemented method of embodiment 39 or embodiment 41, further comprising developing an anti-viral vaccine, an anti-bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof. 44. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the computer-implemented method of any one of embodiments 1 to 43. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 99 of 108 45. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform the computer-implemented method of any one of embodiments 1 to 43. [0258] The description provides preferred example embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the description of the preferred example embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims. sf-6628464

Claims

ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 100 of 108 CLAIMS What is claimed is: 1. A computer-implemented method for generating a plurality of compatible structures for a peptide-protein complex, the method comprising: inputting peptide sequence data for at least one peptide and protein sequence data for at least one protein into a trained machine learning model to determine: predicted pairwise distance data for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and predicted dihedral angle data for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex; identifying an initial structure for the peptide-protein complex based on the predicted pairwise distance data and the predicted dihedral angle data; and generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data. 2. The computer-implemented method of claim 1, further comprising predicting at least one auxiliary feature of the peptide-protein complex based on the plurality of compatible structures for the peptide-protein complex. 3. The computer-implemented method of claims 1 or claim 2, wherein the predicted pairwise distance data comprises one or more predicted posterior distributions of pairwise distances and the predicted dihedral angle data comprises one or more predicted posterior distributions of dihedral angles, and each compatible structure for the peptide-protein complex conforms with the one or more posterior distributions of the predicted pairwise distance data and with the one or more posterior distributions of the predicted dihedral angle data output by the trained machine learning model. 4. The computer-implemented method of any one of claims 1 to 3, wherein the trained machine learning model has been trained at least in part to predict at least one auxiliary feature of input peptide-protein complexes. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 101 of 108 5. The computer-implemented method of any one of claim 4, wherein the at least one auxiliary feature comprises a prediction of binding of the at least one peptide to the at least one protein. 6. The computer-implemented method of any one of claims 4 to 5, wherein the at least one auxiliary feature comprises a prediction of at least one interface feature for the peptide- protein complex. 7. The computer-implemented method of any one of claims 1 to 6, wherein the peptide sequence data comprises a peptide embedded representation for the at least one peptide generated using a trained peptide embedding machine learning model. 8. The computer-implemented method of any one of claims 1 to 7, wherein the protein sequence data comprises a protein embedded representation for the at least one protein generated using a trained protein embedding machine learning model. 9. The computer-implemented method of claim 8, wherein the trained peptide embedding machine learning model and the trained protein embedding machine learning model are the same model. 10. The computer-implemented method of claim 8 or claim 9, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise foundational models, or machine learning models that have been trained in a self-supervised manner to learn an embedding from an input sequence using unlabeled protein and/or peptide sequence data. 11. The computer-implemented method of any one of claims 8 to 10, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise deep learning models, deep neural networks, or large language models (LLM). 12. The computer-implemented method of any one of claims 8 to 11, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model comprise an Evolutionary Scale Model (ESM). sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 102 of 108 13. The computer-implemented method of any one of claims 8 to 12, wherein the trained peptide embedding machine learning model and/or the trained protein embedding machine learning model are integrated with, and re-trained simultaneously with, the trained machine learning model. 14. The computer-implemented method of any one of claims 1 to 13, wherein the trained machine learning model comprises a trained neural network. 15. The computer-implemented method of claim 14, wherein the trained machine learning model comprises a trained multilayer perceptron (MLP), a trained attention-based neural network, or a trained state space model. 16. The method of any one of claims 1 to 15, wherein the trained machine learning model is trained using a training data set that includes incomplete structural data for one or more peptide-protein complexes. 17. The computer-implemented method of any one of claims 1 to 16, wherein the trained machine learning model is configured to predict parameters characterizing: a distribution of pairwise distance for at least one peptide amino acid residue of the at least one peptide and at least one protein amino acid residue of the at least one protein in the peptide-protein complex; and a joint distribution of dihedral angles ψ and φ for at least one peptide amino acid residue in the at least one peptide in the peptide-protein complex.. 18. The computer-implemented method of any one of claims 1 to 17, wherein the machine learning model is trained by: receiving peptide sequence data for a plurality of peptides; receiving protein sequence data for a plurality of proteins; receiving peptide-protein binding data for a plurality of peptide-protein combinations; training a classifier to predict peptide-protein binding for specific combinations of a peptide and a protein based on the peptide sequence data, the protein sequence data, and the peptide-protein binding data; and appending a regressor head to at least a portion of the trained classifier and training the regressor head to predict probabilistic distributions of pairwise distance data and dihedral angle data for peptide-protein complexes based on pairwise distance data and dihedral angle data derived from crystal structure data for a set of example peptide-protein complexes. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 103 of 108 19. The computer-implemented method of claim 18, wherein the pairwise distance data and dihedral angle data derived from the crystal structure data for each combination of a single peptide amino acid residue and a single protein amino acid residue in an example peptide- protein complex is presented as a separate training instance during training of the regressor head. 20. The computer-implemented method of claim 18 or claim 19, wherein the regressor is further trained on augmented pairwise distance data and dihedral angle data. 21. The computer-implemented method of claim 20, wherein the augmented pairwise distance data and dihedral angle data is data that has not been experimentally determined, and/or data that is derived by obtaining a predicted crystal structure for a peptide and protein sequence known to bind to each other to form a peptide-protein complex using a trained structure prediction machine learning model, and/or is data that is derived by: receiving peptide sequence data and protein sequence data for at least one peptide and at least one protein that are known to bind and form a peptide-protein complex; processing the peptide sequence data and protein sequence data for the at least one peptide and the at least one protein using a previous version of the trained machine learning model to determine an uncertainty in pairwise distance data and an uncertainty in dihedral angle data for the peptide-protein complex; applying the determined uncertainty in the pairwise distance data and the determined uncertainty in the dihedral angle data to the structure for an example peptide-protein complex to generate a plurality of compatible structures for the peptide-protein complex; and extracting pairwise distance data and dihedral angle data for the plurality of compatible peptide-protein complex structures for use as augmented training data for the regressor. 22. The computer-implemented method of any one of claims 1 to 21, wherein identifying the initial structure for the peptide-protein complex based on the predicted pairwise distance data and predicted dihedral angle data comprises: obtaining a plurality of crystal structures for peptide-protein complexes from a protein structure database; and selecting a crystal structure from the plurality of crystal structures based on a sequence similarity or a similarity between pairwise distance data and dihedral angle data determined for sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 104 of 108 the crystal structure and the predicted pairwise distance data and the predicted dihedral angle data determined by the trained model. 23. The computer-implemented method of claim 22, wherein the protein structure database comprises the Protein Data Bank (PDB), and/or wherein selecting an experimental structure comprises: (i) for each example peptide-protein experimental structure in the plurality of experimental structures, determining a of the distances and dihedral angles in the peptide- protein experimental structure using the predicted distributions for pairwise distances and dihedral angles for the peptide-protein complex output by the trained machine learning model; and (ii) selecting the example peptide-protein experimental structure that has the highest likelihood. 24. The computer-implemented method of any one of claims 1 to 23, wherein generating the plurality of compatible structures for the peptide-protein complex based on the initial structure, the predicted pairwise distance data, and the predicted dihedral angle data comprises: inputting the initial structure for the peptide-protein complex into a structural modeling software package; inputting the predicted pairwise distance data and the predicted dihedral angle data determined by the trained machine learning model for the peptide-protein complex into the structural modeling software package; and iteratively sampling from a plurality of possible structures for the peptide-protein complex and evaluating a structure scoring function to identify the plurality of compatible structures for the peptide-protein complex, wherein compatible structures for the peptide- protein complex are structures that: (i) are compatible with one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data, and (ii) optimize the structure scoring function. 25. The computer-implemented method of claim 24, wherein the iterative sampling from the plurality of possible structures for the peptide-protein complex is performed using a Markov Chain – Monte Carlo (MCMC) sampling algorithm. 26. The computer-implemented method of claim 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Metropolis – Hastings algorithm. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 105 of 108 27. The computer-implemented method of claim 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Hamiltonian Monte Carlo algorithm. 28. The computer-implemented method of claim 25, wherein the Markov Chain – Monte Carlo (MCMC) sampling algorithm comprises a Langevin Monte Carlo algorithm. 29. The computer-implemented method of any one of claims 24 to 28, wherein the structure scoring function comprises a proxy for a free energy calculation. 30. The computer-implemented method of any one of claims 24 to 29, wherein the structure scoring function comprises a proxy for a free energy calculation of peptide-protein binding, a proxy for a free energy calculation for at least one peptide-protein interface feature, or any combination thereof. 31. The computer-implemented method of any one of claims 24 to 30, wherein the scoring function is based on one or more of ab initio quantum mechanical calculations, density functional theory (DFT) calculations, semiempirical calculations, molecular mechanics force field calculations, statistical potential calculations, neural potential calculations, or machine learning models trained on structural data. 32. The computer-implemented method of any one of claims 24 to 31, wherein the structural modeling software package comprises FlexPepDock and/or wherein the one or more constraints imposed by the predicted pairwise distance data and predicted dihedral angle data are included in the structure scoring function. 33. The computer-implemented method of any one of claims 1 to 32, further comprising iterating over a plurality of candidate peptides to identify at least one peptide that has a maximum likelihood for formation of the peptide-protein complex. 34. The computer-implemented method of any one of claims 1 to 33, wherein the at least one peptide comprises at least a portion of a peptide that is expressed in normal cells in a subject. 35. The computer-implemented method of any one of claims 1 to 34, wherein the at least one peptide comprises at least a portion of a neoantigen expressed in tumor cells from a cancer patient. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 106 of 108 36. The computer-implemented method of any one of claims 1 to 35, wherein the at least one peptide comprises at least a portion of a peptide that is expressed in a virus or bacterium. 37. The computer-implemented method of any one of claims 1 to 36, wherein the at least one protein comprises a Major Histocompatibility Complex (MHC) class I protein. 38. The computer-implemented method of any one of claims 1 to 36, wherein the at least one protein comprises a Major Histocompatibility Complex (MHC) class II protein. 39. The computer-implemented method of claim 37 or claim 38, further comprising iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in normal cells from the subject for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide- MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 40. The computer-implemented method of claim 37 or claim 38, further comprising iterating over a plurality of candidate neoantigen sequences, or portions thereof, expressed in tumor cells from the cancer patient for at least one MHC class I or class II protein to rank order the neoantigen sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 41. The computer-implemented method of claim 37 or claim 38, further comprising iterating over a plurality of candidate peptide sequences, or portions thereof, expressed in a virus or bacterium for at least one MHC class I or class II protein to rank order the peptide sequences, or portions thereof, according to a likelihood of formation of a peptide-MHC class I (pMHC-I) or a peptide-MHC class II (pMHC-II) complex. 42. The computer-implemented method of claim 40, further comprising developing a personalized anti-cancer vaccine based on the rank ordering of the neoantigen sequences, or portions thereof. 43. The computer-implemented method of claim 39 or claim 41, further comprising developing an anti-viral vaccine, an anti-bacterial vaccine, or an auto-reactive vaccine based on the rank ordering of the peptide sequences, or portions thereof. sf-6628464 ATTORNEY DOCKET PATENT APPLICATION 146392067640-P38819-WO 107 of 108 44. A system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to perform the computer-implemented method of any one of claims 1 to 43. 45. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions which, when executed by one or more processors of a system, cause the system to perform the computer-implemented method of any one of claims 1 to 43. sf-6628464
PCT/US2025/020470 2024-03-19 2025-03-18 MACHINE LEARNING-BASED METHODS FOR MODELING pMHC CONFORMERS Pending WO2025199168A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202463567360P 2024-03-19 2024-03-19
US63/567,360 2024-03-19

Publications (1)

Publication Number Publication Date
WO2025199168A1 true WO2025199168A1 (en) 2025-09-25

Family

ID=95446619

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/020470 Pending WO2025199168A1 (en) 2024-03-19 2025-03-18 MACHINE LEARNING-BASED METHODS FOR MODELING pMHC CONFORMERS

Country Status (1)

Country Link
WO (1) WO2025199168A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20250308627A1 (en) * 2024-04-02 2025-10-02 Nec Laboratories America, Inc. T-cell receptor-peptide interaction prediction for medical decision making

Non-Patent Citations (31)

* Cited by examiner, † Cited by third party
Title
"Modeling Peptide-Protein Interactions", 25 February 2017, SPRINGER NEW YORK, ISBN: 978-1-4939-6798-8, article ALAM NAWSAD ET AL: "Modeling Peptide-Protein Structure and Binding Using Monte Carlo Sampling Approaches: Rosetta FlexPepDock and FlexPepBind", pages: 139 - 169, XP093289675, DOI: 10.1007/978-1-4939-6798-8_9 *
ABDIN OSAMA ET AL: "PepNN: a deep attention model for the identification of peptide binding sites", COMMUNICATIONS BIOLOGY, vol. 5, no. 1, 26 May 2022 (2022-05-26), pages 1 - 10, XP093071275, Retrieved from the Internet <URL:https://www.nature.com/articles/s42003-022-03445-2> [retrieved on 20250624], DOI: 10.1038/s42003-022-03445-2 *
ABELLA JAYVEE R. ET AL: "Markov state modeling reveals alternative unbinding pathways for peptide-MHC complexes", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES (PNAS), vol. 117, no. 48, 12 November 2020 (2020-11-12), pages 30610 - 30618, XP093289683, ISSN: 0027-8424, DOI: 10.1073/pnas.2007246117 *
ALFORD ET AL., J. CHEM. THEORY COMPUT, vol. 13, no. 6, 2017, pages 3031 - 3048
CHU ET AL., IMMUNE EPITOPE DATABASE & TOOLS, 2022, Retrieved from the Internet <URL:www.iedb.org/home_v3.php>
CHU ET AL.: "A Transformer-Based Model to Predict Peptide-HLA Class I Binding and Optimize Mutated Peptides for Vaccine Design", NATURE MACHINE INTELLIGENCE, vol. 4, no. 3, 2022, pages 300 - 311, XP093092969, DOI: 10.1038/s42256-022-00459-7
CHU., NATURE MACHINE INTELLIGENCE, vol. 4, no. 3, pages 300 - 311
CORNELL ET AL., J AM CHEM SOC, vol. 118, no. 9, 1996, pages 2309 - 2309
DELAUNAY ANTOINE P. ET AL: "Peptide-MHC Structure Prediction With Mixed Residue and Atom Graph Neural Network", BIORXIV, 24 November 2022 (2022-11-24), XP093289268, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2022.11.23.517618v1.full.pdf> [retrieved on 20250624], DOI: 10.1101/2022.11.23.517618 *
HASHEMI NASSER ET AL: "Improved prediction of MHC-peptide binding using protein language models", FRONTIERS IN BIOINFORMATICS, vol. 3, 17 August 2023 (2023-08-17), XP093289282, ISSN: 2673-7647, DOI: 10.3389/fbinf.2023.1207380 *
ITO ET AL.: "Regulation of the Induction and Function of Cytotoxic T Lymphocytes by Natural Killer T Cell", JOURNAL OF BIOMEDICINE AND BIOTECHNOLOGY, vol. 2010, 2010, pages 641757
JONES ET AL.: "Markov Chain Monte Carlo in Practice", ANNU. REV. STAT. APPL., vol. 9, 2022, pages 557 - 578
JUMPER, J.M ET AL., PLOS COMPUTATIONAL BIOLOGY, vol. 14, no. 12, 2018, pages 1006342
LEMAN ET AL., NATURE METHODS, vol. 17, no. 7, 2020, pages 665 - 680
LI ET AL.: "BioSeq-BLM: A Platform for Analyzing DNA, RNA and Protein Sequences Based on Biological Language Models", NUCLEIC ACIDS RES, vol. 49, 2021, pages 129 - 129
MA, J. ET AL., NATURE COMMUNICATIONS, vol. 16, no. 1, 2025, pages 849
MOTMAEN ET AL., PROC. NAT. AC. SCI., vol. 120, no. 9, 2023, pages 2216697120
MOTMAEN, A ET AL., PNAS, vol. 120, no. 9, 2023, pages 2216697120
NAVEED ET AL.: "A Comprehensive Overview of Large Language Models", ARXIV.ORG/PDF/2307.06435.PDF, 2023
NAVEED: "A Comprehensive Overview of Large Language Models", ARXIV.ORG/PDF/2307.06435 PDF, 2023
NIELSEN ET AL.: "NetMHCpan, a Method for Quantitative Predictions of Peptide Binding to Any HLA-A and -B Locus Protein of Known Sequence", PLOS ONE, vol. 2, no. 8, 29 August 2007 (2007-08-29), pages 796
NORTH ET AL.: "A New Clustering of Antibody CDR Loop Conformations", J. MOL BIOL., vol. 406, no. 2, 2011, pages 228 - 256, XP028129711, DOI: 10.1016/j.jmb.2010.10.030
NUCLEIC ACIDS RES, vol. 36, 2008, pages 190 - 195
OFER ET AL.: "The Language of Proteins: NLP, Machine Learning & Protein Sequences", COMPUTATIONAL AND STRUCTURAL BIOTECHNOLOGY JOURNAL, vol. 19, 2021, pages 1750 - 1758, XP093222546, DOI: 10.1016/j.csbj.2021.03.022
RAVEH ET AL., PROTEINS, vol. 78, July 2010 (2010-07-01), pages 2029 - 2040
RAVEH ET AL., PROTEINS: STRUCTURE, FUNCTION, AND BIOINFORMATICS, vol. 78, no. 9, 2010, pages 2029 - 2040
RAVEH ET AL.: "Sub-angstrom modeling of complexes between flexible peptides and globular proteins", PROTEINS, vol. 78, July 2010 (2010-07-01), pages 2029 - 2040
RAVENZWAAIJ ET AL.: "A Simple Introduction to Markov Chain Monte-Carlo Sampling", PSYCHON BULL REV, vol. 25, 2018, pages 143 - 154, XP036464204, DOI: 10.3758/s13423-016-1015-8
RIVES ET AL., PNAS, vol. 118, no. 15, 2021, pages 2016239118, Retrieved from the Internet <URL:https://github.com/facebookresearch/esm>
RIVES ET AL.: "Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences", PNAS, vol. 118, no. 15, 2021, pages 2016239118, XP093038975, DOI: 10.1073/pnas.2016239118
VERKUL ET AL.: "Language Models Generalize Beyond Natural Proteins", BIORXIV, December 2022 (2022-12-01)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20250308627A1 (en) * 2024-04-02 2025-10-02 Nec Laboratories America, Inc. T-cell receptor-peptide interaction prediction for medical decision making

Similar Documents

Publication Publication Date Title
Simonovsky et al. DeeplyTough: learning structural comparison of protein binding sites
Thomas et al. Artificial intelligence in vaccine and drug design
Vishnoi et al. Artificial intelligence and machine learning for protein toxicity prediction using proteomics data
Wardah et al. Predicting protein-peptide binding sites with a deep convolutional neural network
EP4677594A2 (en) Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models
Zhang et al. Epitope-anchored contrastive transfer learning for paired CD8+ T cell receptor–antigen recognition
Xu et al. NetBCE: an interpretable deep neural network for accurate prediction of linear B-cell epitopes
Nedyalkova et al. Sequence-based prediction of plant allergenic proteins: machine learning classification approach
US20250014683A1 (en) Systems and methods for polymer sequence prediction
Wang et al. RACER-m leverages structural features for sparse T cell specificity prediction
Chuang et al. DeepNeoAG: neoantigen epitope prediction from melanoma antigens using a synergistic deep learning model combining protein language models and multi-window scanning convolutional neural networks
Zhang et al. PreAlgPro: prediction of allergenic proteins with pre-trained protein language model and efficient neutral network
Weber et al. Unsupervised and supervised AI on molecular dynamics simulations reveals complex characteristics of HLA-A2-peptide immunogenicity
Weeder et al. pepsickle rapidly and accurately predicts proteasomal cleavage sites for improved neoantigen identification
Yasser et al. Predicting MHC-II binding affinity using multiple instance regression
Shen et al. Self-iterative multiple-instance learning enables the prediction of CD4+ T cell immunogenic epitopes
US20260004873A1 (en) Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models
Arowolo et al. An adaptive genetic algorithm with recursive feature elimination approach for predicting malaria vector gene expression data classification using support vector machine kernels
Slone et al. STAG-LLM: Predicting TCR-pHLA binding with protein language models and computationally generated 3D structures
Gao et al. Structure-directed pan-specific t-cell receptor–peptide-major histocompatibility complex interaction prediction
Liu et al. AlphaFlex: Ensembles of the human proteome representing disordered regions
WO2025160309A1 (en) Systems and methods for dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models
Wang et al. Sabre: Self-attention based model for predicting t-cell receptor epitope specificity
Qayyum et al. ImmUQBench: a benchmark on uncertainty quantification of protein immunogenicity prediction
Zhou et al. GPT2-ICC: A data-driven approach for accurate ion channel identification using pre-trained large language models

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25719490

Country of ref document: EP

Kind code of ref document: A1