EP4612693A1 - Protein solutions - Google Patents

Protein solutions

Info

Publication number
EP4612693A1
EP4612693A1 EP23801725.5A EP23801725A EP4612693A1 EP 4612693 A1 EP4612693 A1 EP 4612693A1 EP 23801725 A EP23801725 A EP 23801725A EP 4612693 A1 EP4612693 A1 EP 4612693A1
Authority
EP
European Patent Office
Prior art keywords
viscosity
protein solution
concentration
protein
neural network
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23801725.5A
Other languages
German (de)
French (fr)
Inventor
Christoph GRAPENTIN
Jonathan Schmitt
Abbas RAZVI
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Lonza AG
Original Assignee
Lonza AG
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Lonza AG filed Critical Lonza AG
Publication of EP4612693A1 publication Critical patent/EP4612693A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
    • G16B40/20Supervised data analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/048Activation functions
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B35/00ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/30Prediction of properties of chemical compounds, compositions or mixtures
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/60In silico combinatorial chemistry
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/70Machine learning, data mining or chemometrics

Definitions

  • the present invention refers to a method for providing a computer-implemented neural network configured for predicting a concentration-dependent viscosity of a protein solution. Further, the present invention refers to methods for predicting a concentration-dependent viscosity of a protein solution, for determining a concentration-dependent viscosity of a protein solution, and for providing a drug product comprising a protein solution, by using such a computer-implemented neural network.
  • Therapeutic proteins such as monoclonal antibodies (mAbs) have become an important factor in the treatment of a broad variety of diseases, such as cancer, immune-mediated disorders and infectious diseases.
  • mAbs protein-based drug products
  • DPs protein-based drug products
  • Subcutaneous (s.c.) injection allows patients to self-administer protein-based drug products, such as drug products based on monoclonal antibodies, by use of prefilled syringes, auto-injectors or other delivery devices, and by this often increasing quality of life and compliance for patients with chronic conditions.
  • protein-based drug products such as drug products based on monoclonal antibodies
  • a single injection volume is limited to about less than 2 mL, determined by the available subcutaneous space and sensation of tolerable pain by the patient. Due to this volume limitation, the use of drug products with high drug substance concentrations are required.
  • mAbs have a high specificity but they also require considerable therapeutic dosages. This consequently results in high concentrations of drug products (DP) comprising solutions of proteins such as mAbs, often exceeding 100 mg/mL protein in solution for s.c. administration.
  • DP drug products
  • PPIs protein-protein-interactions
  • proteins e.g. mAbs'
  • solubility e.g. aggregation
  • viscosity e.g. mAbs'
  • PPIs are inter alia determined by a protein’s (e.g. mAb's) primary amino acid sequence and the resulting three-dimensional structure with charged or hydrophobic patches.
  • Solution conditions can also influence protein-protein-interactions, e.g., by modulating size of charged patches via pH, shielding of charged patches via short-ranged electrostatic interaction using salts, buffer substances, amino acids or other charged excipients.
  • Arginine is a common excipient tested for viscosity reduction, its dual mode of action being both the shielding of charged as well as of hydrophobic patches.
  • 17 of 34 FDA approved drug products with high mAb concentration use salts or amino acids as excipients, likely with the aim to reduce protein-protein-interactions and thus to lower mAb solution viscosity.
  • Highly viscous solutions can be a major roadblock in the development of proteinbased drug products. Disadvantages may be high costs, due to high loss and low recovery in purification, difficult manufacturing or filling, and ultimately low administrability due to need for high injection forces and slow administration with potential sensation of pain. In general, solutions with a dynamic viscosity above about 15 to 30 mPa*s or even above about 15 to 20 mPa*s are considered problematic.
  • the desire to develop a high concentration formulation may not be apparent a priori for a new molecule, it may appear also as a consequence of need for unexpectedly high doses or change of target route of administration. Clinical trials are usually started in phase 1 with lower concentrated ( ⁇ 50 mg/mL protein in solution) formulations, whereas high protein concentrations (>100 mg/mL protein in solution) are typically explored in later phases once safety and efficacious dose levels are established.
  • SCM spatial charge map
  • the deep learning algorithm that Lai established used the preprocessed antibody sequences as input and the SCM scores obtained from MD simulations as output for model training; and a DeepSCM surrogate model for SCM scores was developed based on the convolutional neural network (CNN) architecture. The SCM scores then are used for predicting high concentration viscosity. Lai reports an accuracy of 0.9.
  • the present invention refers to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated to said protein solution.
  • the present invention in particular the method, may comprise the provision and use of a computer-implemented neural network, in particular of a computer-implemented neural network to be trained.
  • the provision and use may be performed to provide a trained computer-implemented neural network, in particular which is configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated to said protein solution.
  • Said method specifically using the computer-implemented neural network, in particular using the computer-implemented neural network to be trained, comprises the steps of:
  • the input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof; ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • HIC hydrophobic interaction chromatography
  • kD diffusion interaction parameter
  • A2 second virial coefficient
  • computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof
  • a further embodiment of the invention comprises a computer-implemented method for predicting a concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, in particular a trained computer- implemented neural network, more particularly which is provided upon provision and use of a computer-implemented neural network to be trained as described above.
  • the computer-implemented method for predicting the concentrationdependent viscosity of a protein solution of the invention has a high accuracy, which may be reflected by R2 values greater than 0.95, and even greater than 0.99.
  • a still further embodiment of the invention comprises a method for determining a concentration-dependent viscosity of a protein solution by using such a computer- implemented neural network, in particular such a trained computer-implemented neural network.
  • a still further embodiment of the invention comprises a method for providing a drug product comprising a protein solution, which enables effective and efficient assessment of the suitability of said protein solution for use as a drug product already at an early stage of development.
  • any of the methods disclosed herein which use a computer- implemented method for predicting or determining a concentration-dependent viscosity of a protein solution use the computer-implemented neural network which is initially configured (sometimes hereinafter referred to as "trained"), for predicting a concentration-dependent viscosity of a protein solution as described herein, also with all its embodiments.
  • the task of a computer-implemented neural network of calculating a concentration-dependent viscosity may be called determining or predicting the concentration-dependent viscosity; the terms "determining” and “predicting” in connection with the calculation of the computer- implemented neural network can be used interchangeably.
  • the suitability of the protein to be used as a drug product in form of a solution is determined by the viscosity of the protein solution at a high concentration, in particular at a concentration of more than 100 mg/mL of the protein in the solution, which should not exceed 15 to 30 mPa*s, preferably 15 to 20 mPa*s, in particular 15 mPa*s.
  • the present invention also comprises a method for providing a computer-implemented neural network (sometimes hereinafter referred to “the neural network”), preferably an artificial neural network (ANN), also referred to as trained ANN.
  • the present invention comprises the provision and use of such a computer-implemented neural network.
  • the neural network in particular the trained neuronal network, is configured for predicting or determining a concentration-dependent viscosity of a protein solution in dependence on a plurality of input parameters associated to said protein solution.
  • the method comprises a step of providing a plurality of training data sets, each of which is associated to a specific protein solution and comprises a plurality of input parameters being indicative of the specific protein solution and at least one associated output parameter being indicative of a concentration-dependent viscosity of the specific protein solution; and a step of training a neural network, i.e. a neural network to be trained or an untrained neural network, based on the training data sets to provide the computer-implemented neural network configured for predicting the concentration-dependent viscosity, i.e. to provide the trained ANN.
  • a neural network i.e. a neural network to be trained or an untrained neural network
  • protein solution refers to an aqueous solution of a protein, preferably a therapeutic protein (often referred to as Active Pharmaceutical Ingredient ‘API’).
  • the protein solution refers to any solution comprising at least one protein dissolved or substantially dissolved in an aqueous medium such as water, a buffer, cell culture medium and the like.
  • protein means a polypeptide composed of a sequence of amino acids.
  • a protein may be a therapeutic protein, e.g., used in diagnosis, treatment, and/or prevention of a disease or disorder.
  • a protein can be a native protein, that is, a protein produced by a naturally-occurring and non-recombinant cell; or it can be produced by a genetically-engineered or recombinant cell and may comprise molecules having the amino acid sequence of the native protein, or molecules having deletions from, additions to, and/or substitutions of one or more amino acids of the native sequence, or molecules having the amino acid sequence of a protein with no relation to a native protein.
  • the term also includes amino acid polymers in which one or more amino acids are chemical analogues of a corresponding naturally-occurring amino acid polymer.
  • Therapeutic proteins included in the protein solution may be antibodies, which may include monoclonal antibodies and polyclonal antibodies, whole antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments, e.g., Fab, Fab', F(ab')2, fragments and the like, including hybrid fragments of any immunoglobulin or any natural, synthetic or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex in one embodiment monoclonal antibodies.
  • epipe denotes a protein determinant capable of specifically binding to an antibody.
  • Epitopes usually consist of chemically active surface groupings of molecules such as amino acids or sugar side chains and usually epitopes have specific three-dimensional structural characteristics, as well as specific charge characteristics. Conformational and non-conformational epitopes are distinguished in that the binding to the former but not the latter is lost in the presence of denaturing solvents.
  • Preferred therapeutic proteins are monoclonal antibodies.
  • the term ‘monoclonal antibody’ (mAb) as used herein refers to an antibody obtained from a population of substantially homogeneous antibodies, i.e. the individual antibodies comprising the population are identical except for possible naturally occurring mutations that may be present in minor amounts. Monoclonal antibodies are highly specific, being directed against a single antigenic site. Furthermore, in contrast to polyclonal antibody preparations, which include different antibodies directed against different antigenic sites (determinants or epitopes), each monoclonal antibody is directed against a single antigenic site on the antigen.
  • the modifier “monoclonal” indicates the character of the antibody as being obtained from a substantially homogeneous population of antibodies and is not to be construed as requiring production of the antibody by any particular method.
  • Preferred monoclonal antibodies are those of class Immunoglobulin G (IgG) such as lgG1 , lgG2, lgG3, and lgG4.
  • aqueous media of the protein solution are known to the skilled artisan.
  • the aqueous medium is or comprises water or an aqueous buffer such as Histidine-HCI buffer.
  • buffers which are commonly used for the formulation of a drug product are known to the skilled person, such as acetate buffer, phosphate buffer, succinate buffer or citrate buffer.
  • the aqueous medium is a buffer having a pH of 5 to 8, more preferably of 5 to 7.5, even more preferably of 5.5 to 6.5.
  • the aqueous medium may also have other pH values suitable for providing a protein solution, the common pH values which the aqueous mediums of drug products have, are known to the skilled person.
  • any input parameter and also any of the at least one associated output parameter, the latter being provided for the configuring of the computer- implemented neural network is determined in a similar aqueous medium that is also used in the intended drug product comprising said protein solution, or more preferably even in the same aqueous medium, such as with same pH, with same buffer, same viscosity, in particular same buffer viscosity etc.
  • At least one pharmaceutically acceptable excipient may also be comprised in the protein solution.
  • Suitable excipients are known to the skilled artisan such as stabilizers, pH modifying agents, such as one or more buffers, etc.
  • a polysorbate such as polysorbate 20, 40, 60 or 80 may be used as stabilizing agent, which can improve stability of a protein in an aqueous solution.
  • Other stabilizer can be for example sugars or sugar alcohols, surfactants, salts or antioxidants, such as for example L-methionine.
  • the method for providing the computer-implemented neural network comprises the step of providing the plurality of training data sets.
  • the use of the computer-implemented neural network, in particular of the neural network to be trained may comprise the step of providing the plurality of training data sets.
  • Each of the training data sets is associated to a specific protein solution.
  • at least two of the plurality of training data sets are associated to different protein solutions.
  • each of the training data sets may be associated to a specific protein, i.e. the protein included in the protein solution.
  • the plurality of training data sets may be associated to different proteins, which in particular constitute a set of proteins. In other words, each protein included in the set of proteins is associated to at least one training data set.
  • the proteins underlying each training data set are different proteins, but similar to each other in at least one aspect.
  • the proteins in the set of proteins may be any of the above-mentioned types of proteins. More preferably, all the proteins in the set of proteins of the training data sets are antibodies, even more preferably all proteins in the set of proteins are monoclonal antibodies.
  • the number of training data sets, and hence the number of proteins may be more than two.
  • the accuracy of prediction (or determination) of the viscosity can be expressed, for example, by the accuracy value which is comparable to the value of the coefficient of determination R2; these values range from 0 (reflecting no accuracy of determination, which can also be called accuracy of prediction) to 1 (reflecting 100% accurate determination).
  • the number of training data sets (and hence proteins in the set of proteins) may be at least 15, preferably at least 20, more preferably at least 25.
  • higher numbers of training data sets result in better accuracy.
  • the accuracy of determination no longer increases linearly, but approaches the upper limit of 1 . For example, where 25 training data sets are provided, the accuracy of determination is usually quite high.
  • concentration-dependent viscosity refers to a parameter being indicative of the protein solution's viscosity, in particular the protein solution's dynamic viscosity, at one or more protein concentrations of the protein solution.
  • the concentration-dependent viscosity associates at least one protein concentration of the protein solution to a corresponding viscosity of the protein solution.
  • the concentrationdependent viscosity may be or may indicate a viscosity in form of a viscosity value c s of the protein solution at a selected protein concentration, in particular at a single protein concentration.
  • the concentration-dependent viscosity can by expressed and/or determined as a viscosity value cs of the protein solution at a selected protein concentration.
  • the concentration-dependent viscosity may indicate a viscosity of the protein solution at more than one protein concentration.
  • the concentration-dependent viscosity may be represented by a function f v , in particular a mathematical function, associating a viscosity of the protein solution to a protein concentration of the protein solution.
  • the concentrationdependent viscosity may be represented by a curve diagram associating viscosity values to protein concentrations of the protein solution.
  • Equation (2) linearized viscosity behavior; wherein ?] rei refers to a relative viscosity; A refers to an intercept; B refers to a slope; c refers to the protein concentration.
  • the concentration-dependent viscosity of the protein solution may be represented by at least one, preferably both, of the constants A and B of the above equation (1 ).
  • the computer-implemented neural network in particular the trained computer-implemented neural network, more specifically the trained ANN, and/or the computer-implemented neural network to be trained, may be configured so as to determine at least one, preferably both, of constants A and B of the above equation (1).
  • the term computer-implemented neural network may be an artificial neural network.
  • the term neural network used in the present disclosure may be used interchangeably with the term “artificial neural network” (ANN), in particular with the term ANN to be trained or trained ANN.
  • ANN artificial neural network
  • an artificial neural network refers to a computing system employing interconnected nodes, which in particular refer to functional units, which make use of machine learning.
  • Artificial neural networks are a subset of machine learning, also known as deep learning, and use unstructured data sets and hidden layers. Similar to normal machine learning, data sets for establishing the neural network are usually divided into the training data sets and validation data sets.
  • the interconnected nodes which may also be referred to as artificial neurons, receive signals, process these signals and can signal nodes connected to them based on the received and processed signals.
  • the signals transmitted to nodes are referred to as inputs and usually represent a real number.
  • the output of each node is calculated by a non-linear function, also referred to as activation function, typically based on the sum of its inputs.
  • the nodes and their connections typically have weights that are adjusted during a training phase based on the training data sets. Then for validating the trained ANN, the validation data sets are used allowing to determine accuracy of the trained ANN.
  • nodes are aggregated into layers as will be described hereinafter.
  • signals travel from a first layer, i.e. the input layer, via hidden layers to a last layer, i.e. the output layer, possibly after traversing the layers multiple times.
  • an ANN to be trained in particular an untrained ANN, may be provided and thereafter properly trained based on the training data sets.
  • the method may further comprise a step of providing a neural network to be trained, that preferably is an untrained ANN.
  • the neural network in particular the neural network to be trained and/or the trained neural network, may comprise an input layer having a plurality of input nodes, each of which receives one input parameter.
  • each one of the plurality of input nodes preferably receives one input parameter.
  • the received input parameters preferably differ among the input nodes.
  • the neural network in particular the neural network to be trained and/or the trained neural network, may comprise at least one hidden layer, in particular only one hidden layer, having a plurality of hidden nodes, in particular four or more hidden nodes.
  • the neural network in particular the neural network to be trained and/or the trained neural network, may comprise a single hidden layer with four hidden nodes.
  • the hidden nodes may use tan h as an activation function. Tan h transforms values to be between -1 and 1 .
  • sigmoid may be used as the activation function of the hidden nodes.
  • the neural network in particular the neural network to be trained and/or the trained neural network, may comprise an output layer having at least one output node for providing output parameters.
  • the output layer may have a plurality of output nodes, each of which determines and outputs a single output parameter.
  • the output parameters provided by the output layer may be used to determine the concentration-dependent viscosity.
  • the output parameters provided by the output layer may indicate or represent the concentration-dependent viscosity.
  • the output layer may be configured to provide at least one output parameter which indicates at least one of constant A and constant B of the above equation (1 ).
  • the output layer may be configured to provide at least one value as an output parameter which is used as constant A or constant B of the above equation (1 ).
  • the output layer may be configured to provide a first output parameter by a first output node which is used as constant A and a second output parameter by a second output node which is used as constant B to determine the concentrationdependent viscosity represented by the function specified in above equation (1).
  • the neural network receives a plurality of input parameters based on which the at least one output parameter is calculated.
  • the training data sets include or consist of the input parameters and the at least one output parameter.
  • the training data sets are used to implement or establish the neural network, in particular by adapting the weights of the neural network’s nodes and their connections in dependence on the training data sets.
  • the neural network to be trained becomes the trained neural network, i.e. the neural network provided by the above method.
  • the training data sets are further specified which allow to train and thus provide an ANN enabling the prediction of a viscosity of a protein solution at a high accuracy.
  • Each set of the training data set preferably has the same number and type of parameters. Further, as set forth above, each training data set is associated to a specific protein solution, in particular to a specific protein.
  • the training data set may be subdivided into input parameters, i.e. parameters representing inputs to the neural network, and output parameters, i.e. parameters representing outputs of the neural network.
  • the input parameters are indicative of the specific protein solution that the corresponding training data set is associated to. As such, the input parameters may indicate properties of the specific protein solution or of the protein included therein.
  • the output parameters are indicative of the concentration-dependent viscosity of the specific protein solution. As such, the output parameters may represent or indicate the concentration-dependent viscosity of the specific protein solution.
  • the input parameters included in the training data sets are indicative of protein-protein-interactions in the associated protein solution.
  • protein-protein-interactions are influenced by the protein’s amino acid sequence and resulting three-dimensional structure with charged or hydrophobic patches.
  • solution conditions may influence protein-protein-interactions by modulating size of charged patches via pH, shielding of charged patches via short- ranged electrostatic interaction using salts, buffer substances, amino acids or other charged excipients.
  • the plurality of input parameters include at least one of: i) experimental data obtained by detecting parameters from a provided protein solution, preferably wherein the experimental data expresses parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; ii) computational data, preferably calculated from the primary amino acid sequence of said protein at a pH of 5.0 to 7.0, preferably wherein the computational data expresses parameters selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), Fv symmetry parameter, Fv hydrophobicity, Vh-charge, Vl-charge, hinge-charge and hydrophobic solvent accessible surface area; and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, and combinations thereof.
  • the experimental data expresses parameters selected from protein hydrophobicity
  • the plurality of input parameters may include at least one selected from the group consisting of said i) experimental data, said ii) computational data and said iii) in silico data.
  • the computational data and the in silico data may be data calculated in dependence on the primary amino acid sequence of the protein included in the protein solution.
  • the i) experimental data are preferably parameters selected from at least one of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof.
  • the i) experimental data may comprise at least one parameter selected from the group consisting of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD) and second virial coefficient (A2). More preferably, the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2).
  • the ii) computational data are preferably parameters selected from at least one of the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof.
  • the ii) computational data may comprise at least one parameter selected from at least one of the protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge). More preferably, the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge).
  • the iii) in silico data are preferably parameters selected from at least one of hydrophobic and charged patch sizes such as score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. More preferably, the iii) in silico data consists of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • the iii) in silico data comprises at least one parameter selected from the group consisting of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • DLS dynamic light scattering
  • Virial coefficients may be determined in various ways, such as static light scattering (SLS), in this case the virial coefficients are commonly denoted with A2 for the second virial coefficient and A3 for the third virial coefficient, or by measurement of the osmolarity, then the virial coefficients are commonly denoted with B2 for the second virial coefficient and B3 for the third virial coefficient.
  • SLS static light scattering
  • a suitable method may exemplarily be via static light scattering (SLS), e.g. following the below further outlined procedure.
  • SLS static light scattering
  • a suitable method may exemplarily be via hydrophobic interaction chromatography (HIC), e.g. following the below further outlined procedure.
  • HIC hydrophobic interaction chromatography
  • the respective parameters such as Fv-charge are calculated at a pH of 5 to 8, or of 5 to 7.5, or of 5.5 to 6.5, or of about 5.5, or of about 6.0, or of about 6.5, preferably at pH 6.0.
  • the person skilled in the art may revert to commonly known platforms such as Prot pi
  • the isoelectric point (pl) can be calculated "manually" from the pKa values of the amino acid residues of the primary sequence.
  • Methods for determining the in silico data are known to the person skilled in the art. Suitably, they may be determined using software BioLuminate (version 3.80; Schrodinger, LLC, New York, NY), e.g. following the below further outlined procedure.
  • the neural network in particular the neural network to be trained and/or the trained neural network, may receive input parameters selected from all three data groups i) to iii) above, but a reliable prediction of the concentration dependent viscosity is also possible if the neural network receives input parameters from only two of the three data groups, such as the neural network receives input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii).
  • the neural network receives the plurality of input parameters based on which the at least one output parameter is calculated.
  • the output parameters may represent or indicate the concentration-dependent viscosity of the specific protein solution.
  • the output parameter may indicate a viscosity value cs associating a viscosity of the protein solution to a specific protein concentration of the protein solution.
  • the at least one output parameter may indicate a viscosity of the protein solution at more than one protein concentration.
  • the at least one output parameter may indicate the function f v represented by above equation (1 ). More specifically, the at least one output parameter may indicate at least one, preferably both, of constants A and B of the above equation (1 ).
  • an example is provided how to determine the at least one output parameter included in the training data sets.
  • different protein solutions referred to as protein solution samples
  • concentration-dependent viscosity data may be generated, e.g., by using a VROC viscometer (Rheosense) for each concentration.
  • VROC viscometer Heosense
  • a plurality of value pairs may be provided, each of which associates a viscosity of the protein solution to a protein concentration in the protein solution.
  • measured viscosity data in particular experimentally measured viscosity data, in a next step, may be processed based on the following equations (3) to (5).
  • equation (3) may be used to calculate the relative viscosity of the different protein solution samples. For doing so, at first, the viscosity of the buffer comprised in the protein solution may be determined, preferably at a defined temperature. Knowing the viscosity of the buffer, the relative viscosity of the protein solution samples may then be calculated based on the measured viscosity data according to equation (3). By doing so, value pairs for concentration and relative viscosity are determined for each protein solution sample.
  • the viscosity descriptors A and B for each protein may then be determined based on equation (4) and/or (5) by applying, e.g., curve fitting techniques, such as least square fitting.
  • a set of a number (num-conc) of protein solutions having different protein concentrations are provided and each of their sample viscosity q is determined; with num-conc being at least two, preferably 2 to 12, such as 4, 5, 6, 7, 8, 9, 10 11 or 12, a typical value is 6; these values pairs of protein concentration and sample viscosity serve for determining the at least one associated output parameter of the training and validation data set as described above.
  • the different concentrations in said set of protein solutions can for example range from a minimum concentration (mini-conc) of 10 mg/mL to a maximum concentration (max-conc) of 320 mg/mL, or from 10 to 280 mg/mL, or from 10 to 240 mg/mL, or from 10 to 210 mg/mL; a typical value for min-conc is 30 mg/mL and for max-conc is 180 mg/mL.
  • the set of protein solutions contains 3 or more protein solutions, each with a different protein concentration, the difference of every two concentrations adjacent to each other, that is the difference of two subsequent concentrations, is the same, so
  • the set of protein solutions for each protein in the set of proteins contains 6 protein solutions with the 6 different concentrations of 30 mg/mL, 60 mg/mL, 90 mg/mL, 120 mg/mL, 150 mg/mL and 180 mg/mL, so max cone is 180 mg/mL, min cone is 30 mg/mL, num cone is 6, diff cone is 30 mg/mL, n integer from 0 to 5.
  • a method according to the present invention may further comprise a step of validating the trained ANN, i.e. after the step of configuring, also called training herein, the neural network has been performed.
  • a plurality of validation data sets may be provided.
  • the method may comprise a further of validating the trained neural network based on a plurality of validation data sets, each of which comprises a plurality of input parameters and at least one associated output parameter.
  • the validation data sets may have the same or different number and type of parameters as the training data sets, but it is preferred that the validation data sets use the same type of parameters as the training data sets. As such, each validation data set may be divided into input parameters and associated output parameters, likewise to the training data sets.
  • the validation data sets may be used to validate proper implementation or training of the neural network.
  • the trained neural network may receive the input parameters of a validation data set to compute the corresponding at least one associated output parameter.
  • the computed output parameter may then be compared to the at least one associated output parameter included in the validation data set. This may be performed for each set of the validation data set to determine whether proper configuring of the neural network is established.
  • the trained neural network may be stored on a data medium, e.g., a non-transitory storage medium configured to store digital data of the sort that will be apparent to one of skill in the art in view of the present disclosure.
  • a further aspect of the invention comprises a computer-implemented method for determining or predicting a concentration-dependent viscosity of a protein solution by using a computer-implemented neural network, in particular a trained neural network, more particularly a trained neural network as described above.
  • the further aspect of the invention may comprise the provision and use of a computer-implemented method for determining or predicting a concentrationdependent viscosity of a protein solution by using a computer-implemented neural network.
  • the neural network in particular the trained neural network, is configured for predicting a concentration-dependent viscosity of the protein solution in dependence on a plurality of input parameters associated to said protein solution, preferably wherein the input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof; ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • HIC hydrophobic interaction chromatography
  • kD diffusion interaction parameter
  • the input parameters may include at least one selected from the group consisting of said i) experimental data, said ii) computational data and said iii) in silico data.
  • the computer-implemented method of the present invention makes use of the neural network, in particular of the trained neural network described above, i.e. which is provided by the above described method.
  • the technical features described herein in connection with the method for providing a computer- implemented neural network in particular in connection with the provision and use of a computer-implemented neural network, may thus also apply and refer to the computer-implemented method for predicting a concentration-dependent viscosity of a protein solution, and vice versa.
  • the methods of the present invention may be used to predict a viscosity of any suitable protein solution.
  • the protein solution may comprise a therapeutic protein, preferably an antibody, more preferably a monoclonal antibody (mAb), such as a monoclonal antibody of lgG1 or lgG2 subtype.
  • mAb monoclonal antibody
  • the protein solution may be an aqueous medium including water, preferably a buffer, more preferably a Histidine-HCI buffer.
  • the neural network in particular the trained neural network, used in the methods of the present invention is configured for determining or predicting the concentrationdependent viscosity in dependence on a plurality of input parameters.
  • the concentration-dependent viscosity can by expressed and/or determined as a viscosity value cs of the protein solution at a selected protein concentration.
  • the concentration-dependent viscosity can by expressed and/or determined as a function f v associating viscosity of the protein solution to protein concentration.
  • the concentrationdependent viscosity may be indicative of or may be represented by the function f v depicted by above equation (1).
  • the neural network may be configured to calculate at least one of constants A and B of the function f v .
  • the neural network receives the input parameters and, based thereupon, determines the concentration-dependent viscosity of the protein solution.
  • the input parameters are preferably associated to the protein solution, the viscosity of which is to be predicted.
  • the input parameters preferably refer to the protein included in the protein solution. More preferably, the input parameters are indicative of protein-protein-interactions in the protein solution.
  • the input parameters may correspond to those input parameters included in the training data sets as described above.
  • the technical features described above in connection with parameters contained in the training data sets may apply and refer likewise to the input parameters to be received by the neural network for calculating the concentration-dependent viscosity, and vice versa.
  • the neural network receives the input parameters as described above.
  • the input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof.
  • the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2);
  • ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof.
  • the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge); and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • the input parameters may include at least one selected from the group consisting of said i) experimental data, ii) computational data, and iii) in silico data.
  • the input parameters may be selected from all three data groups i) to iii) above, but a reliable prediction of the concentration dependent viscosity is also possible if the neural network receives input parameters from only two of the three data groups, such as the neural network receives input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii).
  • a method for determining a concentration-dependent viscosity of a protein solution.
  • the method comprises a step of predicting a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the concentration-dependent viscosity with a target viscosity, in particular to determine a viscosity characteristic of the protein solution.
  • the proposed method may make use of the above described neural network and the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution.
  • technical features described above in particular in connection with the neural network and its use for predicting a concentration-dependent viscosity of a protein solution, may thus also apply and refer to the method for determining a concentration-dependent viscosity of a protein solution, and vice versa.
  • the proposed method may be used for assessing whether the protein solution is subjected to a high or too high viscosity, in particular at a predefined protein concentration.
  • the step of comparing the concentration-dependent viscosity with the target viscosity may be performed to determine whether the protein solution is subjected to a high or too high viscosity, in particular whether it exceeds the target viscosity, in particular at a predefined protein concentration.
  • a viscosity of a protein solution may be considered too high if it renders the protein solution unsuited as a drug product for the reasons given herein above, i.e. , if it’s production produces high costs, due to high loss and low recovery in purification, difficult manufacturing or filling, and ultimately low administrability due to need for high injection forces and slow administration with potential sensation of pain.
  • solutions with a dynamic viscosity above 15 to 30 mPa*s or even above 15 to 20 mPa*s are considered problematic, and hence “too high”.
  • the method may be used for determining suitability as a drug product of a protein solution, in particular of a protein solution to be used or prepared as a drug product. For doing so, the step of comparing the concentrationdependent viscosity with the target viscosity is performed to determine suitability as a drug product of the protein solution.
  • the suggested method makes it possible to easily and reliably take into account the concentration-dependent viscosity of a protein solution dependent on the concentration of the protein in the protein solution, in particular at early stages of a development process. In this way, an improved validation of a protein solution to be used as a drug product is enabled at early stages, which has an impact on the drug product to be produced.
  • the proposed method may be used for assessing suitability of a protein solution to be used in any drug products which comprise or are constituted by a protein solution.
  • the method comprises the step of comparing the determined concentration-dependent viscosity with a target viscosity.
  • target viscosity refers to a desired or predefined viscosity of the protein solution; or the target viscosity may represent an upper threshold for the viscosity of the protein solution.
  • the target viscosity may indicate a maximum value for the viscosity of the protein solution at a given concentration.
  • the target viscosity may indicate a maximum value for the viscosity of the protein solution between 15 mPa*s to 30 mPa*s, or between 15 mPa*s to 25 mPa*s, or between 15 mPa*s to 20 mPa*s, for example 15 mPa*s or 18 mPa*s or 20 mPa*s or 25 mPa*s.
  • a viscosity above the maximum value may be regarded as problematic when using the protein solution as a drug product.
  • this step may be performed to determine suitability of the protein solution as a drug product.
  • the concentration-dependent viscosity may be assessed whether the protein solution has a concentration-dependent viscosity which is favorable when being used as a drug product.
  • Suitability of the protein solution may be determined when the concentration-dependent viscosity complies with or is below of the target viscosity, i.e. lies within a desired viscosity range even if the protein is used at a high concentration, e.g., of > 50 mg/mL, or > 60 mg/mL, > 70 mg/mL, or > 80 mg/mL, or > 90 mg/mL, or > 100 mg/mL.
  • application-specific viscosity of the protein solution may be forecasted to assess whether the protein solution is suitable for being used as a drug product in the clinical therapy.
  • a protein concentration of the protein solution is taken into account.
  • the concentration-dependent viscosity and the target viscosity may be compared at a specific protein concentration or within a protein concentration range.
  • a protein concentration of the protein solution may be determined.
  • a maximum protein concentration may be determined which may refer to a protein concentration the protein solution is expected to not exceed when being used as a drug product.
  • the maximum protein concentration may be in the range from 100 mg/mL to 220 mg/mL or from 120 mg/mL to 180 mg/mL, for example 120 mg/mL or 150 mg/mL or 180 mg/mL.
  • the viscosity of the protein solution at the maximum protein concentration may be determined and compared to the target viscosity.
  • a realistic protein concentration range of the protein in the protein solution for clinical therapy may be determined.
  • the protein concentration range may span from 20 mg/mL or 50 mg/mL or 100 mg/mL to the maximum protein concentration. Then, it may be determined whether the concentration-dependent viscosity of protein solution exceeds the target viscosity within said protein concentration range.
  • a method for providing (or preparing) a drug product comprising a protein solution.
  • the method comprises a step of predicting (or determining) a concentration-dependent viscosity of a protein solution by using a computer-implemented neural network; a step of comparing the concentration-dependent viscosity with a target viscosity to determine suitability of the protein solution as a drug product; and an optional step of preparing the drug product if suitability of the protein solution is determined.
  • the proposed method may make use of the above described neural network, the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution, and the above described method for determining a concentration-dependent viscosity.
  • technical features described above may thus also apply and refer to the method for providing, and optionally preparing, a drug product, and vice versa.
  • the step of preparing the drug product may be performed. If suitability of the protein solution as a drug product is not determined, then, the protein solution may be adapted or changed and the step of determining a concentration-dependent viscosity of a protein solution and the step of comparing the determined concentration-dependent viscosity with a target viscosity may be performed again based on the adapted or changed protein solution.
  • a method for determining a suitable concentration of a protein in a protein solution of a drug product; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine the upper limit of the concentration of the protein in the protein solution which is still acceptable for a drug product.
  • a method for identifying an upper limit of the concentration of the protein in a protein solution which should not be exceeded in order to avoid unacceptable high viscosity of the protein solution; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine if the protein in the protein solution induces above a certain concentration a viscosity of the protein solution which is not acceptable for a drug product.
  • a method for identifying a protein having an unacceptable concentration-dependent viscosity in solution; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine if the protein in the protein solution induces above a certain concentration a viscosity of the protein solution which is not acceptable for a drug product.
  • the present invention provides a use of a computer-implemented neural network for facilitating the preparation of a drug product, wherein the neural network is configured to predict a concentration-dependent viscosity of a protein solution which is used to determine suitability of the protein solution to be prepared as the drug product.
  • the invention provides a method for determining the suitability of a protein solution as a drug product, comprising: a step of experimentally detecting parameters from a provided protein solution, preferably wherein the obtained experimental data expresses parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine suitability of the protein solution as a drug product.
  • a step of experimentally detecting parameters from a provided protein solution preferably wherein the obtained experimental data expresses parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3)
  • the invention provides a method for providing a trained artificial neural network (ANN) for determining a concentration-dependent viscosity of a protein solution by training the ANN with input parameters describing proteins contained in a set of proteins.
  • ANN artificial neural network
  • the invention provides a method for determining (or predicting) a concentration-dependent viscosity of a solution of a protein by providing the trained ANN with input parameters being indicative of or of the protein and letting the ANN calculate an output parameter indicating or being the concentration-dependent viscosity.
  • the present invention provides an apparatus for predicting a concentration-dependent viscosity of an experimental protein solution, the apparatus comprising: an experimental protein solution; a computer component comprising a neural network configured to predict the viscosity of the experimental protein solution: wherein the neural network is trained using a plurality of predetermined data sets; wherein each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical characteristic of a predefined protein solution, and (ii) at least one output parameter indicative of a concentration-dependent viscosity of the predefined protein solution; and wherein the plurality of input parameters comprise at least one of experimentally-derived data, computationally-derived data, and in silico-derived data.
  • the proposed apparatus may make use of the above described neural network and the above described computer-implemented method for predicting a concentration-dependent viscosity of a protein solution.
  • technical features described above, in particular in connection with the neural network, more particularly in connection with the above-described trained neural network, and its use for predicting a concentration-dependent viscosity of a protein solution may thus also apply and refer to the proposed apparatus, and vice versa.
  • the present invention provides a system for determining a concentration-dependent viscosity of a protein solution, the system comprising: a protein solution; and a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution depends on a plurality of input parameters (IP) associated with the protein solution, and further wherein the input parameters (IP) include at least one from the group consisting of (i) experimental data; (ii) computational data; and (iii) in silico data.
  • IP input parameters
  • the proposed system may make use of the above described neural network and the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution.
  • technical features described above in particular in connection with the neural network, more particularly in connection with the above-described trained neural network, and its use for predicting a concentration-dependent viscosity of a protein solution, may thus also apply and refer to the proposed system, and vice versa.
  • Figure 1 a shows an exemplary computing component that may be used to implement various features of the embodiments of the present invention
  • Figure 1 b shows a flow diagram illustrating a method for providing a drug product according to an embodiment of the present invention
  • Figure 2 schematically shows a computer-implemented neural network used in the method depicted in Figure 1 ;
  • Figure 3 schematically shows a further computer-implemented neural network used in the method depicted in Figure 1 ;
  • Figure 4 illustrates training data sets and validation data sets used for training and implementing a neural network used in the method depicted in Figure 1 ;
  • Figure 5 depicts a diagram illustrating a comparison between viscosity values calculated by a suggested neural network and measured viscosity values
  • Figure 7 shows the difference of predicted viscosity values, calculated from the predicted viscosity curves using the predicted values for the intercept and slope, to measured viscosity values at the same concentration (the percentage of difference is presented with bars and the absolute difference is presented with points);
  • Figure 8 illustrates the input variables and setup of the artificial neural network as exemplified (inputs are combined in one layer with 4 hidden nodes with tan h activation functions and target determining viscosity descriptors, the intercept A or slope B);
  • Figures 9a and 9b depict examples of the training and validation data (9a: model xx-A training and validation data, and 9b: model xx-B training and validation data).
  • Figure 1a is a schematic view showing an exemplary computing component 2 that may be used to implement various features of the embodiments of the present invention.
  • computing component 2 generally comprises a bus 3 connected to (i) a processor 4, (ii) memory 5 for storing information and instructions to be executed by processor 4 (e.g., random access memory (RAM) or other dynamic memory, read-only memory (ROM) or any other static storage device for storing static information and instructions for processor 4, etc.), (iii) a storage device 6 comprising a non-transitory medium (e.g., a hard disk drive, solid state disk drive, optical storage, etc.) for the long-term storage of digital information (e.g., computer software, digital data, etc.), and (iv) a communications interface 7 for allowing software and data to be transferred between computing component 2 and external devices via a communications channel 8, as will be apparent to one of skill in the art in view of the present disclosure.
  • processor 4 e.g., random access memory (RAM) or other dynamic memory, read-only memory (ROM) or any other static storage device for storing static information and instructions for processor 4, etc.
  • RAM random access memory
  • ROM read-
  • the computing component 2 may be part of or may constitute an apparatus for predicting a concentration-dependent viscosity of an experimental protein solution.
  • Figure 1 b depicts a method for providing and optionally preparing a drug product comprising a protein solution.
  • a computer-implemented neural network 10 is provided which is configured for predicting a concentration-dependent viscosity of a protein solution in dependence on a plurality of input parameters associated to said protein solution.
  • Step SO represents a sub-method, i.e. a method as such, included in the method for preparing a drug product.
  • computer- implemented neural network 10 which preferably is a trained neural network, may be implemented using the aforementioned computing component 2, or using any appropriate computing system that will be apparent to one of skill in the art in view of the present disclosure.
  • the protein solution comprises a protein, preferably a therapeutic protein such as an antibody, which may include monoclonal antibodies, polyclonal antibodies, whole antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments such as Fab, Fab', F(ab')2, fragments including hybrid fragments of any immunoglobulin or any natural, synthetic or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex.
  • the therapeutic protein is a monoclonal antibody (mAb), preferably monoclonal antibodies of lgG1 or lgG2 subtype.
  • the protein solution comprises an aqueous medium including water, preferably a buffer, more preferably a Histidine-HCI buffer.
  • a neural network 10 to be trained in particular an untrained neural network
  • the neural network 10 to be trained may be configured to receive the plurality of input parameters and based thereupon to compute at least one output parameter.
  • the neural network 10 is configured to calculate the concentration-dependent viscosity of the protein solution in dependence on the input parameters.
  • the concentrationdependent viscosity associates at least one protein concentration of the protein solution to a corresponding viscosity, in particular dynamic viscosity, of the protein solution.
  • the neural network 10 in particular the neural network to be trained and the trained neural network, is configured to compute a viscosity value of the protein solution at a selected protein concentration c s .
  • the selected protein concentration c s may refer to a maximum protein concentration which is expected to be relevant for the drug product.
  • the maximum protein concentration may be in the range from 100 mg/mL to 220 mg/mL ,or from 120 mg/mL to 180 mg/mL, for example 120 mg/mL or 150 mg/mL or 180 mg/mL.
  • the concentrationdependent viscosity is expressed as cs , meaning the viscosity value of the protein solution at the selected protein concentration c s .
  • the neural network 10 in particular the neural network to be trained and the trained neural network, is configured to compute a mathematical function associating viscosity values of the protein solution to protein concentrations of the protein solution.
  • the concentration-dependent viscosity is represented by the function ⁇ as specified in the above equation (1 ).
  • the concentration-dependent viscosity is provided in the form of the function f v .
  • constants A and B of the above equation (1 ) are calculated by the neural network 10.
  • the neural network 10 to be trained will be described in more detail with reference to the embodiments depicted in Fig. 2 and Fig. 3.
  • Fig. 2 depicts an embodiment of the neural network 10, in particular of the neural network to be trained and of the trained neural network, used in the method depicted in Fig. 1.
  • the neural network 10 comprises an input layer 12 having a plurality of input nodes I Ni -I Ni, wherein the index "/" refers to a positive integer number.
  • the input layer 10 comprises / different input nodes IN.
  • each input node INi-IN n receives a corresponding input parameter IPi-IPi.
  • the neural network 10 comprises a hidden layer 14 having a plurality of hidden nodes HNi-HNj, wherein the index ' ' refers to a positive integer number, preferably j is 4 and the neural network has four hidden nodes HN1-HN4, each of which is connected to each input node INi-INi.
  • the hidden nodes HNi-HNj Use tan h as an activation function. Alternatively, sigmoid may be used as the activation function.
  • the neural network 10 comprises an output layer 16 having one single output node ON which is connected to all hidden nodes HNi-HNj.
  • the output node ON provides a single output parameter in the form of the parameter cs which is the viscosity of the protein solution at the selected protein concentration c s .
  • the parameter cs constitutes the concentration-dependent viscosity.
  • Fig. 3 depicts the further embodiment of the neural network 10 according to the present invention, in particular of the neural network to be trained and of the trained neural network.
  • the output layer 16 of the neural network 10 is provided with two output nodes ON1 and ON2 providing different output parameters.
  • the first output node ON1 provides the output parameter OP1, which is A
  • a second output node ON2 provides the output parameter OP2, which is B.
  • These two parameters represent the constants of function f v .
  • the function f v represented by the constants A and B constitutes the concentration-dependent viscosity.
  • a plurality of training data sets ST is provided, each of which is associated to a specific protein solution and comprises a plurality of input parameters IP being indicative of the specific protein solution and at least one associated output parameter OP being indicative of the concentration-dependent viscosity of the specific protein solution.
  • validation data sets Sv are provided which have the same number and types of input parameters and output parameters.
  • Fig. 4 depicts a table generically illustrating the training data sets ST and validation data sets Sv used for properly implementing the neural network 10.
  • Each training data set ST and each validation data set Sv comprises input parameters I Pi -IPi representing values to be received by the input nodes I Ni -INi and the two output parameters OPi and OP2 representing values to be provided by the two output nodes ON1 and ON2.
  • a number of m different training data sets ST are provided based on which the computational model underlying the neural network 10 is adapted during a training phase in the next sub-step S0.3.
  • the parameters "m" indicates a positive integer number.
  • the neural network 10 in particular the neural network to be trained, is trained to provide the computer-implemented neural network 10, in particular the trained neural network, which is suitable and configured for predicting the concentration-dependent viscosity.
  • all input parameters IP and output parameters OP of the training data sets ST are fed to the neural network 10 to adapt the computational model underlying the neural network 10, in particular to adapt weights of the nodes and their connections.
  • a number of n different validation data sets Sv are used to validate the neural network 10, in particular the trained neural network.
  • the parameters "n" indicates a positive integer number.
  • each one of the validation data sets Sv differs from each one of the training data sets ST.
  • the input parameters associated to the validation data sets Sv are fed to the neural network 10 which, based thereupon, calculates the respective output parameters for each validation data set Sv.
  • the calculated output parameters are then compared to the output parameters included in the validation data sets Sv.
  • 20 or more, e.g. 24 different training data sets ST may be used.
  • two or more, e.g. three, different validation data sets Sv may be used.
  • the neural model may comprise a number of / different input nodes and accordingly may receive / different input parameters IP per data set.
  • the input parameters IP preferably are indicative of protein-protein-interactions in the protein solution associated thereto.
  • the input parameters IP may include different types of input parameters which can be categorized into i) experimental data, ii) computational data and iii) in silico data.
  • Experimental data may be obtained by detecting parameters from a provided protein solution.
  • the i) experimental data expresses parameters selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof. More preferably, the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2).
  • Computational data may be calculated from the primary sequence of said protein at a pH of 5.0 to 7.0.
  • the computational data expresses parameters selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof. More preferably, the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv- charge).
  • the descriptors derived from modelling are thus score pos/neg/hyd Fv total, size pos/neg/hyd Fv total, count pos/neg/hyd Fv total; and any combination thereof.
  • the iii) in silico data consists of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
  • the neural network 10 in particular of the trained neural network, is described which is equipped with 14 different input nodes IN1-IN14 and 14 different input parameters IP1-IP14. It will be obvious for a person skilled in the art that these input parameters only depict examples of a plurality of possibilities. Hence, the input parameters described hereinafter should not be understood to form a limitation. Thus, more or less input parameters may be used.
  • a neural network may make use of input parameters including at least one of input parameters IP1 to IP14.
  • input parameters IP1-IP3 are i) experimental data
  • input parameters IP4 and IP5 are ii) computational data
  • input parameters IPe to IP14 are iii) in silico data.
  • IP1 refers to HIC RT [min]
  • IP4 refers to pl
  • IPs refers to Fv charge at pH 6
  • IPe refers to Score pos Fv total
  • IP7 refers to Size pos Fv total
  • IPs refers to Count pos Fv total
  • IP9 refers to Score neg Fv total
  • IP10 refers to Size neg Fv total
  • IP11 refers to Count neg Fv total
  • IP12 refers to Score hyd Fv total
  • I Pi 3 refers to Size hyd Fv total; and IP14 refers to Count hyd Fv total.
  • the combination of i) experimental data consisting of input parameters IP1-IP3, ii) computational data consisting of input parameters IP4 and IPs and iii) in silico data consisting of input parameters IPs to IP14 is used.
  • the neural network 10 in another embodiment, only the combination of ii) computational data consisting of input parameters IP4 and IPs and iii) in silico data consisting of input parameters IPe to IP14; in this embodiment the neural network 10 is equipped with 11 different input nodes IN4-IN14 and 11 different input parameters I P4-I Pi 4.
  • the neural network 10 in particular the trained neural network, was used to model the viscosity of solutions of monoclonal antibodies.
  • Fig. 5 depicts a diagram in which output values provided by the neural network 10 to model the viscosity of solutions of monoclonal antibodies are compared to measured viscosity values.
  • the neural network 10 in particular the trained neural network, validly, and reliably, predicts the concentration-dependent viscosity of the protein solution, in particular by calculating a viscosity curve.
  • a protein solution is identified based on which the drug product may be produced.
  • step S2 a concentration-dependent viscosity, in particular a protein- concentration-dependent viscosity, is predicted by using a computer-implemented neural network 10, in particular by using a trained neural network.
  • step S2 represents a sub-method, i.e. a computer-implemented method as such for predicting a concentration-dependent viscosity of a protein solution.
  • the neural network 10 in particular the trained neural network, provided in step SO is used to compute the concentration-dependent viscosity of a protein solution.
  • the input parameters IP in particular at least one of, preferably all of IPi to IP14, associated to the protein solution identified in step S1 are determined and provided to the neural network 10 to, based thereupon, compute the at least one output parameter OP indicating the concentrationdependent viscosity.
  • Step S3 the concentration-dependent viscosity determined in step S2 is compared with a target viscosity to determine suitability of the protein solution as a drug product.
  • Steps S2 and S3 together represent a sub-method, i.e. a method as such for determining a concentration-dependent viscosity of a protein solution.
  • the target viscosity indicates a maximum value for the viscosity of the protein solution. Specifically, the target viscosity indicates a maximum value of between 15 mPa*s to 30 mPa*s or of between 15 mPa*s to 25 mPa*s or between 15 mPa*s to 20 mPa*s, for example 15 mPa*s or 18 mPa*s or 20 mPa*s or 25 mPa*s.
  • suitability of the protein solution is determined if the concentration-dependent viscosity, in particular at the selected protein concentration c s or for a selected protein concentration range, for example between 20 mg/mL to the selected protein concentration c s , does not exceed the target viscosity, i.e. the threshold or maximum value of 15 mPa*s or 20 mPa*s. However, if the viscosity of the protein solution exceeds the target viscosity, suitability of the protein solution is not determined (i.e., the protein solution is determined to not be suitable as a drug product, or the protein solution is determined to potentially not be suitable as a drug product).
  • step S3 the parameter cs is compared to the target viscosity. If the parameter cs does not exceed the target viscosity, suitability of the protein solution as a drug product is determined (i.e., the protein solution is determined to be suitable as a drug product, or the protein solution is determined to be potentially suitable as a drug product).
  • step S3 the derived function f v is used to determine whether the viscosity of the protein solution exceeds the target viscosity in a selected protein concentration range, for example spanning between 20 mg/mL to the selected protein concentration c s .
  • a viscosity cs of the protein solution at the selected protein concentration c s may be determined based on function f v , before determining whether the viscosity cs , i.e. the viscosity at the selected protein concentration c s , exceeds the target viscosity.
  • suitability of the protein solution as a drug product is determined (i.e., the protein solution is determined to be suitable as a drug product, or the protein solution is determined to be potentially suitable as a drug product).
  • step S4 in Fig. 1 if suitability of the protein solution is determined in step S3, then the method may proceed to optional step S5 in which the drug product is prepared (or produced) based on the protein solution identified in step S1 , in particular at a desired protein concentration. If suitability of the protein solution is not determined in step S3, then the method may proceed to step S1 in which a new protein solution is identified, in particular by adapting or changing the initial protein solution, before performing steps S2 to S4 again.
  • step SO by screening potential protein solutions using the method described above (i.e., by screening potential protein solutions using the neural network 10 provided in step SO) it is possible to identify candidate protein solutions for use as a drug product without having to actually prepare and empirically test the potential protein solutions.
  • the foregoing in silico method for screening potential protein solutions permits a large number of candidate protein solutions to be screened without the necessity of empirical measurement of the qualities of every candidate protein solution.
  • the present invention solves the long- unmet need for a quick, high-throughput method for screening candidate protein solutions for use as a drug product that avoids the time and labor inherent in empirical (i.e., laboratory-based) screening.
  • the validation sets contained as data only the input variables of the remaining 7-5 mAbs, with the goal to predict their viscosity descriptors. These viscosity descriptors are derived from the linearization of the viscosity-concentration-curves of each mAb, the intercept A and the slope B.
  • the R 2 values of the created models show the interdependency of validation and training set (table 1 ; exemplary graphs are depicted in Figures 9a and 9b), as in several cases the quality of one set is excellent (R 2 > 0.99) while the second set is less good (R 2 ⁇ 0.95).
  • the highest quality model was achieved for the artificial neural network in which both variables (experimental and in silico) were used, with the lowest being R 2 > 0.92 (intercept A of training set).
  • the artificial neural network using only in silico-derived inputs achieves a minimally lower "worst" R 2 of > 0.90 (slope B of training set).
  • the artificial neural network containing only experimental data achieves the lowest R 2 of > 0.75, which is comparable to results published using linear correlations of kD to viscosity, and better than a linear correlation of A2 and kD to the intercept and slope parameter done for this data set.
  • the slope parameter B describes the steepness of the viscosity curve so the exponential increase of viscosity with the protein concentration, which is important to describe potentially ..problematic" mAbs.
  • Such ..problematic" mAbs may show a moderate viscosity at low protein concentration but experience substantial increase of viscosity above values usually regarded as acceptable (i.e. , values above about 15-25 mPa*s) for drug products. Consequently, the artificial neural network was used with all input variables in the following, as it provided the best prediction for the slope B.
  • mAbs 26 and 27 were not used in artificial neural network model creation but were kept separately for additional verification tests.
  • Table 1 shows a comparison of models created using different inputs predicting either the intercept (A) or slope (B) of the concentration-dependent viscosity curve for the mAbs. “x” indicates that a specific set of inputs was used for the model, while indicates that the set of input parameters was not used.
  • Experimental inputs refer to HIC retention time, kD and A2, while in silico input includes the data from the surface patch analysis which is the information of size, score and count of positive, negative and hydrophobic surface patches. Fv-charge and pl was included regardless of the selected inputs.
  • the quality parameters of the models shown were obtained from plotting the values of A or B that were calculated from the measured viscosity against the predicted values of the respective models.
  • models for categorical classification were created by training the models on whether the mAbs show a viscosity above a threshold of 15 mPa*s at a certain concentration. This was done for concentrations of 120, 150 and 180 mg/mL, where of the 25 mAbs used in the model creation at 120 mg/mL 3 mAbs, at 150 mg/mL 6 mAbs and at 180 mg/mL 15 mAbs exhibited a viscosity of above 15 mPa*s.
  • mAb 26 shows unproblematic behavior at 120 and 150 mg/mL, but exceeds 15 mPa*s at 180 mg/mL, which was correctly predicted by the models disclosed herein (No/No/Yes).
  • mAb 27 shows no problematic behavior at either of the concentrations which was again correctly predicted by the models presented herein (No/No/No).
  • Table 2 Confusion matrices of the categorical models using experimental and in silico data as inputs. Results of the validation sets are in brackets.
  • the percentage and absolute differences of predicted compared to measured viscosities of all mAbs is presented in Figure 7. Calculated average differences across all 27 mAbs between predicted and measured viscosity values are in table 3.
  • the average absolute difference, calculated to measured, in mPa*s is between 0.1 -4.1 mPa*s, with a gradual increase of difference with increasing protein concentration.
  • the relative % difference are between 8.2 - 26.2 %, but follow a curved function, with the highest differences at concentrations of 90-120 mg/mL and decreasing difference with decreasing and increasing protein concentration.
  • the present invention provides artificial neural networks and methods to predict and determine the viscosity of proteins such as mAbs in solutions.
  • the models can be used to predict a viscosity categorization, above or below a given threshold of e.g. 15 mPa*s, and to predict viscosity curves. Whilst already the categ. 15 mPa*s is good, the models for viscosity curve prediction show a high power and good viscosity curve forecast.
  • the use of a multitude of input variables, derived from experimental data as well as from computational data and in silico modeling, is advantageous. Monoclonal antibodies mAbs were obtained.
  • Double gene vectors containing the heavy and light chains were transfected into CHOK1 SV GS-KO cells and cultured under selection conditions as stable pooled cultures. Clarified supernatant was obtained by centrifugation followed by filter sterilization using 0.22 pm filters. Protein A chromatography was used for mAb purification. All proteins were concentrated to a final concentration of 10 mg/mL, and the buffer exchanged into the formulation buffer (protein solution) (20 mM histidine-HCI, pH 6.0) by tangential flow filtration. mAbs were of different subtypes IgG 1 or lgG2 (see tables 4 and 5 below).
  • hydrophobic surface properties of all mAbs were determined by hydrophobic interaction chromatography (HIC). Proteins were analyzed at 10 mg/mL in formulation buffer, 5 pL were injected on a ProPac Hic-10 column (ThermoScientific) and separated using a Waters HPLC system. The start condition of 95 % mobile phase A (1 M ammonium sulfate in 20 mM sodium phosphate pH 7.0) was linearly reduced over 39 min to 95 % mobile phase B (20 mM sodium phosphate pH 7.0). Flow rate was set to 1 mL/min at a column temperature of 24 °C. Dynamic and static light scattering
  • Dynamic light scattering (DLS) and static light scattering (SLS) measurements were performed on a DynaPro PlateReader III (with software Dynamics; Wyatt Technologies).
  • Stock solutions of the antibodies were filtered through 0.22 pm PVDF filters (Millex GV) and serial dilutions with seven concentrations from 10-2 mg/mL protein with the formulation buffer were prepared.
  • Samples were transferred to 384 well plates (Aurora) in triplicate and the plates were centrifuged at 750*g for 2 min to remove air bubbles. The temperature for the measurement was set at 25 °C. Laser power was set to 20 % and attenuation level to 0 %, 20 acguisitions of 5 s length were made for each well.
  • the diffusion interaction parameter kD was performed via DLS.
  • the mutual diffusion coefficient Dm (m2/s) was plotted against the protein concentration (g/mL) and kD was obtained from the slope of a linear fit.
  • the second virial coefficient A2 (mol*mL/g) was obtained from SLS measurements.
  • Calibration of the plates was performed using Dextran (Sigma) with a predetermined molecular weight of 36.9 ⁇ 0.1 kDa.
  • Solvent offsets were measured in triplicates for the formulation buffer.
  • the reciprocal molecular weight (mol/g) was plotted against the protein concentration (g/mL) and A2 was obtained from the slope of a linear fit.
  • the descriptors derived from modelling are thus score pos/neg/hyd Fv total, size pos/neg/hyd Fv total, count pos/neg/hyd Fv total.
  • variable heavy- and variable light-chains of each antibody were analyzed with the prot pi protein tool (https://www.protpi.ch/Calculator/ProteinTool).
  • the two chains were each defined as a subunit of the entire protein.
  • the set modifier for post translational modifications was global disulfide bridges for the cysteine residues in the Fv. Charge was calculated at pH 6.0. pl calculation
  • the pl was calculated "manually" from the pKa values of the amino acid residues of the primary sequence.
  • equation (3) was used to calculate the relative viscosity of the mAb samples. For doing so, at first, the viscosity of the buffer was measured to be 0.92 mPa*s at 25°C. Knowing the viscosity of the buffer, the relative viscosity of the mAb samples were then calculated based on the measured viscosity values according to equation (3). Accordingly, value pairs for concentration and relative viscosity were determined for each mAb sample.
  • the ANN model creation follows the approach illustrated in Figure 8. All input parameters (experimental data, calculated values from sequence, in silico-demed data, and viscosity descriptors) are listed in table 4 and 5. Each model is trained on a categorical response aiming to identify mAbs which show a viscosity value above the threshold value of 15 mPa*s at a distinct protein concentration.
  • the ANNs are generated using software JMP v.16.0.0 (SAS Institute Inc.).
  • the activation function used for all nodes is the tan h-function, which transforms values to be between -1 and 1 . For all models one hidden layer is sufficient with the number of nodes being four.
  • the data was split into a training and a validation set.
  • the method used was K-fold where the 25 mAb and their data sets are split into K sets.
  • Each of the K sets contained the data sets of all 25 mAb, but in each K set the two subsets training set and validation set was different, the mAb were distributed randomly in each K set to the two subsets training set and validation set.
  • Each of the K sets was used to validate the model fit on the rest of the data, fitting a total of K sets. The value of K was set to 5 for each model.
  • the distribution of the mAb to the K sets was also not correlated between the three models but was done randomly, so each model has a different K set.
  • the model reported by the JMP software is based on the best log likelihood.
  • the quality of the ANNs was determined using the coefficient of determination (R2), the Standard Square Error (SSE) and the Root- mean-square Error (RMSE) for training and validation datasets.
  • Table 4 Data included in artificial neural network modelling. Experimental inputs (retention time in hydrophobic interaction chromatography (HIC); diffusion interaction parameter kD; second virial coefficient A2) and computational inputs derived input from amino acid sequence (isoelectric point (pl) and Fv-charge).
  • Table 5 Data included in artificial neural network modelling. In silico simulation- derived input (scores and areas for positively and negatively charged as well as hydrophobic patches). The viscosity descriptors intercept A and slope B derived from linearization of viscosity data (original viscosity data in tables 6 and 7). Data are fed into an artificial neural network following the scheme in Figure 8.
  • Table 6 Viscosity raw data. For every mAb a dilution series in 20 mM histidine-HCI pH 6 buffer was prepared with target concentrations (target c) of 180, 150, 120 mg/mL. The table contains values of dynamic viscosity in mPa*s ⁇ SD of 10 measurements as well as the measured protein concentration in mg/mL.
  • Table 7 Viscosity raw data. For every mAb a dilution series in 20 mM histidine-HCI pH 6 buffer was prepared with target concentrations (target c) of 90, 60, 30 mg/mL. The table contains values of dynamic viscosity in mPa*s ⁇ SD of 10 measurements as well as the measured protein concentration in mg/mL.
  • the viscosity descriptors intercept A and slope B are calculated from the values shown in tables 6 and 7. For doing so, at first, viscosity of the buffer is determined which was measured to be 0.92 mPa*s. Then, the buffer viscosity is used to calculate the relative viscosity for each mAb and for each value pair for concentration and viscosity in tables 6 and 7 based on above equation (3). Thereafter, a least square fit (or any other mathematical procedure for finding a relation, in particular for finding a best-fitting curve, to given sets of points) is done with the 6 value pairs for concentration and viscosity (with the relative viscosity) for each mAb, thereby fitting the 6 value pairs to the exponential function of above equation (4). Then, the calculated function is linearized as shown in above equation (5) to provide the A and B values for each mAb which are included in table 5.

Landscapes

  • Engineering & Computer Science (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Computing Systems (AREA)
  • Medical Informatics (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Biophysics (AREA)
  • Software Systems (AREA)
  • Chemical & Material Sciences (AREA)
  • Molecular Biology (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Biotechnology (AREA)
  • Evolutionary Biology (AREA)
  • General Physics & Mathematics (AREA)
  • Bioethics (AREA)
  • Epidemiology (AREA)
  • Public Health (AREA)
  • Library & Information Science (AREA)
  • Biochemistry (AREA)
  • Medicinal Chemistry (AREA)
  • Investigating Or Analysing Biological Materials (AREA)
  • Peptides Or Proteins (AREA)

Abstract

The present invention refers to a method for providing a computer-implemented neural network configured for predicting a concentration-dependent viscosity of a protein solution. Further, the present invention refers to methods for predicting a concentration-dependent viscosity of a protein solution, for determining a concentration-dependent viscosity of a protein solution, and for providing a drug product comprising a protein solution, by using such a computer-implemented neural network.

Description

Protein Solutions
Technical Field
The present invention refers to a method for providing a computer-implemented neural network configured for predicting a concentration-dependent viscosity of a protein solution. Further, the present invention refers to methods for predicting a concentration-dependent viscosity of a protein solution, for determining a concentration-dependent viscosity of a protein solution, and for providing a drug product comprising a protein solution, by using such a computer-implemented neural network.
Technological Background
Therapeutic proteins, such as monoclonal antibodies (mAbs), have become an important factor in the treatment of a broad variety of diseases, such as cancer, immune-mediated disorders and infectious diseases. The most common route of administration of protein-based drug products (DPs) is the intravenous route, which however requires the patient to be hospitalized and the drug product to be administered by a healthcare professional. For patients with chronic diseases the need for repeated drug administration and thus hospitalization poses a significant burden and stress which may put the success of the intended therapy at risk. Subcutaneous (s.c.) injection allows patients to self-administer protein-based drug products, such as drug products based on monoclonal antibodies, by use of prefilled syringes, auto-injectors or other delivery devices, and by this often increasing quality of life and compliance for patients with chronic conditions.
However, there are certain limitations with subcutaneous administration. Usually, a single injection volume is limited to about less than 2 mL, determined by the available subcutaneous space and sensation of tolerable pain by the patient. Due to this volume limitation, the use of drug products with high drug substance concentrations are required.
In general, mAbs have a high specificity but they also require considerable therapeutic dosages. This consequently results in high concentrations of drug products (DP) comprising solutions of proteins such as mAbs, often exceeding 100 mg/mL protein in solution for s.c. administration.
With increasing concentration, however, inter-molecular distances are reduced and proteins may interact with one another. These protein-protein-interactions (PPIs) then exponentially influence and determine the proteins’ (e.g. mAbs') solubility, aggregation, and also viscosity. PPIs are inter alia determined by a protein’s (e.g. mAb's) primary amino acid sequence and the resulting three-dimensional structure with charged or hydrophobic patches. Solution conditions can also influence protein-protein-interactions, e.g., by modulating size of charged patches via pH, shielding of charged patches via short-ranged electrostatic interaction using salts, buffer substances, amino acids or other charged excipients. Arginine is a common excipient tested for viscosity reduction, its dual mode of action being both the shielding of charged as well as of hydrophobic patches. 17 of 34 FDA approved drug products with high mAb concentration use salts or amino acids as excipients, likely with the aim to reduce protein-protein-interactions and thus to lower mAb solution viscosity. More explorative excipients with demonstrated potential for viscosity reduction, but no application in commercialized drug products, mostly due to lack of approval as excipients for parenteral administration or concerns on toxicity, are poly-l-glutamic acid, caffeine, hydrophobic salts, or amino acid derivates.
Highly viscous solutions can be a major roadblock in the development of proteinbased drug products. Disadvantages may be high costs, due to high loss and low recovery in purification, difficult manufacturing or filling, and ultimately low administrability due to need for high injection forces and slow administration with potential sensation of pain. In general, solutions with a dynamic viscosity above about 15 to 30 mPa*s or even above about 15 to 20 mPa*s are considered problematic. The desire to develop a high concentration formulation may not be apparent a priori for a new molecule, it may appear also as a consequence of need for unexpectedly high doses or change of target route of administration. Clinical trials are usually started in phase 1 with lower concentrated (<50 mg/mL protein in solution) formulations, whereas high protein concentrations (>100 mg/mL protein in solution) are typically explored in later phases once safety and efficacious dose levels are established.
To evaluate CMC (chemistry, manufacturing and control) issues related to a drug product later in development when dose ranges are established, it is essential to forecast the viscosity of a new protein, in particular a mAb, at high protein concentrations during an early development stage.
For high concentration mAb solutions, extensive work has been done in the past two decades to understand the factors resulting in high viscosity.
In early development, often multiple candidates are available in small quantities and need to be tested in pre-formulation studies for their stability and solubility. A substantial amount of work has been done in the last years to use experimental data from low concentration experiments to predict the viscosity of high concentration solutions.
To forecast the viscosity of a protein solution being used as a drug product in the clinical therapy, different techniques are known, which often rely on experimental parameters like the diffusion interaction parameter (kD) or on computational tools harnessing information derived from the primary sequence of the protein.
Up until the work of Roberts (Woldeyes, M. A., Qi, W., Razinkov, V. I., Furst, E. M. & Roberts, C. J. How Well Do Low- and High-Concentration Protein Interactions Predict Solution Viscosities of Monoclonal Antibodies? J Pharm Sci 108, 142-154 (2019)) experimental data of colloidal interaction, mainly the diffusion interaction parameter (kD) and the second virial coefficient (A2), were found to at least qualitatively predict mAbs with potentially high viscosity (problematic mAbs). Though Roberts highlights many cases in which prediction failed. This questions the validity of the use of low concentration experimental data as a predictive tool. Most publications also refer to a rather low number of samples and describe the relationship between the experimental data and the viscosity linearly. The complexity of the origin of the solution viscosity at high mAb concentrations suggests that a non-linear modelling approach may be more appropriate. In recent years experimental approaches to forecast a concentration-dependent viscosity of antibodies have been complemented by in silico methods, which aim at identifying the decisive molecular descriptors such as solvent exposure, local charge, hydrophobic effects, and surface patches. One model to predict the viscosity is known from Agrawal et al. (Agrawal, N. J. et al. Computational tool for the early screening of monoclonal antibodies for their viscosities, MAbs 8, 43-48 (2016)), which uses a spatial charge map (SCM) for identifying mAbs with high viscosity. SCM applies molecular dynamics (MD) simulations to calculate a score for the screening of antibody viscosity at high concentrations. However, molecular dynamics simulations are computationally costly and require structural information, a significant application bottleneck. The principle of SCM was also used in an approach testing machine learning to predict mAb solution viscosity (Lai, P.-K. DeepSCM: An efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity. Comput Struct Biotechnol J, 20, 2143- 2152 (2022)). The deep learning algorithm that Lai established used the preprocessed antibody sequences as input and the SCM scores obtained from MD simulations as output for model training; and a DeepSCM surrogate model for SCM scores was developed based on the convolutional neural network (CNN) architecture. The SCM scores then are used for predicting high concentration viscosity. Lai reports an accuracy of 0.9.
Summary of the invention
The present invention refers to a method for providing a computer-implemented neural network configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated to said protein solution. Further, the present invention, in particular the method, may comprise the provision and use of a computer-implemented neural network, in particular of a computer-implemented neural network to be trained. Specifically, the provision and use may be performed to provide a trained computer-implemented neural network, in particular which is configured to predict the concentration-dependent viscosity of a protein solution based on a plurality of input parameters associated to said protein solution. Said method, specifically using the computer-implemented neural network, in particular using the computer-implemented neural network to be trained, comprises the steps of:
- providing a plurality of training data sets, each of which is associated to a specific protein solution and comprises a plurality of input parameters being indicative of the specific protein solution and at least one associated output parameter being indicative of a concentration-dependent viscosity of the specific protein solution; and
- training a neural network based on the training data sets to provide the computer-implemented neural network configured for predicting the concentrationdependent viscosity. The input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof; ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
A further embodiment of the invention comprises a computer-implemented method for predicting a concentration-dependent viscosity of a protein solution by using such a computer-implemented neural network, in particular a trained computer- implemented neural network, more particularly which is provided upon provision and use of a computer-implemented neural network to be trained as described above. The computer-implemented method for predicting the concentrationdependent viscosity of a protein solution of the invention has a high accuracy, which may be reflected by R2 values greater than 0.95, and even greater than 0.99.
A still further embodiment of the invention comprises a method for determining a concentration-dependent viscosity of a protein solution by using such a computer- implemented neural network, in particular such a trained computer-implemented neural network.
A still further embodiment of the invention comprises a method for providing a drug product comprising a protein solution, which enables effective and efficient assessment of the suitability of said protein solution for use as a drug product already at an early stage of development. Preferably, any of the methods disclosed herein which use a computer- implemented method for predicting or determining a concentration-dependent viscosity of a protein solution use the computer-implemented neural network which is initially configured (sometimes hereinafter referred to as "trained"), for predicting a concentration-dependent viscosity of a protein solution as described herein, also with all its embodiments.
In the context of the present disclosure, the task of a computer-implemented neural network of calculating a concentration-dependent viscosity may be called determining or predicting the concentration-dependent viscosity; the terms "determining" and "predicting" in connection with the calculation of the computer- implemented neural network can be used interchangeably.
In the context of the present disclosure, the suitability of the protein to be used as a drug product in form of a solution is determined by the viscosity of the protein solution at a high concentration, in particular at a concentration of more than 100 mg/mL of the protein in the solution, which should not exceed 15 to 30 mPa*s, preferably 15 to 20 mPa*s, in particular 15 mPa*s.
These objectives are solved by the subject matter of the independent claims. Preferred embodiments are set forth in the present specification, the Figures, and the dependent claims.
Accordingly, the present invention also comprises a method for providing a computer-implemented neural network (sometimes hereinafter referred to “the neural network”), preferably an artificial neural network (ANN), also referred to as trained ANN. In particular, the present invention comprises the provision and use of such a computer-implemented neural network. The neural network, in particular the trained neuronal network, is configured for predicting or determining a concentration-dependent viscosity of a protein solution in dependence on a plurality of input parameters associated to said protein solution. The method comprises a step of providing a plurality of training data sets, each of which is associated to a specific protein solution and comprises a plurality of input parameters being indicative of the specific protein solution and at least one associated output parameter being indicative of a concentration-dependent viscosity of the specific protein solution; and a step of training a neural network, i.e. a neural network to be trained or an untrained neural network, based on the training data sets to provide the computer-implemented neural network configured for predicting the concentration-dependent viscosity, i.e. to provide the trained ANN.
The term "protein solution" refers to an aqueous solution of a protein, preferably a therapeutic protein (often referred to as Active Pharmaceutical Ingredient ‘API’). As such, the protein solution refers to any solution comprising at least one protein dissolved or substantially dissolved in an aqueous medium such as water, a buffer, cell culture medium and the like.
The term "protein" means a polypeptide composed of a sequence of amino acids. A protein may be a therapeutic protein, e.g., used in diagnosis, treatment, and/or prevention of a disease or disorder. A protein can be a native protein, that is, a protein produced by a naturally-occurring and non-recombinant cell; or it can be produced by a genetically-engineered or recombinant cell and may comprise molecules having the amino acid sequence of the native protein, or molecules having deletions from, additions to, and/or substitutions of one or more amino acids of the native sequence, or molecules having the amino acid sequence of a protein with no relation to a native protein. The term also includes amino acid polymers in which one or more amino acids are chemical analogues of a corresponding naturally-occurring amino acid polymer.
Therapeutic proteins included in the protein solution may be antibodies, which may include monoclonal antibodies and polyclonal antibodies, whole antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments, e.g., Fab, Fab', F(ab')2, fragments and the like, including hybrid fragments of any immunoglobulin or any natural, synthetic or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex in one embodiment monoclonal antibodies. The term “epitope” denotes a protein determinant capable of specifically binding to an antibody. Epitopes usually consist of chemically active surface groupings of molecules such as amino acids or sugar side chains and usually epitopes have specific three-dimensional structural characteristics, as well as specific charge characteristics. Conformational and non-conformational epitopes are distinguished in that the binding to the former but not the latter is lost in the presence of denaturing solvents.
Preferred therapeutic proteins are monoclonal antibodies. The term ‘monoclonal antibody’ (mAb) as used herein refers to an antibody obtained from a population of substantially homogeneous antibodies, i.e. the individual antibodies comprising the population are identical except for possible naturally occurring mutations that may be present in minor amounts. Monoclonal antibodies are highly specific, being directed against a single antigenic site. Furthermore, in contrast to polyclonal antibody preparations, which include different antibodies directed against different antigenic sites (determinants or epitopes), each monoclonal antibody is directed against a single antigenic site on the antigen. The modifier “monoclonal” indicates the character of the antibody as being obtained from a substantially homogeneous population of antibodies and is not to be construed as requiring production of the antibody by any particular method. Preferred monoclonal antibodies are those of class Immunoglobulin G (IgG) such as lgG1 , lgG2, lgG3, and lgG4.
Suitable aqueous media of the protein solution are known to the skilled artisan. Preferably, the aqueous medium is or comprises water or an aqueous buffer such as Histidine-HCI buffer.
Other buffers can be used, buffers which are commonly used for the formulation of a drug product are known to the skilled person, such as acetate buffer, phosphate buffer, succinate buffer or citrate buffer.
Preferably, the aqueous medium is a buffer having a pH of 5 to 8, more preferably of 5 to 7.5, even more preferably of 5.5 to 6.5.
The aqueous medium may also have other pH values suitable for providing a protein solution, the common pH values which the aqueous mediums of drug products have, are known to the skilled person.
Preferably, any input parameter and also any of the at least one associated output parameter, the latter being provided for the configuring of the computer- implemented neural network, is determined in a similar aqueous medium that is also used in the intended drug product comprising said protein solution, or more preferably even in the same aqueous medium, such as with same pH, with same buffer, same viscosity, in particular same buffer viscosity etc.
At least one pharmaceutically acceptable excipient may also be comprised in the protein solution. Suitable excipients are known to the skilled artisan such as stabilizers, pH modifying agents, such as one or more buffers, etc. In some embodiments, a polysorbate such as polysorbate 20, 40, 60 or 80 may be used as stabilizing agent, which can improve stability of a protein in an aqueous solution. Other stabilizer can be for example sugars or sugar alcohols, surfactants, salts or antioxidants, such as for example L-methionine.
As set forth above, the method for providing the computer-implemented neural network, in particular the trained computer-implemented neural network, also referred to as the trained ANN, comprises the step of providing the plurality of training data sets. Further, the use of the computer-implemented neural network, in particular of the neural network to be trained, may comprise the step of providing the plurality of training data sets. Each of the training data sets is associated to a specific protein solution. Specifically, at least two of the plurality of training data sets are associated to different protein solutions. More specifically, each of the training data sets may be associated to a specific protein, i.e. the protein included in the protein solution. Preferably, the plurality of training data sets may be associated to different proteins, which in particular constitute a set of proteins. In other words, each protein included in the set of proteins is associated to at least one training data set.
Preferably, the proteins underlying each training data set are different proteins, but similar to each other in at least one aspect. The proteins in the set of proteins may be any of the above-mentioned types of proteins. More preferably, all the proteins in the set of proteins of the training data sets are antibodies, even more preferably all proteins in the set of proteins are monoclonal antibodies.
For improving training of the computer-implemented neural network, in particular of the ANN, more particular of the computer-implemented neural network to be trained, the number of training data sets, and hence the number of proteins may be more than two. The accuracy of prediction (or determination) of the viscosity can be expressed, for example, by the accuracy value which is comparable to the value of the coefficient of determination R2; these values range from 0 (reflecting no accuracy of determination, which can also be called accuracy of prediction) to 1 (reflecting 100% accurate determination).
For example, the number of training data sets (and hence proteins in the set of proteins) may be at least 15, preferably at least 20, more preferably at least 25. In general, higher numbers of training data sets result in better accuracy. With an increasing number of proteins in the set of proteins, the accuracy of determination no longer increases linearly, but approaches the upper limit of 1 . For example, where 25 training data sets are provided, the accuracy of determination is usually quite high.
In the context of the present disclosure, the term "concentration-dependent viscosity" refers to a parameter being indicative of the protein solution's viscosity, in particular the protein solution's dynamic viscosity, at one or more protein concentrations of the protein solution. As such, the concentration-dependent viscosity associates at least one protein concentration of the protein solution to a corresponding viscosity of the protein solution. For example, the concentrationdependent viscosity may be or may indicate a viscosity in form of a viscosity value cs of the protein solution at a selected protein concentration, in particular at a single protein concentration. In other words, the concentration-dependent viscosity can by expressed and/or determined as a viscosity value cs of the protein solution at a selected protein concentration.
Alternatively, the concentration-dependent viscosity may indicate a viscosity of the protein solution at more than one protein concentration. Specifically, the concentration-dependent viscosity may be represented by a function fv, in particular a mathematical function, associating a viscosity of the protein solution to a protein concentration of the protein solution. For example, the concentrationdependent viscosity may be represented by a curve diagram associating viscosity values to protein concentrations of the protein solution.
Further, the concentration-dependent viscosity may be represented by the following function represented by equation (1 ): f^fc) = A * eB*c (1 ) wherein fv refers to a function associating viscosity of the protein solution to a protein concentration; c refers to a protein concentration; A and B refer to constants.
This equation (1) can be linearized to obtain the intercept A and the slope B using the natural logarithm as in equation (2): lii7]re; = In?! + B * c (2)
Equation (2): linearized viscosity behavior; wherein ?]rei refers to a relative viscosity; A refers to an intercept; B refers to a slope; c refers to the protein concentration.
In a further development, the concentration-dependent viscosity of the protein solution may be represented by at least one, preferably both, of the constants A and B of the above equation (1 ). In other words, the computer-implemented neural network, in particular the trained computer-implemented neural network, more specifically the trained ANN, and/or the computer-implemented neural network to be trained, may be configured so as to determine at least one, preferably both, of constants A and B of the above equation (1).
The term computer-implemented neural network (also referred to as “the neural network” herein) may be an artificial neural network. Accordingly, the term neural network used in the present disclosure may be used interchangeably with the term "artificial neural network" (ANN), in particular with the term ANN to be trained or trained ANN. In general, an artificial neural network refers to a computing system employing interconnected nodes, which in particular refer to functional units, which make use of machine learning. Artificial neural networks are a subset of machine learning, also known as deep learning, and use unstructured data sets and hidden layers. Similar to normal machine learning, data sets for establishing the neural network are usually divided into the training data sets and validation data sets.
In the ANN, the interconnected nodes, which may also be referred to as artificial neurons, receive signals, process these signals and can signal nodes connected to them based on the received and processed signals. Typically, the signals transmitted to nodes are referred to as inputs and usually represent a real number. The output of each node is calculated by a non-linear function, also referred to as activation function, typically based on the sum of its inputs. The nodes and their connections typically have weights that are adjusted during a training phase based on the training data sets. Then for validating the trained ANN, the validation data sets are used allowing to determine accuracy of the trained ANN. Typically, in the ANN, nodes are aggregated into layers as will be described hereinafter. Usually, signals travel from a first layer, i.e. the input layer, via hidden layers to a last layer, i.e. the output layer, possibly after traversing the layers multiple times.
The general function and operation of such ANN are well known to a person skilled in the art and are thus not further specified. Rather, technical features of the ANN interlinked with the present invention are addressed in the following.
Typically, for providing an ANN suitable for predicting the concentration-dependent viscosity, at first, an ANN to be trained, in particular an untrained ANN, may be provided and thereafter properly trained based on the training data sets.
Accordingly, the method may further comprise a step of providing a neural network to be trained, that preferably is an untrained ANN.
Specifically, the neural network, in particular the neural network to be trained and/or the trained neural network, may comprise an input layer having a plurality of input nodes, each of which receives one input parameter. In other words, each one of the plurality of input nodes preferably receives one input parameter. The received input parameters preferably differ among the input nodes.
Alternatively or additionally, the neural network, in particular the neural network to be trained and/or the trained neural network, may comprise at least one hidden layer, in particular only one hidden layer, having a plurality of hidden nodes, in particular four or more hidden nodes. According to one configuration, the neural network, in particular the neural network to be trained and/or the trained neural network, may comprise a single hidden layer with four hidden nodes. The hidden nodes may use tan h as an activation function. Tan h transforms values to be between -1 and 1 . Alternatively, sigmoid may be used as the activation function of the hidden nodes.
Alternatively or additionally, the neural network, in particular the neural network to be trained and/or the trained neural network, may comprise an output layer having at least one output node for providing output parameters. Specifically, the output layer may have a plurality of output nodes, each of which determines and outputs a single output parameter. The output parameters provided by the output layer may be used to determine the concentration-dependent viscosity. Alternatively, the output parameters provided by the output layer may indicate or represent the concentration-dependent viscosity. For example, the output layer may be configured to provide at least one output parameter which indicates at least one of constant A and constant B of the above equation (1 ). In other words, the output layer may be configured to provide at least one value as an output parameter which is used as constant A or constant B of the above equation (1 ). For example, the output layer may be configured to provide a first output parameter by a first output node which is used as constant A and a second output parameter by a second output node which is used as constant B to determine the concentrationdependent viscosity represented by the function specified in above equation (1).
In general, the neural network, in particular the neural network to be trained and/or the trained neural network, receives a plurality of input parameters based on which the at least one output parameter is calculated. Thus, the training data sets include or consist of the input parameters and the at least one output parameter.
In the step of training the neural network, preferably the training data sets are used to implement or establish the neural network, in particular by adapting the weights of the neural network’s nodes and their connections in dependence on the training data sets. By doing so, the neural network to be trained becomes the trained neural network, i.e. the neural network provided by the above method. In the following, the training data sets are further specified which allow to train and thus provide an ANN enabling the prediction of a viscosity of a protein solution at a high accuracy.
Each set of the training data set preferably has the same number and type of parameters. Further, as set forth above, each training data set is associated to a specific protein solution, in particular to a specific protein. Specifically, the training data set may be subdivided into input parameters, i.e. parameters representing inputs to the neural network, and output parameters, i.e. parameters representing outputs of the neural network. The input parameters are indicative of the specific protein solution that the corresponding training data set is associated to. As such, the input parameters may indicate properties of the specific protein solution or of the protein included therein. The output parameters are indicative of the concentration-dependent viscosity of the specific protein solution. As such, the output parameters may represent or indicate the concentration-dependent viscosity of the specific protein solution.
Preferably, the input parameters included in the training data sets are indicative of protein-protein-interactions in the associated protein solution.
Inter-molecular distances reduce with increasing concentrations, and protein- protein-interactions then exponentially influence and determine protein's solubility, aggregation, and also protein solution's viscosity. Thus, by providing input parameters to the neural network which are indicative of the protein solution's protein-protein-interactions, the proposed method takes those properties into account which substantially influence the protein solution's viscosity at increasing protein concentrations. By doing so, an effective and efficient prediction of a protein solution’s viscosity may be achieved.
In general, protein-protein-interactions are influenced by the protein’s amino acid sequence and resulting three-dimensional structure with charged or hydrophobic patches. Further, solution conditions may influence protein-protein-interactions by modulating size of charged patches via pH, shielding of charged patches via short- ranged electrostatic interaction using salts, buffer substances, amino acids or other charged excipients.
As set forth above, the plurality of input parameters include at least one of: i) experimental data obtained by detecting parameters from a provided protein solution, preferably wherein the experimental data expresses parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; ii) computational data, preferably calculated from the primary amino acid sequence of said protein at a pH of 5.0 to 7.0, preferably wherein the computational data expresses parameters selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), Fv symmetry parameter, Fv hydrophobicity, Vh-charge, Vl-charge, hinge-charge and hydrophobic solvent accessible surface area; and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, and combinations thereof.
In other words, the plurality of input parameters may include at least one selected from the group consisting of said i) experimental data, said ii) computational data and said iii) in silico data.
Specifically, the computational data and the in silico data may be data calculated in dependence on the primary amino acid sequence of the protein included in the protein solution.
The i) experimental data are preferably parameters selected from at least one of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof. In other words, the i) experimental data may comprise at least one parameter selected from the group consisting of apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD) and second virial coefficient (A2). More preferably, the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2).
The ii) computational data are preferably parameters selected from at least one of the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof. In other words, the ii) computational data may comprise at least one parameter selected from at least one of the protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge). More preferably, the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge).
The iii) in silico data are preferably parameters selected from at least one of hydrophobic and charged patch sizes such as score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. More preferably, the iii) in silico data consists of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. In other words, the iii) in silico data comprises at least one parameter selected from the group consisting of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
Using computational data and in silico modelling has the advantage of being relatively easily accessible and does not require material and laboratory work.
Methods for determining the diffusion interaction parameter (kD) are known to the person skilled in the art. A suitable method may exemplarily be via dynamic light scattering (DLS), e.g. following the procedure further outlined below.
Virial coefficients may be determined in various ways, such as static light scattering (SLS), in this case the virial coefficients are commonly denoted with A2 for the second virial coefficient and A3 for the third virial coefficient, or by measurement of the osmolarity, then the virial coefficients are commonly denoted with B2 for the second virial coefficient and B3 for the third virial coefficient.
Methods for determining the second virial coefficient (A2) or third virial coefficient (A3) are known to the person skilled in the art. A suitable method may exemplarily be via static light scattering (SLS), e.g. following the below further outlined procedure.
Methods for determining the apparent surface hydrophobicity are known to the person skilled in the art. A suitable method may exemplarily be via hydrophobic interaction chromatography (HIC), e.g. following the below further outlined procedure.
Methods for determining the computational data are known to the person skilled in the art. Suitably, the respective parameters such as Fv-charge are calculated at a pH of 5 to 8, or of 5 to 7.5, or of 5.5 to 6.5, or of about 5.5, or of about 6.0, or of about 6.5, preferably at pH 6.0. The person skilled in the art may revert to commonly known platforms such as Prot pi | Protein Tool, https://www.protpi.ch/Calculator/ProteinToo, site operator Roland Josuran Prot pi, 8820 Wadenswil, Switzerland, following the below further outlined procedure.
The isoelectric point (pl) can be calculated "manually" from the pKa values of the amino acid residues of the primary sequence. Methods for determining the in silico data are known to the person skilled in the art. Suitably, they may be determined using software BioLuminate (version 3.80; Schrodinger, LLC, New York, NY), e.g. following the below further outlined procedure.
Specifically, the neural network, in particular the neural network to be trained and/or the trained neural network, may receive input parameters selected from all three data groups i) to iii) above, but a reliable prediction of the concentration dependent viscosity is also possible if the neural network receives input parameters from only two of the three data groups, such as the neural network receives input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii).
As set forth above, the neural network, in particular the neural network to be trained and/or the trained neural network, receives the plurality of input parameters based on which the at least one output parameter is calculated. The output parameters may represent or indicate the concentration-dependent viscosity of the specific protein solution. As such, the output parameter may indicate a viscosity value cs associating a viscosity of the protein solution to a specific protein concentration of the protein solution. Alternatively or additionally, the at least one output parameter may indicate a viscosity of the protein solution at more than one protein concentration. Specifically, the at least one output parameter may indicate the function fv represented by above equation (1 ). More specifically, the at least one output parameter may indicate at least one, preferably both, of constants A and B of the above equation (1 ).
In the following, an example is provided how to determine the at least one output parameter included in the training data sets. At first, different protein solutions, referred to as protein solution samples, may be provided which comprise the same type of protein, but at different concentrations. Based on these samples, concentration-dependent viscosity data may be generated, e.g., by using a VROC viscometer (Rheosense) for each concentration. In this way, a plurality of value pairs may be provided, each of which associates a viscosity of the protein solution to a protein concentration in the protein solution. These measured viscosity data, in particular experimentally measured viscosity data, in a next step, may be processed based on the following equations (3) to (5). As to substance, equation (3) may be used to calculate the relative viscosity of the different protein solution samples. For doing so, at first, the viscosity of the buffer comprised in the protein solution may be determined, preferably at a defined temperature. Knowing the viscosity of the buffer, the relative viscosity of the protein solution samples may then be calculated based on the measured viscosity data according to equation (3). By doing so, value pairs for concentration and relative viscosity are determined for each protein solution sample.
O)
Equation (3): Relative viscosity; qrei = relative viscosity ; qo = buffer viscosity; q = sample viscosity
Equation (4) may then be used to describe the exponential concentrationdependent viscosity of the protein solutions. This equation can be linearized to obtain the intercept A and the slope B using the natural logarithm as in equation (5). r]rei = A * eB*cs (4)
Equation (4): Exponential viscosity behaviour; qrei = relative viscosity; A = intercept; B = slope; cs = selected mAb concentration
In ?]re; = In + B * cs (5)
Equation (5): linearized viscosity behavior; qrei = relative viscosity; A = intercept; B = slope; cs = selected mAb concentration
Based on the determined value pairs for concentration and relative viscosity of the protein solution samples, the viscosity descriptors A and B for each protein may then be determined based on equation (4) and/or (5) by applying, e.g., curve fitting techniques, such as least square fitting.
With each protein in the set of proteins of the training and of the validation data set a set of a number (num-conc) of protein solutions having different protein concentrations are provided and each of their sample viscosity q is determined; with num-conc being at least two, preferably 2 to 12, such as 4, 5, 6, 7, 8, 9, 10 11 or 12, a typical value is 6; these values pairs of protein concentration and sample viscosity serve for determining the at least one associated output parameter of the training and validation data set as described above.
The different concentrations in said set of protein solutions can for example range from a minimum concentration (mini-conc) of 10 mg/mL to a maximum concentration (max-conc) of 320 mg/mL, or from 10 to 280 mg/mL, or from 10 to 240 mg/mL, or from 10 to 210 mg/mL; a typical value for min-conc is 30 mg/mL and for max-conc is 180 mg/mL. Preferably, in case that the set of protein solutions contains 3 or more protein solutions, each with a different protein concentration, the difference of every two concentrations adjacent to each other, that is the difference of two subsequent concentrations, is the same, so
(max-conc - min-conc) I (num-conc - 1 ) = diff-conc; and concn+i - concn = diff-conc. max-conc maximum concentration in said set of protein solutions min-conc minimum concentration in said set of protein solutions num-conc number of protein solutions with different concentration in the set of protein solutions diff-conc difference between two subsequent concentrations n integer from 0 to (num-conc - 1 ) concn concentration of the (n+1 ) protein solution in the set of protein solutions
In one embodiment, the set of protein solutions for each protein in the set of proteins contains 6 protein solutions with the 6 different concentrations of 30 mg/mL, 60 mg/mL, 90 mg/mL, 120 mg/mL, 150 mg/mL and 180 mg/mL, so max cone is 180 mg/mL, min cone is 30 mg/mL, num cone is 6, diff cone is 30 mg/mL, n integer from 0 to 5.
A method according to the present invention may further comprise a step of validating the trained ANN, i.e. after the step of configuring, also called training herein, the neural network has been performed. In this "validation" step, a plurality of validation data sets may be provided. In other words, the method may comprise a further of validating the trained neural network based on a plurality of validation data sets, each of which comprises a plurality of input parameters and at least one associated output parameter. The validation data sets may have the same or different number and type of parameters as the training data sets, but it is preferred that the validation data sets use the same type of parameters as the training data sets. As such, each validation data set may be divided into input parameters and associated output parameters, likewise to the training data sets. Then, the validation data sets may be used to validate proper implementation or training of the neural network. For doing so, the trained neural network may receive the input parameters of a validation data set to compute the corresponding at least one associated output parameter. The computed output parameter may then be compared to the at least one associated output parameter included in the validation data set. This may be performed for each set of the validation data set to determine whether proper configuring of the neural network is established.
In a further development, the trained neural network may be stored on a data medium, e.g., a non-transitory storage medium configured to store digital data of the sort that will be apparent to one of skill in the art in view of the present disclosure.
A further aspect of the invention comprises a computer-implemented method for determining or predicting a concentration-dependent viscosity of a protein solution by using a computer-implemented neural network, in particular a trained neural network, more particularly a trained neural network as described above. Specifically, the further aspect of the invention may comprise the provision and use of a computer-implemented method for determining or predicting a concentrationdependent viscosity of a protein solution by using a computer-implemented neural network. The neural network, in particular the trained neural network, is configured for predicting a concentration-dependent viscosity of the protein solution in dependence on a plurality of input parameters associated to said protein solution, preferably wherein the input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof; ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
In other words, the input parameters may include at least one selected from the group consisting of said i) experimental data, said ii) computational data and said iii) in silico data.
Preferably, the computer-implemented method of the present invention makes use of the neural network, in particular of the trained neural network described above, i.e. which is provided by the above described method. Accordingly, technical features described herein in connection with the method for providing a computer- implemented neural network, in particular in connection with the provision and use of a computer-implemented neural network, may thus also apply and refer to the computer-implemented method for predicting a concentration-dependent viscosity of a protein solution, and vice versa. This particularly applies to the abovedescribed features of the neural network, in particular of the neural network to be trained or the trained neural network, of the protein solution, of any input parameter and of any output parameter.
The methods of the present invention may be used to predict a viscosity of any suitable protein solution. As described above, the protein solution may comprise a therapeutic protein, preferably an antibody, more preferably a monoclonal antibody (mAb), such as a monoclonal antibody of lgG1 or lgG2 subtype. Accordingly, the protein solution may be an aqueous medium including water, preferably a buffer, more preferably a Histidine-HCI buffer.
The neural network, in particular the trained neural network, used in the methods of the present invention is configured for determining or predicting the concentrationdependent viscosity in dependence on a plurality of input parameters.
As described above, the concentration-dependent viscosity can by expressed and/or determined as a viscosity value cs of the protein solution at a selected protein concentration. Alternatively or additionally, the concentration-dependent viscosity can by expressed and/or determined as a function fv associating viscosity of the protein solution to protein concentration. Specifically, the concentrationdependent viscosity may be indicative of or may be represented by the function fv depicted by above equation (1). More specifically, the neural network may be configured to calculate at least one of constants A and B of the function fv.
In a method according to the present invention, the neural network, in particular the trained neural network, receives the input parameters and, based thereupon, determines the concentration-dependent viscosity of the protein solution. The input parameters are preferably associated to the protein solution, the viscosity of which is to be predicted. Specifically, the input parameters preferably refer to the protein included in the protein solution. More preferably, the input parameters are indicative of protein-protein-interactions in the protein solution. Further, the input parameters may correspond to those input parameters included in the training data sets as described above. Thus, the technical features described above in connection with parameters contained in the training data sets may apply and refer likewise to the input parameters to be received by the neural network for calculating the concentration-dependent viscosity, and vice versa.
More specifically, in the method, the neural network, in particular the trained neural network, receives the input parameters as described above. Accordingly, the input parameters include at least one of: i) experimental data preferably selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof. More preferably, the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2); ii) computational data preferably selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof. More preferably, the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv-charge); and iii) in silico data preferably selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
In other words, the input parameters may include at least one selected from the group consisting of said i) experimental data, ii) computational data, and iii) in silico data.
Specifically, as set forth above, the input parameters may be selected from all three data groups i) to iii) above, but a reliable prediction of the concentration dependent viscosity is also possible if the neural network receives input parameters from only two of the three data groups, such as the neural network receives input parameters from data groups i) and ii), or from data groups ii) and iii), or from data groups i) and iii) above, preferably from data groups ii) and iii).
In a further aspect of the invention, a method is provided for determining a concentration-dependent viscosity of a protein solution. The method comprises a step of predicting a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the concentration-dependent viscosity with a target viscosity, in particular to determine a viscosity characteristic of the protein solution.
The proposed method may make use of the above described neural network and the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution. Thus, technical features described above, in particular in connection with the neural network and its use for predicting a concentration-dependent viscosity of a protein solution, may thus also apply and refer to the method for determining a concentration-dependent viscosity of a protein solution, and vice versa. The proposed method may be used for assessing whether the protein solution is subjected to a high or too high viscosity, in particular at a predefined protein concentration. In other words, the step of comparing the concentration-dependent viscosity with the target viscosity may be performed to determine whether the protein solution is subjected to a high or too high viscosity, in particular whether it exceeds the target viscosity, in particular at a predefined protein concentration. A viscosity of a protein solution may be considered too high if it renders the protein solution unsuited as a drug product for the reasons given herein above, i.e. , if it’s production produces high costs, due to high loss and low recovery in purification, difficult manufacturing or filling, and ultimately low administrability due to need for high injection forces and slow administration with potential sensation of pain. In general, solutions with a dynamic viscosity above 15 to 30 mPa*s or even above 15 to 20 mPa*s, are considered problematic, and hence “too high”.
In a further development, the method may be used for determining suitability as a drug product of a protein solution, in particular of a protein solution to be used or prepared as a drug product. For doing so, the step of comparing the concentrationdependent viscosity with the target viscosity is performed to determine suitability as a drug product of the protein solution.
By providing a method in which suitability of a protein solution as a drug product is assessed based on a concentration-dependent viscosity determined by a computer implemented neural network, in particular by a trained neural network, the suggested method makes it possible to easily and reliably take into account the concentration-dependent viscosity of a protein solution dependent on the concentration of the protein in the protein solution, in particular at early stages of a development process. In this way, an improved validation of a protein solution to be used as a drug product is enabled at early stages, which has an impact on the drug product to be produced.
The proposed method may be used for assessing suitability of a protein solution to be used in any drug products which comprise or are constituted by a protein solution.
As set forth above, the method comprises the step of comparing the determined concentration-dependent viscosity with a target viscosity. In the context of the present disclosure, the term "target viscosity" refers to a desired or predefined viscosity of the protein solution; or the target viscosity may represent an upper threshold for the viscosity of the protein solution. For example, the target viscosity may indicate a maximum value for the viscosity of the protein solution at a given concentration. Specifically, the target viscosity may indicate a maximum value for the viscosity of the protein solution between 15 mPa*s to 30 mPa*s, or between 15 mPa*s to 25 mPa*s, or between 15 mPa*s to 20 mPa*s, for example 15 mPa*s or 18 mPa*s or 20 mPa*s or 25 mPa*s. A viscosity above the maximum value may be regarded as problematic when using the protein solution as a drug product.
Therefore, this step may be performed to determine suitability of the protein solution as a drug product. In other words, by comparing the concentrationdependent viscosity with the target viscosity, it may be assessed whether the protein solution has a concentration-dependent viscosity which is favorable when being used as a drug product. Suitability of the protein solution may be determined when the concentration-dependent viscosity complies with or is below of the target viscosity, i.e. lies within a desired viscosity range even if the protein is used at a high concentration, e.g., of > 50 mg/mL, or > 60 mg/mL, > 70 mg/mL, or > 80 mg/mL, or > 90 mg/mL, or > 100 mg/mL. By doing so, application-specific viscosity of the protein solution may be forecasted to assess whether the protein solution is suitable for being used as a drug product in the clinical therapy.
Preferably, in the step of comparing the concentration-dependent viscosity with the target viscosity, a protein concentration of the protein solution is taken into account. In other words, the concentration-dependent viscosity and the target viscosity may be compared at a specific protein concentration or within a protein concentration range.
When comparing the concentration-dependent viscosity and the target viscosity at a specific protein concentration, at first, a protein concentration of the protein solution may be determined. For example, a maximum protein concentration may be determined which may refer to a protein concentration the protein solution is expected to not exceed when being used as a drug product. The maximum protein concentration may be in the range from 100 mg/mL to 220 mg/mL or from 120 mg/mL to 180 mg/mL, for example 120 mg/mL or 150 mg/mL or 180 mg/mL. Then, based on the determined concentration-dependent viscosity, the viscosity of the protein solution at the maximum protein concentration may be determined and compared to the target viscosity.
When comparing the concentration-dependent viscosity and the target viscosity within a protein concentration range, at first, a realistic protein concentration range of the protein in the protein solution for clinical therapy may be determined. For example, the protein concentration range may span from 20 mg/mL or 50 mg/mL or 100 mg/mL to the maximum protein concentration. Then, it may be determined whether the concentration-dependent viscosity of protein solution exceeds the target viscosity within said protein concentration range.
In a further aspect of the invention, a method is given for providing (or preparing) a drug product comprising a protein solution. The method comprises a step of predicting (or determining) a concentration-dependent viscosity of a protein solution by using a computer-implemented neural network; a step of comparing the concentration-dependent viscosity with a target viscosity to determine suitability of the protein solution as a drug product; and an optional step of preparing the drug product if suitability of the protein solution is determined.
The proposed method may make use of the above described neural network, the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution, and the above described method for determining a concentration-dependent viscosity. Thus, technical features described above may thus also apply and refer to the method for providing, and optionally preparing, a drug product, and vice versa.
If in the method, as a result of the comparison of the concentration-dependent viscosity with a target viscosity, suitability of the protein solution as a drug product is determined, then the step of preparing the drug product may be performed. If suitability of the protein solution as a drug product is not determined, then, the protein solution may be adapted or changed and the step of determining a concentration-dependent viscosity of a protein solution and the step of comparing the determined concentration-dependent viscosity with a target viscosity may be performed again based on the adapted or changed protein solution.
Still further, a method is provided for determining a suitable concentration of a protein in a protein solution of a drug product; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine the upper limit of the concentration of the protein in the protein solution which is still acceptable for a drug product.
Still further, a method is provided for identifying an upper limit of the concentration of the protein in a protein solution which should not be exceeded in order to avoid unacceptable high viscosity of the protein solution; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine if the protein in the protein solution induces above a certain concentration a viscosity of the protein solution which is not acceptable for a drug product.
Still further, a method is provided for identifying a protein having an unacceptable concentration-dependent viscosity in solution; the method comprises a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine if the protein in the protein solution induces above a certain concentration a viscosity of the protein solution which is not acceptable for a drug product.
Still further, the present invention provides a use of a computer-implemented neural network for facilitating the preparation of a drug product, wherein the neural network is configured to predict a concentration-dependent viscosity of a protein solution which is used to determine suitability of the protein solution to be prepared as the drug product.
Still further, the invention provides a method for determining the suitability of a protein solution as a drug product, comprising: a step of experimentally detecting parameters from a provided protein solution, preferably wherein the obtained experimental data expresses parameters selected from protein hydrophobicity, diffusion interaction parameter (kD), net protein charge, Zeta potential, second virial coefficient (A2), third virial coefficient (A3), apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), and combinations thereof; a step of determining a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network; and a step of comparing the determined concentration-dependent viscosity with a target viscosity to determine suitability of the protein solution as a drug product.
Still further, the invention provides a method for providing a trained artificial neural network (ANN) for determining a concentration-dependent viscosity of a protein solution by training the ANN with input parameters describing proteins contained in a set of proteins.
Still further, the invention provides a method for determining (or predicting) a concentration-dependent viscosity of a solution of a protein by providing the trained ANN with input parameters being indicative of or of the protein and letting the ANN calculate an output parameter indicating or being the concentration-dependent viscosity.
Still further, the present invention provides an apparatus for predicting a concentration-dependent viscosity of an experimental protein solution, the apparatus comprising: an experimental protein solution; a computer component comprising a neural network configured to predict the viscosity of the experimental protein solution: wherein the neural network is trained using a plurality of predetermined data sets; wherein each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical characteristic of a predefined protein solution, and (ii) at least one output parameter indicative of a concentration-dependent viscosity of the predefined protein solution; and wherein the plurality of input parameters comprise at least one of experimentally-derived data, computationally-derived data, and in silico-derived data.
The proposed apparatus may make use of the above described neural network and the above described computer-implemented method for predicting a concentration-dependent viscosity of a protein solution. Thus, technical features described above, in particular in connection with the neural network, more particularly in connection with the above-described trained neural network, and its use for predicting a concentration-dependent viscosity of a protein solution, may thus also apply and refer to the proposed apparatus, and vice versa.
Still further, the present invention provides a system for determining a concentration-dependent viscosity of a protein solution, the system comprising: a protein solution; and a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution depends on a plurality of input parameters (IP) associated with the protein solution, and further wherein the input parameters (IP) include at least one from the group consisting of (i) experimental data; (ii) computational data; and (iii) in silico data.
The proposed system may make use of the above described neural network and the above described computer-implemented method for predicting a concentrationdependent viscosity of a protein solution. Thus, technical features described above, in particular in connection with the neural network, more particularly in connection with the above-described trained neural network, and its use for predicting a concentration-dependent viscosity of a protein solution, may thus also apply and refer to the proposed system, and vice versa.
Brief description of the drawings
The present disclosure will be more readily appreciated by reference to the following detailed description when being considered in connection with the accompanying drawings in which:
Figure 1 a shows an exemplary computing component that may be used to implement various features of the embodiments of the present invention;
Figure 1 b shows a flow diagram illustrating a method for providing a drug product according to an embodiment of the present invention;
Figure 2 schematically shows a computer-implemented neural network used in the method depicted in Figure 1 ;
Figure 3 schematically shows a further computer-implemented neural network used in the method depicted in Figure 1 ; Figure 4 illustrates training data sets and validation data sets used for training and implementing a neural network used in the method depicted in Figure 1 ;
Figure 5 depicts a diagram illustrating a comparison between viscosity values calculated by a suggested neural network and measured viscosity values;
Figure 6 depicts examples for predicted viscosity curves (the predicted versus the measured viscosity of selected mAbs; crosses indicate the measured values, the black line was drawn using the predicted vales for the slope (B) and intercept (A) inserted into Formula X ( y = A * e ( B * x), the dashed line indicates the viscosity threshold for problematic mAbs at 15 mPa*s A) is mAb 13 B) is mAb 2 C) is mAb 17 D) is mAb 24 E) is mAb 26 F) is mAb 27);
Figure 7 shows the difference of predicted viscosity values, calculated from the predicted viscosity curves using the predicted values for the intercept and slope, to measured viscosity values at the same concentration (the percentage of difference is presented with bars and the absolute difference is presented with points);
Figure 8 illustrates the input variables and setup of the artificial neural network as exemplified (inputs are combined in one layer with 4 hidden nodes with tan h activation functions and target determining viscosity descriptors, the intercept A or slope B);
Figures 9a and 9b depict examples of the training and validation data (9a: model xx-A training and validation data, and 9b: model xx-B training and validation data).
Detailed description of preferred embodiments
In the following, the invention will be explained in more detail with reference to the accompanying Figures. In the Figures, like elements are denoted by identical reference numerals and repeated description thereof may be omitted in order to avoid redundancies. Figure 1a is a schematic view showing an exemplary computing component 2 that may be used to implement various features of the embodiments of the present invention. More particularly, computing component 2 generally comprises a bus 3 connected to (i) a processor 4, (ii) memory 5 for storing information and instructions to be executed by processor 4 (e.g., random access memory (RAM) or other dynamic memory, read-only memory (ROM) or any other static storage device for storing static information and instructions for processor 4, etc.), (iii) a storage device 6 comprising a non-transitory medium (e.g., a hard disk drive, solid state disk drive, optical storage, etc.) for the long-term storage of digital information (e.g., computer software, digital data, etc.), and (iv) a communications interface 7 for allowing software and data to be transferred between computing component 2 and external devices via a communications channel 8, as will be apparent to one of skill in the art in view of the present disclosure.
The computing component 2 may be part of or may constitute an apparatus for predicting a concentration-dependent viscosity of an experimental protein solution.
Figure 1 b depicts a method for providing and optionally preparing a drug product comprising a protein solution.
In step SO of the method, a computer-implemented neural network 10 is provided which is configured for predicting a concentration-dependent viscosity of a protein solution in dependence on a plurality of input parameters associated to said protein solution. Step SO represents a sub-method, i.e. a method as such, included in the method for preparing a drug product. It will be appreciated that computer- implemented neural network 10, which preferably is a trained neural network, may be implemented using the aforementioned computing component 2, or using any appropriate computing system that will be apparent to one of skill in the art in view of the present disclosure.
The protein solution comprises a protein, preferably a therapeutic protein such as an antibody, which may include monoclonal antibodies, polyclonal antibodies, whole antibodies, antibody-drug conjugates, chimeric antibodies, humanized antibodies, human antibodies or hybrid antibodies with dual or multiple antigen or epitope specificities, antibody fragments and antibody sub-fragments such as Fab, Fab', F(ab')2, fragments including hybrid fragments of any immunoglobulin or any natural, synthetic or genetically engineered protein that acts like an antibody by binding to a specific antigen to form a complex. In the shown embodiment, the therapeutic protein is a monoclonal antibody (mAb), preferably monoclonal antibodies of lgG1 or lgG2 subtype.
Further, the protein solution comprises an aqueous medium including water, preferably a buffer, more preferably a Histidine-HCI buffer.
In a first sub-step SO.1 , a neural network 10 to be trained, in particular an untrained neural network, is provided. The neural network 10 to be trained may be configured to receive the plurality of input parameters and based thereupon to compute at least one output parameter. In the shown configuration, the neural network 10 is configured to calculate the concentration-dependent viscosity of the protein solution in dependence on the input parameters. In general, the concentrationdependent viscosity associates at least one protein concentration of the protein solution to a corresponding viscosity, in particular dynamic viscosity, of the protein solution.
According to one embodiment of the present invention depicted in Fig. 2, the neural network 10, in particular the neural network to be trained and the trained neural network, is configured to compute a viscosity value of the protein solution at a selected protein concentration cs. The selected protein concentration cs may refer to a maximum protein concentration which is expected to be relevant for the drug product. The maximum protein concentration may be in the range from 100 mg/mL to 220 mg/mL ,or from 120 mg/mL to 180 mg/mL, for example 120 mg/mL or 150 mg/mL or 180 mg/mL. Accordingly, in this embodiment, the concentrationdependent viscosity is expressed as cs, meaning the viscosity value of the protein solution at the selected protein concentration cs.
According to another embodiment of the present invention depicted in Fig. 3, the neural network 10, in particular the neural network to be trained and the trained neural network, is configured to compute a mathematical function associating viscosity values of the protein solution to protein concentrations of the protein solution. Specifically, in this embodiment, the concentration-dependent viscosity is represented by the function ^ as specified in the above equation (1 ). Accordingly, in this embodiment, the concentration-dependent viscosity is provided in the form of the function fv. For doing so, constants A and B of the above equation (1 ) are calculated by the neural network 10. In the following, the neural network 10 to be trained will be described in more detail with reference to the embodiments depicted in Fig. 2 and Fig. 3.
Fig. 2 depicts an embodiment of the neural network 10, in particular of the neural network to be trained and of the trained neural network, used in the method depicted in Fig. 1. The neural network 10 comprises an input layer 12 having a plurality of input nodes I Ni -I Ni, wherein the index "/" refers to a positive integer number. In other words, the input layer 10 comprises / different input nodes IN. As can be depicted in Fig. 2, each input node INi-INn receives a corresponding input parameter IPi-IPi.
Further, the neural network 10 comprises a hidden layer 14 having a plurality of hidden nodes HNi-HNj, wherein the index ' ' refers to a positive integer number, preferably j is 4 and the neural network has four hidden nodes HN1-HN4, each of which is connected to each input node INi-INi. The hidden nodes HNi-HNj Use tan h as an activation function. Alternatively, sigmoid may be used as the activation function.
Still further, the neural network 10 comprises an output layer 16 having one single output node ON which is connected to all hidden nodes HNi-HNj. The output node ON provides a single output parameter in the form of the parameter cs which is the viscosity of the protein solution at the selected protein concentration cs. In this embodiment, the parameter cs constitutes the concentration-dependent viscosity.
Fig. 3 depicts the further embodiment of the neural network 10 according to the present invention, in particular of the neural network to be trained and of the trained neural network. Compared to the configuration depicted in Fig. 2, the output layer 16 of the neural network 10 is provided with two output nodes ON1 and ON2 providing different output parameters. Specifically, the first output node ON1 provides the output parameter OP1, which is A, and a second output node ON2 provides the output parameter OP2, which is B. These two parameters, as set forth above, represent the constants of function fv. In this embodiment, the function fv represented by the constants A and B constitutes the concentration-dependent viscosity.
In a next sub-step SO.2, a plurality of training data sets ST is provided, each of which is associated to a specific protein solution and comprises a plurality of input parameters IP being indicative of the specific protein solution and at least one associated output parameter OP being indicative of the concentration-dependent viscosity of the specific protein solution. Further, validation data sets Sv are provided which have the same number and types of input parameters and output parameters. Fig. 4 depicts a table generically illustrating the training data sets ST and validation data sets Sv used for properly implementing the neural network 10. Each training data set ST and each validation data set Sv comprises input parameters I Pi -IPi representing values to be received by the input nodes I Ni -INi and the two output parameters OPi and OP2 representing values to be provided by the two output nodes ON1 and ON2.
Specifically, a number of m different training data sets ST are provided based on which the computational model underlying the neural network 10 is adapted during a training phase in the next sub-step S0.3. In this context, the parameters "m" indicates a positive integer number.
Specifically, in a sub-step SO.3, based the training data sets ST, the neural network 10, in particular the neural network to be trained, is trained to provide the computer-implemented neural network 10, in particular the trained neural network, which is suitable and configured for predicting the concentration-dependent viscosity. For doing so, all input parameters IP and output parameters OP of the training data sets ST are fed to the neural network 10 to adapt the computational model underlying the neural network 10, in particular to adapt weights of the nodes and their connections.
In a further optional sub-step, a number of n different validation data sets Sv are used to validate the neural network 10, in particular the trained neural network. In this context, the parameters "n" indicates a positive integer number. Preferably, each one of the validation data sets Sv differs from each one of the training data sets ST. For validating the provided neural network 10, in particular the trained neural network, the input parameters associated to the validation data sets Sv are fed to the neural network 10 which, based thereupon, calculates the respective output parameters for each validation data set Sv. The calculated output parameters are then compared to the output parameters included in the validation data sets Sv. For example, for training such a neural network, 20 or more, e.g. 24, different training data sets ST may be used. Further, two or more, e.g. three, different validation data sets Sv may be used.
As set forth above, the neural model may comprise a number of / different input nodes and accordingly may receive / different input parameters IP per data set. The input parameters IP preferably are indicative of protein-protein-interactions in the protein solution associated thereto. Further, the input parameters IP may include different types of input parameters which can be categorized into i) experimental data, ii) computational data and iii) in silico data.
Experimental data may be obtained by detecting parameters from a provided protein solution. Preferably the i) experimental data expresses parameters selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof. More preferably, the i) experimental data consists of the parameters apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), and second virial coefficient (A2). Computational data may be calculated from the primary sequence of said protein at a pH of 5.0 to 7.0.
Preferably the computational data expresses parameters selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof. More preferably, the ii) computational data consists of the parameters protein’s isoelectric point (pl) and variable fragment (Fv)-charge (Fv- charge).
In silico data are preferably selected from hydrophobic and charged patch sizes, and combinations thereof; more preferably from size (in A2), score (sum of all contributing patch scores associated with this patch), and type (pos = positively charged; neg = negatively charged; hyd = hydrophobic) of the patches. The descriptors derived from modelling are thus score pos/neg/hyd Fv total, size pos/neg/hyd Fv total, count pos/neg/hyd Fv total; and any combination thereof. More preferably, the iii) in silico data consists of score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total. In the following, one embodiment of the neural network 10, in particular of the trained neural network, is described which is equipped with 14 different input nodes IN1-IN14 and 14 different input parameters IP1-IP14. It will be obvious for a person skilled in the art that these input parameters only depict examples of a plurality of possibilities. Hence, the input parameters described hereinafter should not be understood to form a limitation. Thus, more or less input parameters may be used. In a further embodiment, a neural network may make use of input parameters including at least one of input parameters IP1 to IP14.
Specifically, input parameters IP1-IP3 are i) experimental data, input parameters IP4 and IP5 are ii) computational data, and input parameters IPe to IP14 are iii) in silico data. As to substance:
IP1 refers to HIC RT [min];
I P2 refers
IP3 refers
IP4 refers to pl;
IPs refers to Fv charge at pH 6;
IPe refers to Score pos Fv total;
IP7 refers to Size pos Fv total;
IPs refers to Count pos Fv total;
IP9 refers to Score neg Fv total;
IP10 refers to Size neg Fv total;
IP11 refers to Count neg Fv total;
IP12 refers to Score hyd Fv total;
I Pi 3 refers to Size hyd Fv total; and IP14 refers to Count hyd Fv total.
In one embodiment, the combination of i) experimental data consisting of input parameters IP1-IP3, ii) computational data consisting of input parameters IP4 and IPs and iii) in silico data consisting of input parameters IPs to IP14 is used.
In another embodiment, only the combination of ii) computational data consisting of input parameters IP4 and IPs and iii) in silico data consisting of input parameters IPe to IP14; in this embodiment the neural network 10 is equipped with 11 different input nodes IN4-IN14 and 11 different input parameters I P4-I Pi 4. According to the former embodiment, the neural network 10, in particular the trained neural network, was used to model the viscosity of solutions of monoclonal antibodies. Fig. 5 depicts a diagram in which output values provided by the neural network 10 to model the viscosity of solutions of monoclonal antibodies are compared to measured viscosity values. As can be seen, the neural network 10, in particular the trained neural network, validly, and reliably, predicts the concentration-dependent viscosity of the protein solution, in particular by calculating a viscosity curve.
In a next step S1 of the method, a protein solution is identified based on which the drug product may be produced.
In step S2, a concentration-dependent viscosity, in particular a protein- concentration-dependent viscosity, is predicted by using a computer-implemented neural network 10, in particular by using a trained neural network. As such, step S2 represents a sub-method, i.e. a computer-implemented method as such for predicting a concentration-dependent viscosity of a protein solution.
In this step, the neural network 10, in particular the trained neural network, provided in step SO is used to compute the concentration-dependent viscosity of a protein solution. For doing so, the input parameters IP, in particular at least one of, preferably all of IPi to IP14, associated to the protein solution identified in step S1 are determined and provided to the neural network 10 to, based thereupon, compute the at least one output parameter OP indicating the concentrationdependent viscosity.
In a next step S3, the concentration-dependent viscosity determined in step S2 is compared with a target viscosity to determine suitability of the protein solution as a drug product. Steps S2 and S3 together represent a sub-method, i.e. a method as such for determining a concentration-dependent viscosity of a protein solution.
In this context, the target viscosity indicates a maximum value for the viscosity of the protein solution. Specifically, the target viscosity indicates a maximum value of between 15 mPa*s to 30 mPa*s or of between 15 mPa*s to 25 mPa*s or between 15 mPa*s to 20 mPa*s, for example 15 mPa*s or 18 mPa*s or 20 mPa*s or 25 mPa*s. In this step, suitability of the protein solution is determined if the concentration-dependent viscosity, in particular at the selected protein concentration cs or for a selected protein concentration range, for example between 20 mg/mL to the selected protein concentration cs, does not exceed the target viscosity, i.e. the threshold or maximum value of 15 mPa*s or 20 mPa*s. However, if the viscosity of the protein solution exceeds the target viscosity, suitability of the protein solution is not determined (i.e., the protein solution is determined to not be suitable as a drug product, or the protein solution is determined to potentially not be suitable as a drug product).
According to the embodiment depicted in Fig. 2, in step S3, the parameter cs is compared to the target viscosity. If the parameter cs does not exceed the target viscosity, suitability of the protein solution as a drug product is determined (i.e., the protein solution is determined to be suitable as a drug product, or the protein solution is determined to be potentially suitable as a drug product).
According to the embodiment depicted in Fig. 3, in step S3, the derived function fv is used to determine whether the viscosity of the protein solution exceeds the target viscosity in a selected protein concentration range, for example spanning between 20 mg/mL to the selected protein concentration cs. Alternatively, in this step at first, a viscosity cs of the protein solution at the selected protein concentration cs may be determined based on function fv , before determining whether the viscosity cs, i.e. the viscosity at the selected protein concentration cs, exceeds the target viscosity. Accordingly, if the target viscosity is not exceeded, suitability of the protein solution as a drug product is determined (i.e., the protein solution is determined to be suitable as a drug product, or the protein solution is determined to be potentially suitable as a drug product).
As indicated by step S4 in Fig. 1 , if suitability of the protein solution is determined in step S3, then the method may proceed to optional step S5 in which the drug product is prepared (or produced) based on the protein solution identified in step S1 , in particular at a desired protein concentration. If suitability of the protein solution is not determined in step S3, then the method may proceed to step S1 in which a new protein solution is identified, in particular by adapting or changing the initial protein solution, before performing steps S2 to S4 again.
Thus it will be seen that by screening potential protein solutions using the method described above (i.e., by screening potential protein solutions using the neural network 10 provided in step SO) it is possible to identify candidate protein solutions for use as a drug product without having to actually prepare and empirically test the potential protein solutions. The foregoing in silico method for screening potential protein solutions permits a large number of candidate protein solutions to be screened without the necessity of empirical measurement of the qualities of every candidate protein solution. As a result, the present invention solves the long- unmet need for a quick, high-throughput method for screening candidate protein solutions for use as a drug product that avoids the time and labor inherent in empirical (i.e., laboratory-based) screening.
It will be apparent for a person skilled in the art that these embodiments and items only depict examples of a plurality of possibilities. Hence, the embodiments shown herein below should not be understood to form a limitation of these features and configurations. Any possible combination and configuration of the described features can be chosen according to the scope of the invention.
Abbreviations Materials and Methods
To evaluate the impact of various input parameters on the predictive power of artificial neural networks, three different models were created containing the data of mAbs 1-25 (table 4 and 5 herein below), in which either only experimental inputs (retention time in HIC; kD, A2), or only in silico-derived inputs (patch sizes), or both inputs were fed into the modelling; in all three models Fv-charge and pl was included regardless of the selected inputs. Each model was created by splitting the input data into a training and a validation set, which contained the information from randomized mAbs for each new artificial neural network model. The training sets contained data, input variables and viscosity descriptors, of 18-20 mAbs. The validation sets contained as data only the input variables of the remaining 7-5 mAbs, with the goal to predict their viscosity descriptors. These viscosity descriptors are derived from the linearization of the viscosity-concentration-curves of each mAb, the intercept A and the slope B.
The R2 values of the created models show the interdependency of validation and training set (table 1 ; exemplary graphs are depicted in Figures 9a and 9b), as in several cases the quality of one set is excellent (R2 > 0.99) while the second set is less good (R2 < 0.95). The highest quality model was achieved for the artificial neural network in which both variables (experimental and in silico) were used, with the lowest being R2 > 0.92 (intercept A of training set). The artificial neural network using only in silico-derived inputs achieves a minimally lower "worst" R2 of > 0.90 (slope B of training set). Lastly, the artificial neural network containing only experimental data achieves the lowest R2 of > 0.75, which is comparable to results published using linear correlations of kD to viscosity, and better than a linear correlation of A2 and kD to the intercept and slope parameter done for this data set.
The slope parameter B describes the steepness of the viscosity curve so the exponential increase of viscosity with the protein concentration, which is important to describe potentially ..problematic" mAbs. Such ..problematic" mAbs may show a moderate viscosity at low protein concentration but experience substantial increase of viscosity above values usually regarded as acceptable (i.e. , values above about 15-25 mPa*s) for drug products. Consequently, the artificial neural network was used with all input variables in the following, as it provided the best prediction for the slope B.
Of all available data, mAbs 26 and 27 were not used in artificial neural network model creation but were kept separately for additional verification tests.
Table 1 :
Table 1 shows a comparison of models created using different inputs predicting either the intercept (A) or slope (B) of the concentration-dependent viscosity curve for the mAbs. “x” indicates that a specific set of inputs was used for the model, while indicates that the set of input parameters was not used. Experimental inputs refer to HIC retention time, kD and A2, while in silico input includes the data from the surface patch analysis which is the information of size, score and count of positive, negative and hydrophobic surface patches. Fv-charge and pl was included regardless of the selected inputs. The quality parameters of the models shown were obtained from plotting the values of A or B that were calculated from the measured viscosity against the predicted values of the respective models. R2 = coefficient of determination; SSE = Standard square error; RMSE = Root-mean- square error
Categorical classification
Using experimental and in silico data as inputs, models for categorical classification were created by training the models on whether the mAbs show a viscosity above a threshold of 15 mPa*s at a certain concentration. This was done for concentrations of 120, 150 and 180 mg/mL, where of the 25 mAbs used in the model creation at 120 mg/mL 3 mAbs, at 150 mg/mL 6 mAbs and at 180 mg/mL 15 mAbs exhibited a viscosity of above 15 mPa*s. The confusion matrices of these models are shown in table 2 below, it is apparent that both the training sets and validation sets contained problematic (in this experiment defined as having a viscosity > 15 mPa*s) and unproblematic (in this experiment defined as having a viscosity < 15 mPa*s) mAbs. The models created have excellent predictive power regardless of the inputs used, with all of them exhibiting a misclassification rate of 0 (Table 8). To further evaluate the predictive power of these models, the two mAbs, which were neither used in the training nor validation sets, were chosen for verification. mAb 26 shows unproblematic behavior at 120 and 150 mg/mL, but exceeds 15 mPa*s at 180 mg/mL, which was correctly predicted by the models disclosed herein (No/No/Yes). mAb 27 shows no problematic behavior at either of the concentrations which was again correctly predicted by the models presented herein (No/No/No).
Table 2: Confusion matrices of the categorical models using experimental and in silico data as inputs. Results of the validation sets are in brackets.
Viscosity curve prediction
To obtain more extensive information of mAbs' viscosity the full concentrationdependent viscosity curves were predicted, or more specifically, the intercept (A) and slope (B) of the linearized exponential function. Using the predicted values for A and B, theoretical viscosity curves can be constructed and compared to the actual measured values, examples of such comparison are in Figure 6 A-D.
Despite not exactly matching the actual values, the predicted values are very close and the predicted curve reflects the actual concentration-dependent viscosity in a similar fashion. The percentage and absolute differences of predicted compared to measured viscosities of all mAbs is presented in Figure 7. Calculated average differences across all 27 mAbs between predicted and measured viscosity values are in table 3. The average absolute difference, calculated to measured, in mPa*s is between 0.1 -4.1 mPa*s, with a gradual increase of difference with increasing protein concentration. The relative % difference are between 8.2 - 26.2 %, but follow a curved function, with the highest differences at concentrations of 90-120 mg/mL and decreasing difference with decreasing and increasing protein concentration.
Table 3: Comparison of predicted and measured viscosity (includes data from all mAbs in this study)
For verification of the model again mAb 26 and 27 were used, in Figure 6 E and F the predicted viscosity curves for both mAbs are drawn. While the course of the predicted curve for mAb 27 (Figure 6 F) seems very similar to the measured values, the predicted curve for mAb 26 (Figure 6 E) doesn’t correctly reflect the steepness of the measured curve above 150 mg/mL.
Using computational data and in silico modelling has the advantage of being relatively easily accessible and do not require material and laboratory work.
The present invention provides artificial neural networks and methods to predict and determine the viscosity of proteins such as mAbs in solutions. The models can be used to predict a viscosity categorization, above or below a given threshold of e.g. 15 mPa*s, and to predict viscosity curves. Whilst already the categorial classification is good, the models for viscosity curve prediction show a high power and good viscosity curve forecast. The use of a multitude of input variables, derived from experimental data as well as from computational data and in silico modeling, is advantageous. Monoclonal antibodies mAbs were obtained. Double gene vectors containing the heavy and light chains were transfected into CHOK1 SV GS-KO cells and cultured under selection conditions as stable pooled cultures. Clarified supernatant was obtained by centrifugation followed by filter sterilization using 0.22 pm filters. Protein A chromatography was used for mAb purification. All proteins were concentrated to a final concentration of 10 mg/mL, and the buffer exchanged into the formulation buffer (protein solution) (20 mM histidine-HCI, pH 6.0) by tangential flow filtration. mAbs were of different subtypes IgG 1 or lgG2 (see tables 4 and 5 below).
Buffer
All lab experiments described (HIC, DLS, SLS, viscosity) were performed in the buffer that serves as the solution for the protein (and later for the drug product), 20 mM histidine-HCI, pH 6.0.
Protein concentration
For concentration determination an Agilent Cary 60 UV-Spectrophotometer with a variable path length extension SoloVPE was used. For each measurement 30 pL of sample were loaded into a cuvette and measured at 280 nm using the appropriate specific extinction coefficient.
Hydrophobic interaction chromatography
The hydrophobic surface properties of all mAbs were determined by hydrophobic interaction chromatography (HIC). Proteins were analyzed at 10 mg/mL in formulation buffer, 5 pL were injected on a ProPac Hic-10 column (ThermoScientific) and separated using a Waters HPLC system. The start condition of 95 % mobile phase A (1 M ammonium sulfate in 20 mM sodium phosphate pH 7.0) was linearly reduced over 39 min to 95 % mobile phase B (20 mM sodium phosphate pH 7.0). Flow rate was set to 1 mL/min at a column temperature of 24 °C. Dynamic and static light scattering
Dynamic light scattering (DLS) and static light scattering (SLS) measurements were performed on a DynaPro PlateReader III (with software Dynamics; Wyatt Technologies). Stock solutions of the antibodies were filtered through 0.22 pm PVDF filters (Millex GV) and serial dilutions with seven concentrations from 10-2 mg/mL protein with the formulation buffer were prepared. Samples were transferred to 384 well plates (Aurora) in triplicate and the plates were centrifuged at 750*g for 2 min to remove air bubbles. The temperature for the measurement was set at 25 °C. Laser power was set to 20 % and attenuation level to 0 %, 20 acguisitions of 5 s length were made for each well. Assessment of the diffusion interaction parameter kD (mL/g) was performed via DLS. The mutual diffusion coefficient Dm (m2/s) was plotted against the protein concentration (g/mL) and kD was obtained from the slope of a linear fit. The second virial coefficient A2 (mol*mL/g) was obtained from SLS measurements. Calibration of the plates was performed using Dextran (Sigma) with a predetermined molecular weight of 36.9 ± 0.1 kDa. Solvent offsets were measured in triplicates for the formulation buffer. The reciprocal molecular weight (mol/g) was plotted against the protein concentration (g/mL) and A2 was obtained from the slope of a linear fit.
In silico modelling patches of mAb Fv
Modelling of mAbs was performed using software BioLuminate (version 3.80; Schrodinger, LLC, New York, NY). Homology modelling of the Fv region was done by use of the antibody prediction tool. Framework templates for isotypes IgG 1 or IgG 2 were selected based on the highest composite score from the pdb database. The best CDR loop cluster was selected automatically. For modelling the standard presets of the software were kept, except the pH was set to 6.0 to represent the experimental settings. The surface of the modelled mAb Fv-regions were analyzed with the protein surface analyzer tool in the BioLuminate software, to obtain size (in A2), score (sum of all contributing patch scores associated with this patch), and type (pos = positively charged; neg = negatively charged; hyd = hydrophobic) of the patches. The descriptors derived from modelling are thus score pos/neg/hyd Fv total, size pos/neg/hyd Fv total, count pos/neg/hyd Fv total. Fv charge calculation
To calculate the Fv charge, the variable heavy- and variable light-chains of each antibody were analyzed with the prot pi protein tool (https://www.protpi.ch/Calculator/ProteinTool). The two chains were each defined as a subunit of the entire protein. The set modifier for post translational modifications was global disulfide bridges for the cysteine residues in the Fv. Charge was calculated at pH 6.0. pl calculation
The pl was calculated "manually" from the pKa values of the amino acid residues of the primary sequence.
Rheometry and viscosity descriptors mAbs were concentrated to approximately 180 mg/mL using spin filters with a 30 kDa molecular weight cut off and then diluted to six concentrations from 180-30 mg/mL. Concentration-dependent viscosity data were generated using a VROC viscometer (Rheosense) for each concentration. To obtain descriptive information of the concentration-dependent viscosity of the mAb samples, the experimentally measured viscosity data were processed based on above equations (3) to (5).
Above equation (3) was used to calculate the relative viscosity of the mAb samples. For doing so, at first, the viscosity of the buffer was measured to be 0.92 mPa*s at 25°C. Knowing the viscosity of the buffer, the relative viscosity of the mAb samples were then calculated based on the measured viscosity values according to equation (3). Accordingly, value pairs for concentration and relative viscosity were determined for each mAb sample.
Above equation (4) was used to describe the exponential concentration-dependent viscosity of mAb solutions. This equation can be linearized to obtain the intercept A and the slope B using the natural logarithm as in equation (5).
Based on the determined value pairs for concentration and relative viscosity of the mAB samples, the viscosity descriptors A and B for each mAb was then determined based on above equation (4) and (5) by applying least square fitting. Data analyses and artificial neural network modelling
The ANN model creation follows the approach illustrated in Figure 8. All input parameters (experimental data, calculated values from sequence, in silico-demed data, and viscosity descriptors) are listed in table 4 and 5. Each model is trained on a categorical response aiming to identify mAbs which show a viscosity value above the threshold value of 15 mPa*s at a distinct protein concentration. The ANNs are generated using software JMP v.16.0.0 (SAS Institute Inc.). The activation function used for all nodes is the tan h-function, which transforms values to be between -1 and 1 . For all models one hidden layer is sufficient with the number of nodes being four. To prevent the network from overfitting the model and in turn losing predictive power the data was split into a training and a validation set. The method used was K-fold where the 25 mAb and their data sets are split into K sets. Each of the K sets contained the data sets of all 25 mAb, but in each K set the two subsets training set and validation set was different, the mAb were distributed randomly in each K set to the two subsets training set and validation set. Each of the K sets was used to validate the model fit on the rest of the data, fitting a total of K sets. The value of K was set to 5 for each model. The distribution of the mAb to the K sets was also not correlated between the three models but was done randomly, so each model has a different K set. The model reported by the JMP software is based on the best log likelihood. The quality of the ANNs was determined using the coefficient of determination (R2), the Standard Square Error (SSE) and the Root- mean-square Error (RMSE) for training and validation datasets.
Table 4: Data included in artificial neural network modelling. Experimental inputs (retention time in hydrophobic interaction chromatography (HIC); diffusion interaction parameter kD; second virial coefficient A2) and computational inputs derived input from amino acid sequence (isoelectric point (pl) and Fv-charge).
Table 5; Data included in artificial neural network modelling. In silico simulation- derived input (scores and areas for positively and negatively charged as well as hydrophobic patches). The viscosity descriptors intercept A and slope B derived from linearization of viscosity data (original viscosity data in tables 6 and 7). Data are fed into an artificial neural network following the scheme in Figure 8.
Table 6: Viscosity raw data. For every mAb a dilution series in 20 mM histidine-HCI pH 6 buffer was prepared with target concentrations (target c) of 180, 150, 120 mg/mL. The table contains values of dynamic viscosity in mPa*s ± SD of 10 measurements as well as the measured protein concentration in mg/mL.
Table 7: Viscosity raw data. For every mAb a dilution series in 20 mM histidine-HCI pH 6 buffer was prepared with target concentrations (target c) of 90, 60, 30 mg/mL. The table contains values of dynamic viscosity in mPa*s ± SD of 10 measurements as well as the measured protein concentration in mg/mL.
As set forth above, the viscosity descriptors intercept A and slope B are calculated from the values shown in tables 6 and 7. For doing so, at first, viscosity of the buffer is determined which was measured to be 0.92 mPa*s. Then, the buffer viscosity is used to calculate the relative viscosity for each mAb and for each value pair for concentration and viscosity in tables 6 and 7 based on above equation (3). Thereafter, a least square fit (or any other mathematical procedure for finding a relation, in particular for finding a best-fitting curve, to given sets of points) is done with the 6 value pairs for concentration and viscosity (with the relative viscosity) for each mAb, thereby fitting the 6 value pairs to the exponential function of above equation (4). Then, the calculated function is linearized as shown in above equation (5) to provide the A and B values for each mAb which are included in table 5.
Table 8: Regardless of the level of input variables, i.e. experimental only, in silico only, or both types of variables (in all three models Fv-charge and pl was included regardless of the selected inputs), models derived from artificial neural networks are able to correctly classify whether a certain mAb above a certain protein concentration threshold value may show a problematic behavior, e.g. a viscosity above 20 mPa*s.

Claims

Claims
1 . Method for providing a computer-implemented neural network (10) configured for predicting a concentration-dependent viscosity of a protein solution in dependence on a plurality of input parameters (IP) associated to said protein solution, comprising the steps of:
- providing a plurality of training data sets (ST), each of which is associated to a specific protein solution and comprises a plurality of input parameters (IP) being indicative of the specific protein solution and at least one associated output parameter (OP) being indicative of a concentration-dependent viscosity of the specific protein solution; and
- training a neural network based on the training data sets (ST) to provide the computer-implemented neural network (10) configured for predicting the concentration-dependent viscosity, wherein the input parameters (IP) include at least one of: i) experimental data; ii) computational data; and iii) in silico data.
2. Method according to claim 1 , wherein the protein solution comprises a therapeutic protein, preferably an antibody.
3. Method according to claim 2, wherein the antibody is a monoclonal antibody (mAb), preferably a monoclonal antibody of lgG1 or lgG2 subtype.
4. Method according to any one of claims 1 to 3, wherein the protein solution comprises an aqueous medium including water, preferably a buffer, more preferably a Histidine-HCI buffer.
5. Method according to any one of claims 1 to 4, wherein the concentrationdependent viscosity represents a value (T]CS) associating a viscosity of the protein solution to a specific protein concentration of the protein solution.
6. Method according to any one of claims 1 to 4, wherein the concentrationdependent viscosity represents a function in particular a mathematical function, associating a viscosity of the protein solution to a protein concentration of the protein solution.
7. Method according to claim 6, wherein the concentration-dependent viscosity is indicative of the following function: wherein fv refers to a function associating viscosity of the protein solution to a protein concentration; c refers to a protein concentration; A and B refer to constants.
8. Method according to claim 7, wherein the concentration-dependent viscosity of a protein solution is represented by at least one of constants A and B.
9. Method according to any one of claims 1 to 8, further comprising a step of providing the neural network to be trained which comprises: an input layer (12) having a plurality of input nodes (IN), each of which receives one input parameter (IP); at least one hidden layer (14), in particular only one hidden layer (14), having a plurality of hidden nodes (HN), in particular four or more hidden nodes (HN); and an output layer (16) having at least one output node (ON) for providing at least one output parameter (OP) being indicative of the concentration-dependent viscosity.
10. Method according to claim 9, wherein the hidden nodes (HN) use tan h as an activation function.
11 . Method according to claim 9 or 10, wherein the at least one output parameter provided by the at least one output node (ON) indicates: a viscosity value ( cs) associating a viscosity of the protein solution to a specific protein concentration of the protein solution; or at least one of the constants A and B of the following function: wherein fv refers to a function associating viscosity of the protein solution to a protein concentration; c refers to a protein concentration; A and B refer to constants.
12. Method according to any one of claims 1 to 11 , wherein the input parameters (IP) are indicative of protein-protein-interactions.
13. Method according to any one of claims 1 to 12, wherein the neural network is an artificial neural network.
14. Method according to any one of claims 1 to 13, further comprising a step of validating the trained neural network (10) based on a plurality of validation data sets (Sv), each of which comprises a plurality of input parameters (IP) and at least one associated output parameter (OP).
15. Computer-implemented method for predicting a concentration-dependent viscosity of a protein solution by using a neural network, wherein the neural network is configured for predicting a concentration-dependent viscosity of the protein solution in dependence on a plurality of input parameters associated to said protein solution, and wherein the input parameters (IP) include at least one of: i) experimental data; ii) computational data; and iii) in silico data.
16. Method according to any one of claims 1 to 15, wherein the input parameters (IP) include at least one of i) experimental data, ii) computational data or iii) in silico data, wherein said experimental data i) are selected from apparent surface hydrophobicity measured by hydrophobic interaction chromatography (HIC), diffusion interaction parameter (kD), second virial coefficient (A2) and combinations thereof; said computational data ii) are selected from the protein’s isoelectric point (pl), variable fragment (Fv)-charge (Fv-charge), and combinations thereof; and said in silico data iii) are selected from hydrophobic and charged patch sizes, in particular score pos Fv total, size pos Fv total, count pos Fv total, score neg Fv total, size neg Fv total, count neg Fv total, score hyd Fv total, size hyd Fv total, and count hyd Fv total.
17. Method according to any one of claims 1 to 14 or 16, wherein the input parameters (IP) are selected from said ii) computational data and said iii) in silico data.
18. Method for determining a concentration-dependent viscosity of a protein solution, comprising: a step of predicting a concentration-dependent viscosity of the protein solution by using a computer-implemented neural network (10); and a step of comparing the concentration-dependent viscosity with a target viscosity.
19. Method according to claim 18, wherein the step of comparing the concentration-dependent viscosity with the target viscosity is performed to determine suitability as a drug product of the protein solution.
20. Method for providing a drug product comprising a protein solution, the method comprising: a step (S2) of predicting a concentration-dependent viscosity of a protein solution by using a computer-implemented neural network (10); a step (S3) of comparing the concentration-dependent viscosity with a target viscosity to determine suitability of the protein solution as a drug product; and optionally a step (S5) of preparing the drug product if suitability of the protein solution is determined.
21 . Apparatus for predicting a concentration-dependent viscosity of an experimental protein solution, the apparatus comprising: an experimental protein solution; a computer component comprising a neural network configured to predict the viscosity of the experimental protein solution: wherein the neural network is trained using a plurality of predetermined data sets; wherein each of the plurality of predetermined data sets comprises (i) a plurality of input parameters indicative of at least one physical characteristic of a predefined protein solution, and (ii) at least one output parameter indicative of a concentration-dependent viscosity of the predefined protein solution; and wherein the plurality of input parameters comprise at least one of experimentally-derived data, computationally-derived data, and in silico-derived data.
22. A system for determining a concentration-dependent viscosity of a protein solution, the system comprising: a protein solution; and a neural network configured to determine the viscosity of the protein solution, wherein determining the viscosity of the protein solution depends on a plurality of input parameters (IP) associated with the protein solution, and further wherein the input parameters (IP) include at least one from the group consisting of (i) experimental data; (ii) computational data; and (iii) in silico data.
EP23801725.5A 2022-11-04 2023-11-03 Protein solutions Pending EP4612693A1 (en)

Applications Claiming Priority (6)

Application Number Priority Date Filing Date Title
EP22205657 2022-11-04
EP22207063 2022-11-11
EP22208344 2022-11-18
EP22209537 2022-11-25
EP23205886 2023-10-25
PCT/EP2023/080727 WO2024094879A1 (en) 2022-11-04 2023-11-03 Protein solutions

Publications (1)

Publication Number Publication Date
EP4612693A1 true EP4612693A1 (en) 2025-09-10

Family

ID=88731345

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23801725.5A Pending EP4612693A1 (en) 2022-11-04 2023-11-03 Protein solutions

Country Status (5)

Country Link
EP (1) EP4612693A1 (en)
JP (1) JP2025538138A (en)
KR (1) KR20250107843A (en)
CN (1) CN120129943A (en)
WO (1) WO2024094879A1 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119470879B (en) * 2025-01-14 2025-05-13 南京微测生物科技有限公司 Preparation method of ochratoxin A monoclonal antibody conjugate

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
NZ799831A (en) * 2020-11-02 2026-02-27 Regeneron Pharma Methods and systems for biotherapeutic development

Also Published As

Publication number Publication date
JP2025538138A (en) 2025-11-26
WO2024094879A1 (en) 2024-05-10
CN120129943A (en) 2025-06-10
KR20250107843A (en) 2025-07-14

Similar Documents

Publication Publication Date Title
Narayanan et al. Design of biopharmaceutical formulations accelerated by machine learning
Borman et al. Selection of analytical technology and development of analytical procedures using the analytical target profile
García-Quintanilla et al. Pharmacokinetics of intravitreal anti-VEGF drugs in age-related macular degeneration
Raut et al. Pharmaceutical perspective on opalescence and liquid–liquid phase separation in protein solutions
Glover et al. Compatibility and stability of pertuzumab and trastuzumab admixtures in iv infusion bags for coadministration
Grunst et al. Structure and inhibition of SARS-CoV-2 spike refolding in membranes
Dear et al. X-ray scattering and coarse-grained simulations for clustering and interactions of monoclonal antibodies at high concentrations
Khraishi et al. Biosimilars: a multidisciplinary perspective
Makowski et al. Reduction of monoclonal antibody viscosity using interpretable machine learning
Singh et al. Determination of protein–protein interactions in a mixture of two monoclonal antibodies
Prass et al. Viscosity prediction of high-concentration antibody solutions with atomistic simulations
Bramham et al. Comprehensive assessment of protein and excipient stability in biopharmaceutical formulations using 1H NMR spectroscopy
EP4612693A1 (en) Protein solutions
Zhang et al. Asymmetric structures and conformational plasticity of the uncleaved full-length human immunodeficiency virus envelope glycoprotein trimer
Shrivastava et al. Rapid estimation of size-based heterogeneity in monoclonal antibodies by machine learning-enhanced dynamic light scattering
Tan et al. Intraglomerular crosstalk elaborately regulates podocyte injury and repair in diabetic patients: insights from a 3D multiscale modeling study
Volkin et al. Two decades of publishing excellence in pharmaceutical biotechnology
Kallewaard et al. Functional maturation of the human antibody response to rotavirus
Rojekar et al. Exploring protein aggregation in biological products: From mechanistic understanding to practical solutions
Kumar et al. Concentration-dependent diffusion of monoclonal antibodies: underlying mechanisms of anomalous diffusion
Kessler et al. Biomarkers to predict the success of treatment with the intravitreal 0.19 mg fluocinolone acetonide implant in uveitic macular edema
Torisu et al. Physicochemical characterization of sabin inactivated poliovirus vaccine for process development
Yang et al. Trimerization dictates solution opalescence of a monoclonal antibody
Thorn et al. Assessing the impact of viscosity lowering excipient on liquid-liquid phase separation for high concentration monoclonal antibody solutions
HK40124248A (en) Protein solutions

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250603

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)