EP4732294A1 - Methods and systems for determining an impact of attributes - Google Patents
Methods and systems for determining an impact of attributesInfo
- Publication number
- EP4732294A1 EP4732294A1 EP24743145.5A EP24743145A EP4732294A1 EP 4732294 A1 EP4732294 A1 EP 4732294A1 EP 24743145 A EP24743145 A EP 24743145A EP 4732294 A1 EP4732294 A1 EP 4732294A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- data
- clinical
- computer
- attributes
- pharmaceutical
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/50—Molecular design, e.g. of drugs
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/90—Programming languages; Computing architectures; Database systems; Data warehousing
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H10/00—ICT specially adapted for the handling or processing of patient-related medical or healthcare data
- G16H10/20—ICT specially adapted for the handling or processing of patient-related medical or healthcare data for electronic clinical trials or questionnaires
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H20/00—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance
- G16H20/10—ICT specially adapted for therapies or health-improving plans, e.g. for handling prescriptions, for steering therapy or for monitoring patient compliance relating to drugs or medications, e.g. for ensuring correct administration to patients
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16H—HEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
- G16H50/00—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
- G16H50/50—ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for simulation or modelling of medical disorders
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Public Health (AREA)
- General Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Epidemiology (AREA)
- Primary Health Care (AREA)
- Theoretical Computer Science (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- Medicinal Chemistry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Crystallography & Structural Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Data Mining & Analysis (AREA)
- Biomedical Technology (AREA)
- Pathology (AREA)
- Physics & Mathematics (AREA)
- Pharmacology & Pharmacy (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Software Systems (AREA)
- Investigating Or Analysing Biological Materials (AREA)
- Medical Treatment And Welfare Office Work (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
A computer-implemented method for determining a correlation status associated with a clinical impact of one or more material attributes, the method comprising: obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and determining, via the one or more processors, the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
Description
METHODS AND SYSTEMS FOR DETERMINING AN IMPACT OF ATTRIBUTES
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Provisional Patent Application Serial No. 63/522,493 filed on June 22, 2023, entitled “METHODS AND SYSTEMS FOR DETERMINING AN IMPACT OF ATTRIBUTES,” which is herein incorporated by reference in its entirety.
FIELD
[0002] Embodiments herein relate to methods and systems for determining an impact of attributes, more specifically, automatically determining a clinical impact of material attributes of pharmaceutical materials.
BACKGROUND
[0003] The native structure of chemistry of biological molecules (such as therapeutic proteins) adapts or alters in response to changes within the molecules’ environment. Other pharmaceutical material, including biological therapies, nucleic acid and cell-based therapies can also undergo changes within their environment. Although this flexibility in structure or chemistry is required for the biological function of most, if not all, biological molecules and cells, it also presents many challenges during the development and manufacture of pharmaceutical material for pharmaceutical applications. For example, therapeutic proteins endure various conditions during the many process steps that lead up to being administered to a patient. The many process steps include, for instance, one or more of protein production (e.g., recombinant production), harvest, purification, formulation, filling, packaging, storage, delivery, and final preparation immediately prior to administration to the patient. During each of these steps, a therapeutic protein is placed in one or more environments that may or may not lead to a change in its structure or chemistry. The change in structure or chemistry can lead to the formation of different species of the pharmaceutical material that results in a heterogeneous product. While some species retain their ability to bind to their targets and therefore maintain therapeutic efficacy, others lose target binding ability and thus become functionally inactive. In order to maximize and maintain quality control of these pharmaceutical material, the biopharmaceutical industry has focused much effort to understand why some species lose activity while others retain activity.
[0004] Material attributes comprise the physicochemical property of therapeutic biological molecules and can therefore impact the drug safety and efficacy. The levels of attributes critical to the drug quality, or critical quality attributes (CQAs), are explicitly defined by the product purity specifications subject to extensive regulatory reviews.
SUMMARY
[0005] In one aspect, a computer-implemented method for determining a correlation status associated with a clinical impact of one or more material attributes comprises: (a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (d) determining, via the one or more processors, the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
[0006] In some embodiments, the one or more material attributes comprise at least one of a molecular attribute, a process related impurity, or a drug substance feature. In some embodiments, the one or more material attributes comprise a molecular attribute, and wherein the molecular attribute comprises at least one of: acidic species, basic species, high molecular weight species, subvisible particle number, visible particles, aggregation, low molecular weight, middle molecular weight, glycosylation, glycation, deamidation, deamination, cyclization, oxidation, sulfation, hydroxylysine, isomerization, fragmentation/clipping, N- terminal and C-terminal variants, signal peptide, reduced and partial species, misfolding, disulfide scrambling , domain swapping, folded structure, surface hydrophobicity, chemical modification, glycation, covalent bond, mutations or misincorporations, a C-terminal amino acid motif PARG, a C-terminal amino acid motif PAR- Amide , drug antibody ratio (DAR), or peptide antibody ratio (PAR). In some embodiments, the one or more material attributes
comprise a process related impurity, and wherein the process related impurity comprises at least one of CHOP, HCP, residual host cell DNA, residual ProA, or a process reagent. [0007] In some embodiments, the one or more material attributes comprise a drug substance feature, and wherein the drug substance feature comprises at least one of a component feature or a drug administration feature. In some embodiments, the pharmaceutical material comprises at least one of a biological therapy, a synthetic small molecule, or a nucleic acid. In some embodiments, the pharmaceutical material comprises a biological therapy, and wherein the biological therapy is selected from the group consisting of: an antibody, an antigen-binding antibody fragment, an antibody protein product, a Bi-specific T cell engager (BiTE®) molecule, a bispecific antibody, a trispecific antibody, an Fc fusion protein, a recombinant protein, a recombinant virus, a recombinant T cell, a synthetic peptide, and an active fragment of a recombinant protein. In some embodiments, the pharmaceutical material comprises a nucleic acid, and wherein the nucleic acid comprises an siRNA, an mRNA or a DNA.
[0008] In some embodiments, the pharmaceutical material comprises a biological therapy, and wherein manufacturing the pharmaceutical material comprises culturing a genetically engineered mammalian host cell comprising one or more nucleic acids encoding the biological therapy. In some embodiments, the pharmaceutical material is in a pharmaceutically acceptable formulation. In some embodiments, the measurement data of the one or more material attributes at the one or more timepoints is determined by at least one of mass spectrometry, chromatography, electrophoresis, spectroscopy, light obscuration, a particle method, analytical centrifugation, imaging or imaging characterization, or immunoassay. In some embodiments, the material attribute data comprises a change of the measurement data of the one or more material attributes at the one or more time points. In some embodiments, the material attribute data comprises a duration that the pharmaceutical material was under storage conditions prior to the administration of the pharmaceutical material.
[0009] In some embodiments, the material attribute data comprises a dose of the pharmaceutical material in the administration. In some embodiments, the material attribute data comprises a level of material attribute exposure received by the one or more subjects at the time of the administration. In some embodiments, the one or more timepoints comprise at least one of a timepoint of manufacture or a timepoint of lot release. In some embodiments, the measurement data of the one or more material attributes is detected at two or more
timepoints under storage conditions. In some embodiments, the one or more timepoints comprise a timepoint of manufacture and at least two subsequent timepoints. In some embodiments, the one or more clinical events comprise one or more clinical adverse events associated with the one or more subjects that have received the administration of the pharmaceutical material.
[0010] In some embodiments, the subject data comprises at least one of preexisting condition, biomarker, laboratory result, metabolomic data, or demographic information of the one or more subjects that have received the administration of the pharmaceutical material. In some embodiments, the one or more transformations comprise cleaning the material attribute data or the clinical data, and wherein the cleaning comprises filtering the material attribute data or the clinical data based on one or more filtering criteria. In some embodiments, the one or more filtering criteria comprise removing one or more orphan adverse events. In some embodiments, the one or more transformations comprise merging the material attribute data and the clinical data, and wherein the merging comprises combining the material attribute data and the clinical data based on a level of similarity between the material attribute data and the clinical data. In some embodiments, the one or more transformations comprise associating the material attribute data and the clinical data, and wherein the associating comprises correlating the material attribute data and clinical data based on one or more associating criteria.
[0011] In some embodiments, the one or more transformations comprise selecting the material attribute data or the clinical data, and wherein the selecting comprises selecting the material attribute data or the clinical data based on one or more selection criteria. In some embodiments, the one or more transformations comprise grouping the material attribute data or the clinical data, and wherein such grouping comprises generating one or more subgroups based on one or more patterns of the material attribute data and the clinical data. In some embodiments, the computational model comprises at least one of: a logistic regression model, a support vector machine model, a multinomial logistic regression model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, or a Markov model. In some embodiments, determining the correlation status comprises using the computational model to identify one or more patterns associated with the modified material attribute data and the modified clinical data.
[0012] In some embodiments, the correlation status comprises a presence of correlation or an absence of correlation. In some embodiments, the computer-implemented method further comprises, if the correlation status comprises an absence of correlation, determining that the one or more material attributes impact neither clinical safety nor efficacy of the pharmaceutical material. In some embodiments, the computer-implemented method further comprises, if the correlation status comprises a presence of correlation, determining that the one or more material attributes impact at least one of safety or efficacy of the pharmaceutical material. In some embodiments, the computer-implemented method further comprises, if the correlation status comprises an absence of correlation, setting a specification for permissible levels of the one or more material attributes of the pharmaceutical material, wherein the permissible levels of the one or more material attributes are based on one or more levels of material attribute exposure received by the subjects.
[0013] In some embodiments, the computer-implemented method further comprises, if the correlation status comprises a presence of correlation, setting a specification for maximum permissible levels of the one or more material attributes of the pharmaceutical material, wherein said maximum permissible levels of the one or more material attributes are based on one or more levels of the one or more material attributes associated with at least one of clinical adverse events or inhibition of efficacy of the pharmaceutical material. In some embodiments, the computer-implemented method further comprises, if the correlation status comprises an absence of correlation, manufacturing production lots of the pharmaceutical material comprising the one or more material attributes at or below a specified permissible level of the one or more material attributes based on one or more levels of material attribute exposure. In some embodiments, the computer-implemented method further comprises, if the correlation status comprises a presence of correlation, setting a specification for levels of the one or more material attributes at a time of manufacture that do not exceed a maximum permissible level of the one or more material attributes of the pharmaceutical material.
[0014] In some embodiments, the computer-implemented method further comprises, if the correlation status comprises an absence of correlation, establishing a manufacturing process to produce levels of the one or more material attributes at or below a permissible level based on the correlation status. In some embodiments, the computer-implemented method further comprises generating a rank of the one or more material attributes based on the correlation status using the computational model. In some embodiments, the computer-implemented method further comprises selecting a subset of the one or more material attributes based on
the rank of the one or more material attributes and setting a specification for permissible levels of the subset of the one or more material attributes. In some embodiments, the computer-implemented method further comprises generating one or more heuristics associated with the one or more subjects based on the modified material attribute data and the modified clinical data.
[0015] In some embodiments, the computer-implemented method further comprises estimating one or more parameters associated with the administration of the pharmaceutical material, wherein the one or more parameters comprise at least one of a stability expanded time-course dosing, a pharmacokinetic expanded time-course dosing, a step dosing, or a multiple dose overlap time-course dosing. In some embodiments, the computer-implemented method further comprises estimating a cooperative effect of the one or more material attributes, wherein the cooperative effect comprises at least one of an additive effect, an inhibitory effect, or a feedback effect.
[0016] In another aspect, a computer-implemented method for training a computational model for determining a correlation status associated with a clinical impact of one or more material attributes comprises: (a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (d) training, via the one or more processors, the computational model using the modified material attribute data and the modified clinical data.
[0017] In another aspect, a computer-implemented method for determining a correlation status associated with a clinical impact of one or more material attributes comprises: (a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical
material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) generating, via the one or more processors, predictive data associated with the pharmaceutical material based on the material attribute data and the clinical data using a predictive model; (d) applying, via the one or more processors, one or more transformations to the material attribute data, the clinical data, and the predictive data to create modified material attribute data, modified clinical data, and modified predictive data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (e) determining, via the one or more processors, the correlation status based on at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data using a computational model.
[0018] In another aspect, a computer system for determining a correlation status associated with a clinical impact of one or more material attributes comprises a memory storing instructions; and one or more processors configured to execute the instructions to perform operations including: (a) obtaining material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) applying one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (d) determining the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
[0019] In another aspect, a non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions for performing a method of determining a correlation status associated with a clinical impact of one or more material attributes comprises: (a) obtaining material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining clinical data associated with the pharmaceutical material, wherein the clinical
data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) applying one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (d) determining the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
BRIEF DESCRIPTION OF DRAWINGS
[0020] The skilled artisan will understand that the figures, described herein, are included for purposes of illustration and are not limiting on the present disclosure. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure. It is to be understood that, in some instances, various aspects of the described implementations may be shown exaggerated or enlarged to facilitate an understanding of the described implementations. In the drawings, like reference characters throughout the various drawings generally refer to functionally similar and/or structurally similar components.
[0021] FIG. 1 is a block diagram of an exemplary system 100 for determining an impact of attributes, in accordance with some embodiments of the technology described herein.
[0022] FIG. 2 is a flowchart of an exemplary process 200 for determining an impact of attributes, in accordance with some embodiments of the technology described herein.
[0023] FIG. 3 is a diagram depicting an illustrative technique 300 for determining an impact of attributes, in accordance with some embodiments of the technology described herein.
[0024] FIG. 4 depicts an exemplary histogram showing number of different adverse events, in accordance with some embodiments of the technology described herein.
[0025] FIG. 5 is a diagram depicting an exemplary statistical T-test plot, in accordance with some embodiments of the technology described herein.
[0026] FIG. 6 is a diagram depicting an exemplary dimensionality analysis of principal component analysis (PCA), in accordance with some embodiments of the technology described herein.
[0027] FIG. 7 is a diagram depicting another exemplary dimensionality analysis of principal component analysis (PCA), in accordance with some embodiments of the technology described herein.
[0028] FIG. 8 is a diagram depicting an exemplary data table, in accordance with some embodiments of the technology described herein.
[0029] FIG. 9 is a diagram depicting two performance metrics of a trained neural network, in accordance with some embodiments of the technology described herein.
[0030] FIG. 10 is a flowchart of another exemplary process 1000 for determining an impact of attributes, in accordance with some embodiments of the technology described herein. [0031] FIG. 11 is a diagram depicting an exemplary dimensionality analysis of two- dimensional linear discriminant analysis (2D-LDA), in accordance with some embodiments of the technology described herein.
[0032] FIG. 12 is a schematic diagram of an illustrative computing device with which aspects described herein may be implemented.
DETAILED DESCRIPTION
[0033] Described herein comprise methods and systems for determining a correlation status associated with a clinical impact of one or more material attributes, and methods of manufacturing a pharmaceutical material (such as a therapeutic protein, nucleic acid, or therapy comprising cells) in which safety and efficacy of the pharmaceutical material are controlled by limiting levels of attribute exposure to a subject. After a production lot of a pharmaceutical material is manufactured, the pharmaceutical material can be stored in formulation for some period of time before it is administered to a subject. During that time in storage, levels of material attributes of the pharmaceutical material may change.
Additionally, levels of material attributes may vary between different productions lots, for example reflecting differences in material attribute levels at different stages of production between initial cell culture and final pharmaceutical material of the pharmaceutical material. For example, levels of attributes such as acidic species, basic species, high molecular weight species, amino acid isomers, or subvisible particles may increase. Such attributes may cause a decrease in efficacy of the pharmaceutical material, and/or may cause adverse events in a subject receiving the pharmaceutical material. As described herein, changes in levels of material attributes in storage can be modeled. Based on the level of material attribute at the time a production lot of a pharmaceutical material is manufactured, the amount of time that the pharmaceutical material in formulation is in storage before administration to a subject, the rate of change of the material attribute in storage, and the dosage of the pharmaceutical material, the level of material attribute at the time of administration to the subject can be
calculated. Furthermore, if the actual level of material attribute at the time of administration is not associated with adverse events or a loss of efficacy, it can be determined that that level of material attributes is safe and effective. Accordingly, using methods described herein, production lots of the pharmaceutical material can be produced at or below specified levels based on the level that is deemed to be safe and effective at the time of administration.
[0034] It is contemplated that while criticality and manufacturing specifications of material attributes have been determined using conventional methods, conventionally determined criticality and specifications for material attributes can present little clinical relevance. The criticality is generally investigated using non-human model systems, and the specification limits reflect the low levels of attributes reasonably achievable upon manufacture and storage. Alternatively, in a prior-knowledge or clinical-experience approach, an attribute is proposed to be not critical when clinical consequences are rarely observed for other pharmaceutical materials containing the attribute. This approach, however, ignores the possible product-specific variability in attribute impact. Also, adverse events can be caused by an attribute and still appear rare, if the lots that contain a high enough level of the attribute to cause the events are infrequently distributed in clinics due to lot-to-lot variability.
[0035] Methods and systems described herein can utilize a data analysis approach that tests whether a correlation exists between the estimated actual level of patient exposure to a given attribute and the extent of the occurrence of a clinical outcome by analyzing the data from clinical studies and product or product-related quality analysis studies. This approach may be referred to as Clinical Impact of Attributes (CIA). The methods and systems disclosed herein associated with CIA can be referred to as CIA tools. CIA can assess any quantifiable pharmaceutical material (or pharmaceutical material-related) or pharmaceutical material manufacturing process characteristic (e.g., host cell protein which is a co-eluting compound that is not the pharmaceutical material itself), and manufacturing process parameters (e.g., a different hold time for a molecule, different temperature used for cell incubation, different parameters (e.g. supplier) for raw materials). Additionally, CIA tools can enable the use of the methods and systems disclosed herein to be applied to expanded non-drug attributes and/or non-in-vivo attributes. The non-drug attributes and/or non-in-vivo attributes can comprise preconditions (e.g., pre-existing diseases, genetics, age), manufacturing site, syringe type, container defects, hospital/clinical trial location, environmental conditions such as time of year and weather and other spatiotemporal events which can have baseline implications on adverse events. For example, if the weather is continually hot in some location, baseline
levels of headache or fatigue may be greater than a different location or during a less hot time period, and as such can be subtracted from analytically-derived assessments.
[0036] The methods and systems described herein can utilize actual clinical data and assessments of levels of attribute exposure at the time that a pharmaceutical material is administered. By applying one or more transformations to the material attribute data and clinical data, compared with conventional method, the methods described herein can provide data structures not feasibly generated or manipulated by manual methods to train, validate or test computational models, automate database generation or data updates, and provide a logical and computationally storable quantitative representation of multi-modal data, making the model training, validating and testing process more time efficient and less costly. These methods provide a real-world assessment of the impact of material attributes, overcoming shortcomings of conventional approaches that have utilized model systems without considering human factors for determining attribute levels and impact, and which have not accounted for attribute exposure at the time of administration.
[0037] Conventional approaches involve manual data processing using copy and paste actions done by a scientist, and the output often involves simple statistical method to compare two population distributions based on the variance between the two means of the populations. The methods and systems disclosed herein (e.g., CIA tools) can automate the process of data organizing and cleaning, all function application, tabulation, iterations, additional cleaning, and visualization without any manual copy, paste, and transpose operations. This automation can prevent human error in the process and improve the quality of the data. Additionally, the methods and systems disclosed herein (e.g., CIA tools) can be a platform on which diverse types of computational, mathematical, and statistical analytical methods can be implemented and tested to provide analytical arguments beyond a statistical one. The diverse types of computational, mathematical, and statistical analytical methods can include, but are not limited to, classification, dimensionality manipulation, machine learning, deep learning, and stochastic modeling. The diverse types of computational, mathematical, and statistical analytical methods (e.g., pharmacokinetic modeling, stability modeling, emergent features, and baseline subtraction) can target data features that cannot be explored by conventional approaches. Examples of some of these computational, mathematical, and statistical analytical methods can include assessing classes of quantifiable relevant data associated with the pharmaceutical material or clinical trial (other than just pharmaceutical material attributes) such as treatment location, weather conditions, specific hospitals, socioeconomic
events, or any other quantifiable categorical or numerical metric associated with the pharmaceutical material or clinical trial.
[0038] Material attributes
[0039] “Material attributes” and variations of this root term has its ordinary and customary meaning as would be understood by one of ordinary skill in the art in view of this disclosure. It refers to a structure or a chemically or physically changed structure on a macromolecule, such as a protein or nucleic acid, and may be characterized in terms of its physicochemical identity or attribute type and location within the sequence of the macromolecule, e.g., the position of the amino acid on which the attribute is present. For example, asparagine and glutamine residues are susceptible to deamidation. A deamidated asparagine at position 10 of a therapeutic protein amino acid sequence is an example of an attribute. Exemplary material attribute types are described elsewhere herein. For conciseness, material attributes may be referred to herein simply as “attributes.” The levels of attributes critical to the drug quality, or critical quality attributes (CQAs), may be explicitly defined by the product purity specifications. These specifications are typically subject to extensive regulatory reviews. In some embodiments, a specification may set the permissible levels of one or more material attributes in the manufacture of a pharmaceutical material.
[0040] In some embodiments, the material attribute comprises or consists of one or more of acidic species, basic species, high molecular weight species, subvisible particle number, low molecular weight, middle molecular weight, glycosylation (such as non-glycosylated heavy chain or high mannose), deamidation, deamination, cyclization, oxidation, isomerization, fragmentation/clipping, N-terminal and C-terminal variants, reduced and partial species, folded structure, surface hydrophobicity, chemical modification, covalent bonds a C-terminal amino acid motif PARG, a C-terminal amino acid motif PAR- Amide, endotoxin, bioburden, viscosity, container closure integrity, clarity, color, or lyophilization cake appearance.
[0041] PARG is an alternative C-terminal variant of antibodies that may occur as a result of alternative splicing. It represents 4 amino acids (Proline, Alanine, Arginine, Glycine) where the “AR” was genetically inserted in the canonical IgG2 C-terminal sequence. PAR-amide is another C-terminal variant that results from further processing of PARG. It refers to the C- terminal glycine getting cleaved off of antibodies ending in PARG, leaving behind an amide group on the C-terminal Arginine.
[0042] In some embodiments, the material attribute comprises or consists of at least one of: acidic species, basic species, high molecular weight species, amino acid isomers, or subvisible particle number.
[0043] Techniques for detecting levels of material attributes
[0044] Any suitable analytical technique for detecting a material attribute may be used with the methods described herein. Techniques for detecting a material attribute include, but are not limited to mass spectrometry, chromatography, electrophoresis, spectroscopy, light obscuration, particle methods (nanoparticle/visible/micron-sized resonant mass or Brownian motion), analytical centrifugation, imaging and imaging characterizations, and immunoassays.
[0045] Example techniques for detecting material attributes include reduced and non-reduced peptide mapping (which may detect chemical modifications), chromatography (such as size exclusion chromatography (SEC), ion exchange chromatography (IEX) such as cation exchange chromatography (CEX), hydrophobic interaction chromatography (HIC), affinity chromatography such as Protein A-column chromatography, or reverse phase (RP) chromatography), capillary isoelectric focusing (cIEF), capillary zone electrophoresis (CZE), field flow fractionation (FFF), or ultracentrifugation (UC), HIAC (such as for detecting subvisible particle count), MFI (such as for detecting subvisible particle count and morphology), visible inspection (visible particles), SDS-PAGE (such as for detecting fragments, covalent aggregates), color analysis (Trp Ox), rCE-SDS and nrCE-SDS (such as for detecting fragments that are partial molecules), nanoparticle sizing methods, spectroscopy methods (such as FTIR, CD, intrinsic fluorescence, or ANS dye binding), an Ellman’s assay (free sulfhydryl’s), SEC-MALS, HILIC (glycan map), ELISA (such as for detecting HCP), and LAL assay for endotoxin, .
[0046] Pharmaceutical materials
[0047] As used herein “pharmaceutical material” and variations of this root term has its ordinary and customary meaning as would be understood by one of ordinary skill in the art in view of this disclosure. It refers to a therapeutic product comprising an active pharmaceutical ingredient (API). A pharmaceutical material may further comprise additional substances such as carriers or excipients. In some embodiments, a pharmaceutical material is subject to regulation and premarket approval by a government regulatory agency, such as the Food and Drug Administration (FDA) or the European Medicines Agency (EMA). In some embodiments, a pharmaceutical material is authorized for administration to a human subject
by such a government regulatory agency. Examples of pharmaceutical materials include biological therapies, small synthetic molecules, and nucleic acids such as small interfering RNA (siRNA) and DNA. In some embodiments, the pharmaceutical material is for medical use. In some embodiments, the pharmaceutical material is for medical use in a human subject.
[0048] As used herein “biological therapy” and variations of this root term has its ordinary and customary meaning as would be understood by one of ordinary skill in the art in view of this disclosure. It refers to a therapeutic composition comprising a biological macromolecule, for example a gene therapy, a therapeutic protein, a nucleic acid, a virus, or a cell or a portion thereof.
[0049] In methods described herein, the biological therapy may be selected from the group consisting of: an antibody, an antigen-binding antibody fragment, an antibody protein product, a Bi-specific T cell engager (BiTE®) molecule, a bispecific antibody, a trispecific antibody, an Fc fusion protein, a recombinant protein, a recombinant virus, a recombinant T cell, a synthetic peptide, deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and an active fragment of a recombinant protein.
[0050] An “antibody” has its customary and ordinary meaning as understood by one of ordinary skill in the art in view of this disclosure. It refers to an immunoglobulin of any isotype with specific binding to the target antigen, and includes, for instance, chimeric, humanized, and fully human antibodies. By way of example, the antibody may be a monoclonal antibody. By way of example, human antibodies can be of any isotype, including IgG (including IgGl, IgG2, IgG3 and IgG4 subtypes). A human IgG antibody generally comprise two full-length heavy chains and two full-length light chains. Antibodies may be derived solely from a single source, or may be “chimeric,” that is, different portions of the antibody may be derived from two or more different antibodies from the same or different species. Once an antibody is obtained from a source, it may undergo further engineering, for example to enhance stability and folding. Accordingly, a “human” antibody may be obtained from a source, and may undergo further engineering, for example in the Fc region. The engineered antibody may still be referred to as a type of human antibody. Similarly, variants of a human antibody, for example those that have undergone affinity maturation can be “human antibodies” unless stated otherwise. In some embodiments, an antibody comprises, consists essentially of, or consists of a human, humanized, or chimeric monoclonal antibody.
[0051] In various aspects, the biological therapy is an antibody protein product. As used herein, the term “antibody protein product” refers to any one of several antibody alternatives which in various instances is based on the architecture of an antibody but is not found in nature. In some aspects, the antibody protein product has a molecular-weight within the range of at least about 12-150 kDa. In certain aspects, the antibody protein product has a valency (n) range from monomeric (n = 1), to dimeric (n = 2), to trimeric (n = 3), to tetrameric (n = 4), if not higher order valency. Antibody protein products in some aspects are those based on the full antibody structure and/or those that mimic antibody fragments which retain full antigen-binding capacity, e.g., scFvs, Fabs and VHH/VH (discussed below). The smallest antigen binding antibody fragment that retains its complete antigen binding site is the Fv fragment, which consists entirely of variable (V) regions. A soluble, flexible amino acid peptide linker is used to connect the V regions to a scFv (single chain fragment variable) fragment for stabilization of the molecule, or the constant (C) domains are added to the V regions to generate a Fab fragment [fragment, antigen-binding]. Both scFv and Fab fragments can be easily produced in host cells, e.g., prokaryotic host cells. Other antibody protein products include disulfide-bond stabilized scFv (ds-scFv), single chain Fab (scFab), as well as di- and multimeric antibody formats like dia-, tria- and tetra-bodies, or minibodies (miniAbs) that comprise different formats consisting of scFvs linked to oligomerization domains. The smallest fragments are VHH/VH of camelid heavy chain Abs as well as single domain Abs (sdAb). The building block that is most frequently used to create novel antibody formats is the single-chain variable (V)-domain antibody fragment (scFv), which comprises V domains from the heavy and light chain (VH and VL domain) linked by a peptide linker of ~15 amino acid residues. A peptibody or peptide-Fc fusion is yet another antibody protein product. The structure of a peptibody consists of a biologically active peptide grafted onto an Fc domain. Peptibodies are well-described in the art. See, e.g., Shimamoto et al.., mAbs 4(5): 586-591 (2012).
[0052] Biological therapies suitable for the methods described herein can include polypeptides, including those that bind to one or more of the following: These include CD proteins, including CD3, CD4, CD8, CD19, CD20, CD22, CD30, CD34, and CD40; including human serum albumin (HSA) and insulin-like growth factor 1 receptor (IGF-1R), including those that interfere with receptor binding. HER receptor family proteins, including HER2, HER3, HER4, and the EGF receptor. Cell adhesion molecules, for example, LFA-I, Mol, pl50, 95, VLA-4, ICAM-I, VCAM, and alpha v/beta 3 integrin. Growth factors, such as
vascular endothelial growth factor (“VEGF”), growth hormone, thyroid stimulating hormone, follicle stimulating hormone, luteinizing hormone, growth hormone releasing factor, parathyroid hormone, Mullerian-inhibiting substance, human macrophage inflammatory protein (MIP-1 alpha), erythropoietin (EPO), nerve growth factor, such as NGF-beta, platelet- derived growth factor (PDGF), fibroblast growth factors, including, for instance, aFGF and bFGF, epidermal growth factor (EGF), transforming growth factors (TGF), including, among others, TGF- a and TGF-P, including TGF- 1, TGF-P2, TGF-P3, TGF- P4, or TGF- P5, insulin-like growth factors-I and -II (IGF-I and IGF-II), des(l-3)-IGF-I (brain IGF-I), and osteoinductive factors. Insulins and insulin-related proteins, including insulin, insulin A- chain, insulin B-chain, proinsulin, and insulin-like growth factor binding proteins. Coagulation and coagulation-related proteins, such as, among others, factor VIII, tissue factor, von Willebrands factor, protein C, alpha- 1 -antitrypsin, plasminogen activators, such as urokinase and tissue plasminogen activator (“t-PA”), bombazine, thrombin, and thrombopoietin; (vii) other blood and serum proteins, including but not limited to albumin, IgE, and blood group antigens. Colony stimulating factors and receptors thereof, including the following, among others, M-CSF, GM-CSF, and G-CSF, and receptors thereof, such as CSF-1 receptor (c-fms). Receptors and receptor-associated proteins, including, for example, flk2/flt3 receptor, obesity (OB) receptor, LDL receptor, growth hormone receptors, thrombopoietin receptors (“TPO-R,” “c-mpl”), glucagon receptors, interleukin receptors, interferon receptors, T-cell receptors, stem cell factor receptors, such as c-Kit, and other receptors. Receptor ligands, including, for example, OX40L, the ligand for the 0X40 receptor. Neurotrophic factors, including bone-derived neurotrophic factor (BDNF) and neurotrophin-3, -4, -5, or -6 (NT-3, NT-4, NT-5, or NT-6). Relaxin A-chain, relaxin B-chain, and prorelaxin; interferons and interferon receptors, including for example, interferon-a, -P, and -y, and their receptors. Interleukins and interleukin receptors, including IL-I to IL-33 and IL-I to IL-33 receptors, such as the IL-8 receptor, among others. Viral antigens, including an AIDS envelope viral antigen. Lipoproteins, calcitonin, glucagon, atrial natriuretic factor, lung surfactant, tumor necrosis factor-alpha and -beta, enkephalinase, RANTES (regulated on activation normally T-cell expressed and secreted), mouse gonadotropin-associated peptide, DNAse, inhibin, and activin. Integrin, protein A or D, rheumatoid factors, immunotoxins, bone morphogenetic protein (BMP), superoxide dismutase, surface membrane proteins, decay accelerating factor (DAF), HIV envelope, transport proteins, homing receptors, addressins, regulatory proteins, immunoadhesins, antibodies. Myostatins, TALL proteins, including
TALL-I, amyloid proteins, including but not limited to amyloid-beta proteins, thymic stromal lymphopoietins (“TSLP”), RANK ligand (“RANKL” or “OPGL”), c-kit, TNF receptors, including TNF Receptor Type 1, TRAIL-R2, angiopoi etins, and biologically active fragments or analogs or variants of any of the foregoing.
[0053] Examples of biological therapies suitable for the methods described herein include antibodies such as infliximab, bevacizumab, cetuximab, ranibizumab, palivizumab, abagovomab, abciximab, actoxumab, adalimumab, afelimomab, afutuzumab, alacizumab, alacizumab pegol, ald518, alemtuzumab, alirocumab, altumomab, amatuximab, anatumomab mafenatox, anrukinzumab, apolizumab, arcitumomab, aselizumab, altinumab, atlizumab, atorolimiumab, tocilizumab, bapineuzumab, basiliximab, bavituximab, bectumomab, belimumab, bemarituzumab, benralizumab, bertilimumab, besilesomab, bevacizumab, bezlotoxumab, biciromab, bivatuzumab, bivatuzumab mertansine, blinatumomab, blosozumab, brentuximab vedotin, briakinumab, brodalumab, canakinumab, cantuzumab mertansine, cantuzumab mertansine, caplacizumab, capromab pendetide, carlumab, catumaxomab, cc49, cedelizumab, certolizumab pegol, cetuximab, citatuzumab bogatox, cixutumumab, clazakizumab, clenoliximab, clivatuzumab tetraxetan, conatumumab, crenezumab, cr6261, dacetuzumab, daclizumab, dalotuzumab, daratumumab, demcizumab, denosumab, detumomab, dorlimomab aritox, drozitumab, duligotumab, dupilumab, ecromeximab, eculizumab, edobacomab, edrecolomab, efalizumab, efungumab, elotuzumab, elsilimomab, enavatuzumab, enlimomab pegol, enokizumab, enoticumab, ensituximab, epitumomab cituxetan, epratuzumab, erenumab, erlizumab, ertumaxomab, etaracizumab, etrolizumab, evolocumab, exbivirumab, fanolesomab, faralimomab, farletuzumab, fasinumab, fbtaO5, felvizumab, fezakinumab, ficlatuzumab, figitumumab, flanvotumab, fontolizumab, foralumab, foravirumab, fresolimumab, fulranumab, futuximab, galiximab, ganitumab, gantenerumab, gavilimomab, gemtuzumab ozogamicin, gevokizumab, girentuximab, glembatumumab vedotin, golimumab, gomiliximab, gs6624, ibalizumab, ibritumomab tiuxetan, icrucumab, igovomab, imciromab, imgatuzumab, inclacumab, indatuximab ravtansine, infliximab, intetumumab, inolimomab, inotuzumab ozogamicin, ipilimumab, iratumumab, itolizumab, ixekizumab, keliximab, labetuzumab, lebrikizumab, lemalesomab, lerdelimumab, lexatumumab, libivirumab, ligelizumab, lintuzumab, lirilumab, lorvotuzumab mertansine, lucatumumab, lumiliximab, mapatumumab, maslimomab, mavrilimumab, matuzumab, mepolizumab, metelimumab, milatuzumab, minretumomab, mitumomab, mogamulizumab, morolimumab, motavizumab, moxetumomab pasudotox,
muromonab-cd3, nacolomab tafenatox, namilumab, naptumomab estafenatox, namatumab, natalizumab, nebacumab, necitumumab, nerelimomab, nesvacumab, nimotuzumab, nivolumab, nofetumomab merpentan, ocaratuzumab, ocrelizumab, odulimomab, ofatumumab, olaratumab, olokizumab, omalizumab, onartuzumab, oportuzumab monatox, oregovomab, orticumab, otelixizumab, oxelumab, ozanezumab, ozoralizumab, pagibaximab, palivizumab, panitumumab, panobacumab, parsatuzumab, pascolizumab, pateclizumab, patritumab, pemtumomab, perakizumab, pertuzumab, pexelizumab, pidilizumab, pintumomab, placulumab, ponezumab, priliximab, pritumumab, PRO 140, quilizumab, racotumomab, radretumab, rafivirumab, ramucirumab, ranibizumab, raxibacumab, regavirumab, reslizumab, rilotumumab, rituximab, robatumumab, roledumab, romosozumab, rontalizumab, rovelizumab, ruplizumab, samalizumab, sarilumab, satumomab pendetide, secukinumab, sevirumab, sibrotuzumab, sifalimumab, siltuximab, simtuzumab, siplizumab, sirukumab, solanezumab, solitomab, sonepcizumab, sontuzumab, stamulumab, sulesomab, suvizumab, tabalumab, tacatuzumab tetraxetan, tadocizumab, talizumab, tanezumab, taplitumomab paptox, tarlatamab, tefibazumab, telimomab aritox, tenatumomab, tefibazumab, teneliximab, teplizumab, teprotumumab, tezepelumab, TGN1412, tremelimumab, ticilimumab, tildrakizumab, tigatuzumab, TNX-650, tocilizumab, toralizumab, tositumomab, tralokinumab, trastuzumab, TRBS07, tregalizumab, tucotuzumab celmoleukin, tuvirumab, ublituximab, urelumab, urtoxazumab, ustekinumab, vapaliximab, vatelizumab, vedolizumab, veltuzumab, vepalimomab, vesencumab, visilizumab, volociximab, vorsetuzumab mafodotin, votumumab, zalutumumab, zanolimumab, zatuximab, ziralimumab, or zolimomab aritox.
[0054] In some embodiments, the biological therapy is a BiTE® molecule. BiTE® molecules are engineered bispecific antigen binding constructs which direct the cytotoxic activity of T cells against cancer cells. They can be the fusion of two single-chain variable fragments (scFvs) of different antibodies, or amino acid sequences from four different genes, on a single peptide chain of about 55 kilodaltons. One of the scFvs binds to T cells via the CD3 receptor, and the other to a tumor cell via a tumor specific molecule. Blinatumomab (BLINCYTO®) is an example of a BiTE® molecule, specific for CD 19. BiTE® molecules that are modified, such as those modified to extend their half-lives, can also be used in the disclosed methods. In various aspects, the polypeptide is an antigen binding protein, e.g., a BiTE® molecule. In some embodiments, an antibody protein product comprises a BiTE® molecule.
[0055] In some embodiments, the biological therapy is in a formulation. The formulation may be a pharmaceutically acceptable formulation. The formulation may comprise the biological therapy together with a pharmaceutically acceptable diluent, carrier, solubilizer, emulsifier, preservative, surfactant, and/or adjuvant.
[0056] Acceptable formulation materials for biological therapies as described herein can preferably be nontoxic to recipients at the dosages and concentrations employed. In certain embodiments, the pharmaceutical composition may contain formulation materials for modifying, maintaining or preserving, for example, the pH, osmolality, viscosity, clarity, color, isotonicity, odor, sterility, stability, rate of dissolution or release, adsorption or penetration of the composition. In such embodiments, suitable formulation materials include, but are not limited to, amino acids (such as glycine, glutamine, asparagine, arginine or lysine); antimicrobials; antioxidants (such as ascorbic acid, sodium sulfite or sodium hydrogen-sulfite); buffers (such as borate, bicarbonate, Tris-HCl, citrates, phosphates or other organic acids); bulking agents (such as mannitol or glycine); chelating agents (such as ethylenediamine tetraacetic acid (EDTA)); complexing agents (such as caffeine, polyvinylpyrrolidone, beta-cyclodextrin or hydroxypropyl-beta-cyclodextrin); fillers; monosaccharides; disaccharides; and other carbohydrates (such as glucose, sucrose, mannose or dextrins); proteins (such as serum albumin, gelatin or immunoglobulins); coloring, flavoring and diluting agents; emulsifying agents; hydrophilic polymers (such as polyvinylpyrrolidone); low molecular weight polypeptides; salt-forming counterions (such as sodium); preservatives (such as benzalkonium chloride, benzoic acid, salicylic acid, thimerosal, phenethyl alcohol, methylparaben, propylparaben, chlorhexidine, sorbic acid or hydrogen peroxide); solvents (such as glycerin, propylene glycol or polyethylene glycol); sugar alcohols (such as mannitol or sorbitol); suspending agents; surfactants or wetting agents (such as pluronics, PEG, sorbitan esters, polysorbates such as polysorbate 20 or polysorbate 80, polysorbate, triton, tromethamine, lecithin, cholesterol, tyloxapal); stability enhancing agents (such as sucrose or sorbitol); tonicity enhancing agents (such as alkali metal halides, preferably sodium or potassium chloride, mannitol sorbitol); delivery vehicles; diluents; excipients and/or pharmaceutical adjuvants. See, e.g., REMINGTON'S PHARMACEUTICAL SCIENCES, 18" Edition, (A. R. Genrmo, ed.), 1990, Mack Publishing Company.
[0057] A suitable vehicle or carrier for the formulation may be water for injection, physiological saline solution or artificial cerebrospinal fluid, possibly supplemented with
other materials common in compositions for parenteral administration. Neutral buffered saline or saline mixed with serum albumin are further exemplary vehicles. In specific embodiments, pharmaceutical compositions comprise Tris buffer of about pH 7.0-8.5, or acetate buffer of about pH 4.0-5.5, and may further include sorbitol or a suitable substitute therefor.
[0058] Methods of manufacturing a pharmaceutical material
[0059] The methods described herein may be used to assess whether material attributes impact efficacy of a pharmaceutical material, and/or impact a safety profile of a pharmaceutical material. The methods can comprise determining an estimated actual level of material attribute exposure at the time that a subject receives an administration of a pharmaceutical material. The material attribute exposure can be a dose of a material attribute being introduced to a subject. The methods and systems disclosed herein can determine an estimated or actual level of material attribute exposure by applying pharmacokinetic profiles to the pharmaceutical material (e.g., in cases of dose escalation a subject may be exposed to something like 10 times of the clinically relevant dose, and as such some level of a prior dose can feasibly be in the bloodstream by the next dose), stability modeling for material attributes for lots without any associated direct measurement or for material attributes that may have gaps in measurement, and stochastic modeling of emergent phenomena such as aggregate dissociation modeling or in-vivo chemical modification modeling which are factors that cannot be directly observed or measured. The methods and systems disclosed herein can identify suitable material attribute levels or ranges for manufacturing, lot release, and/or time of administration.
[0060] In some embodiments, a method of manufacturing a pharmaceutical material is described. The method can comprise detecting a level of a material attribute of the pharmaceutical material in a formulation at one or more timepoints under storage conditions (optionally, at two or more timepoints under storage conditions). Optionally, for example instances in which there is no change in material attributes upon storage (for example for some pharmaceutical materials that are stored frozen), it may be sufficient to detect the level of material attribute at one timepoint. It is further contemplated, that for levels of material attributes that may change during storage, detecting the level of the material attribute at two more timepoints under storage conditions can be used to calculate a rate of change of the material attribute as described herein. By way of example, the timepoints may comprise the time of manufacture. By way of example, the timepoints may comprise the time of
manufacture and at least one other timepoint. The method can comprise determining a rate of change of the material attribute under the storage conditions (for some material attribute of some pharmaceutical material under some storage conditions, for example some pharmaceutical material that are stored frozen, the rate of change may be calculated as zero). The method can comprise obtaining data on in vivo safety and/or efficacy of the pharmaceutical material for subjects that have received an administration of the pharmaceutical material. The method can comprise estimating a level of material attribute exposure received by the subjects at the time of said administration, based on (i) the rate of change of the material attribute during storage and (ii) a duration that the pharmaceutical material in the formulation was under the storage conditions prior to said administration. The estimating may also take into account (iii) the level of the material attribute at the time of manufacture and/or at lot release, and/or (iv) the dose of the pharmaceutical material administered to the subject. The method can comprise determining a presence or absence of a correlation between the estimated level of material attribute exposure received by the subjects and the safety and/or efficacy data for the pharmaceutical material. If the correlation is absent, the method can comprise manufacturing production lots of the pharmaceutical material comprising the material attribute at or below a specified permissible level of the material attribute based on the estimated level of molecule attribute exposure. For example, the specified permissible level may be a level of the material attribute that is calculated to yield a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure in subjects. For example, the specified permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure in subjects. Accordingly, the permissible level is expected to yield levels of attribute exposure that are shown to be safe and effective in subjects throughout the shelf life of the pharmaceutical material. Production lots with levels of the material attribute that exceed the specified permissible level may be rejected.
[0061] If the correlation is present, the method can further comprise setting a specification for levels of the material attribute at the time of manufacture to not exceed a maximum permissible level of the material attribute of the pharmaceutical material. The maximum permissible level of the material attribute can be based on the highest estimated levels of material attribute exposure in subjects that are not associated with adverse events and/or inhibition of efficacy of the pharmaceutical material. The method of manufacturing can
further comprise rejecting production lots of the pharmaceutical material comprising levels of the material attribute exceeding the maximum permissible level (and thus outside of the specification). Production lots with levels of the material attribute that do not exceed the specified maximum permissible level may be accepted. For example, the specified maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects. The calculation may take into account the rate of change of the material attribute in formulation during storage, and the dose of material attribute to be administered to a subject. For example, the specified maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects.
[0062] As used herein a “permissible level” of material attribute refers to a level of material attribute of a pharmaceutical material in formulation that is within specification at the time of manufacture. As described herein, the permissible level may be calculated to yield a specified level of material attribute exposure (to a subject who would be administered the pharmaceutical material in formulation) at end-of-shelf or expiration that is less than or equal to an estimated level of material attribute exposure that was not associated with adverse events and/or loss of efficacy (this may be referred to as a permissible level “based on” the estimated level of attribute exposure). Thus, permissible levels of material attributes can provide confidence that the level of material attribute exposure from the pharmaceutical material in formulation is safe and effective when administered to a subject, even if it is administered very close to its end-of-shelf or expiration date. A “maximum permissible level” refers to a scenario when there is a correlation between the estimated level of material attribute exposure and safety and/or efficacy data. The “maximum permissible level” refers to the highest level of material attribute in formulation that yields a level of material attribute exposure (to a subject who would be administered the pharmaceutical material in formulation) at end-of-shelf or expiration that is less than or equal to the highest level of material attribute exposure not associated with adverse events. In some embodiments, the maximum permissible level can be a maximum level of attribute that was exposed in patients during the clinical trial. The permissible level or maximum permissible level may be calculated based on the estimated level of attribute exposure not associated with adverse
events, the duration until end-of-shelf or expiration, and rate of change of the level of material attribute, and the dose of the pharmaceutical material. That is, using the level of material attribute exposure that was not associated with adverse events and/or inhibition of efficacy, the rate of change in the material attribute in formulation, and this duration, the permissible level (if applicable, the maximum permissible level) may be determined. Thus, if a pharmaceutical material is manufactured with a material attribute at or below a permissible level (if applicable a maximum permissible level), it can be expected that at the time of administration to a subject, the estimated level of material attribute exposure to the subject will be at or below a level that is not associated with adverse events and/or loss of efficacy. [0063] As used herein, a permissible level or maximum permissible level “based on” an estimated level of a material attribute exposure refers to the permissible level (or maximum permissible level) calculated to yield no more than the estimated level of material attribute at end-of-shelf or expiration. If more than one dose is suitable for a pharmaceutical material in formulation, the highest suitable dose may be used to calculate the permissible level or maximum permissible level “based on” the estimated level of a material attribute (as lower doses will have even lower levels of the material attribute exposure). By way of example, the estimated level of attribute exposure may be selected to fall within a confidence interval for a distribution or a spread of estimated material attribute exposure levels that were determined for a group of subjects. An estimated level of attribute exposure received by subjects does not imply that every subject received the same numerical level of attribute exposure. Rather, the estimated level of attribute exposure received by those subjects may comprise a distribution of estimated attribute levels received by individual subjects. Thus, the maximum permissible level may be selected using a confidence interval so that there is at least an 85%, 90%, 95%, 97%, or 99% probability that the level of attribute exposure at end of shelf is less than or equal to the highest level of attribute exposure that was not associated with a loss of safety or efficacy. The actual value of the permissible or maximum permissible level “based on” the level of material attribute exposure may undergo rounding. The rounding may provide administrative or mathematical convenience, such as rounding to one, two, or three significant figures in a suitable unit for measuring the material attribute. For added caution, the rounding may be a rounding down. For example, the permissible level or maximum permissible level of attribute “based on” the estimated level of attribute exposure may be calculated to yield a level of attribute exposure at end-of-shelf that is no more than 99%, 97%, 95%, 90%, 85%, or 80% of the reference level of attribute exposure (that was not
associated with a loss of safety and/or efficacy). This rounding down can provide additional assurance that the level of attribute exposure from the end-of-shelf pharmaceutical material will still be safe and/or effective.
[0064] In some embodiments, production lots are manufactured with levels of more than one material attribute, each at or below its permissible (or maximum permissible) level. For example, the production lots may be manufactured with at least one, two, three, four, five, six, seven, eight, nine, ten, twenty, or more material attributes, including ranges between any two of the listed values, for example one - ten, one - five, two - ten, two - five, three - ten, three - five, or five - ten material attributes, each of which is at or below a specified permissible (or maximum permissible) level.
[0065] Manufacturing techniques
[0066] Methods for manufacturing pharmaceutical material as described herein (e.g., therapeutic proteins) may utilize recombinant DNA technology. Recombinant DNA methods for producing therapeutic proteins such as antibodies or antibody protein products are well- known. DNA may encode the therapeutic protein. For example, DNA may encode the antibodies, for example, DNA encoding a VH domain, a VL domain, a single chain variable fragment (scFv), or fragments and combinations thereof (target polynucleotides), can be inserted into a suitable expression vector, which can then be transfected into a suitable host cell, such as Escherichia coli cells, COS cells, Chinese Hamster Ovary (CHO) cells, or myeloma cells that do not otherwise produce an antibody, to obtain the desired antibodies. [0067] Suitable expression vectors are known in the art, containing, for example a polynucleotide that encodes the target polypeptide linked to a promoter. Such vectors can include the nucleotide sequence encoding the constant region of the antibody molecule, and the variable domain of the antibody can be cloned into such a vector for expression of the heavy chain, the entire light chain, or both the entire heavy and light chains (or fragments thereof). The expression vector can be transferred to a host cell by conventional techniques, and the transfected cells can be cultured to produce the antibodies.
[0068] Any cell line that can express, or is engineered to express, proteins such as functional antibody or antibody fragments, can be used. For example, suitable mammalian cell lines include immortalized cell lines available from the American Type Culture Collection (Manassas, VA), including Chine Hamster Ovary (CHO)) cells, HeLa cells, baby hamster kidney (BHK) cells, monkey kidney cells (COS), human hepatocellular carcinoma cells (e.g., Hep G2), and human epithelial kidney 293 cells. Furthermore, cell lines or host systems can
be chosen to ensure correct modification and processing of antibodies. Eukaryotic host cells that possess the cellular machinery for proper processing of the primary transcript, glycosylation, and phosphorylation of the gene product can be used. These include CHO, VERY, BHK, Hela, COS, MDCK, 293, 3T3, W138, BT483, Hs578T, HTB2, BT20 and T47D, NSO (a murine myeloma cell line that does not endogenously produce any functional immunoglobulin chains), SP20, CRL7030 and HsS78Bst cells. Human cell lines developed by immortalizing human lymphocytes can also be used. The human cell line PER.C6® (Janssen; Titusville, NJ) can be used to recombinantly produce monoclonal antibodies. Examples of non-mammalian cells that can also be used include insect cells (e.g., Sf21/Sf9, Trichoplusia ni Bti-Tn5bl-4), or yeast cells (e.g., Saccharomyces (such as S. cerevisiae, Pichia, etc.), plant cells, or chicken cells.
[0069] Proteins such as antibodies can be stably expressed in a cell line using conventional methods. Stable expression can be used for long-term, high-yield production of recombinant proteins. For stable expression, host cells can be transformed with an appropriately engineered vector that includes expression control elements (e.g., promoter, enhancer, transcription terminators, polyadenylation sites, etc.), and a selectable marker gene. Methods for producing stable cell lines with a high yield are known in the art and reagents are available commercially. Transient expression can also be accomplished using conventional methods.
[0070] A cell line expressing a protein such as an antibody can be maintained in cell culture medium and under culture conditions that result in the expression and production of the antibodies. Cell culture media can be based on commercially available media formulations, including, for example, DMEM or Ham's F12. In addition, the cell culture media can be modified to support increases in both cell growth and biologic protein expression. Of course, cell culture medium can be optimized for a specific cell culture, including cell culture growth medium which is formulated to promote cellular growth or cell culture production medium which is formulated to promote recombinant protein production.
[0071] Many cell culture media and cell culture nutrients and supplements are known. For example, suitable basal media include Dulbecco's Modified Eagle's Medium (DMEM), DME/F12, Minimal Essential Medium (MEM), Basal Medium Eagle (BME), RPMI 1640, F- 10, F-12, a-Minimal Essential Medium (a-MEM), Glasgow's Minimal Essential Medium (G- MEM), PF CHO, and Iscove's Modified Dulbecco's Medium. Other examples of basal media that can be used include BME Basal Medium, Dulbecco's Modified Eagle Medium.
[0072] The basal medium may be serum-free, meaning that the medium contains no serum (e.g., fetal bovine serum (FBS)) or animal protein-free media or chemically-defined media. The basal medium can be modified in order to remove certain non-nutritional components found in basal media, such as various inorganic and organic buffers, surfactant(s), and sodium chloride. The cell culture medium can contain a basal cell medium (modified or not), and at least one of the following: iron source, recombinant growth factor; buffer; surfactant; osmolarity regulator; energy source; and non-animal hydrolysates. In addition, the modified basal cell medium can optionally contain amino acids, vitamins, or a combination of both amino acids and vitamins. A modified basal medium can further contain glutamine, e.g., L- glutamine, and/or methotrexate.
[0073] Once a therapeutic protein (e.g., an antibody or antibody protein product) has been produced, it can be purified by conventional methods, for example, by chromatography (e.g., ion exchange, affinity, particularly by affinity for the specific antigens, Protein A, Protein G, or sizing column chromatography), centrifugation, differential solubility, or by any other standard technique for the purification of proteins. Further, the protein can be fused to heterologous polypeptide sequences (“tags”) to facilitate purification.
[0074] The purified protein can typically be formulated with an excipient to produce a sterile solution that can be injected or infused. For example, the purified protein may be formulated in a formulation as described herein. Formulation may be followed by filling, packaging, storage, delivery, and final preparation immediately prior to administration to the subject. [0075] Methods of developing a manufacturing process for pharmaceutical material [0076] The methods described herein may also be used to develop methods for manufacturing a pharmaceutical material. In some embodiments, a method of developing a manufacturing process for pharmaceutical material is described. The method can comprise detecting a level of a material attribute of the pharmaceutical material in a formulation at one or more timepoints under storage conditions (optionally at two or more timepoints under storage conditions). For example, the timepoints may include the time of manufacture. For example, the timepoints may include the time of manufacture, and at least one other timepoint. The method can comprise determining a rate of change of the material attribute under the storage conditions. The method can comprise obtaining data on safety and/or efficacy of the pharmaceutical material in subjects that have received an administration of the pharmaceutical material. The method can comprise estimating a level of material attribute exposure received by the subjects at the time of administration, based on (i) the rate of
change of the material attribute during storage of the pharmaceutical material in the formulation and (ii) a duration that the pharmaceutical material in the formulation was under the storage conditions prior to said administration. The estimating may also take into account
(iii) the level of the material attribute at the time of manufacture and/or at lot release, and/or
(iv) the dose of the pharmaceutical material administered to the subject. The method can comprise determining a presence or absence of a correlation between the estimated level of material attribute exposure and the safety and/or efficacy data of the pharmaceutical material. [0077] If the correlation is absent, the method can comprise establishing the manufacturing process to produce levels of the material attribute at or below a specified permissible level based on the estimated level of material attribute exposure. For example, the specified permissible levels may be a level of the material attribute that is calculated to yield a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure in subjects. For example, the specified permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end- of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure in subjects. Production lots with levels of the material attribute that exceed the specified permissible level may be rejected.
[0078] If the correlation is present, the method can comprise establishing the manufacturing process to produce levels of the material attribute at or below a specified a maximum permissible level of the material attribute based on the highest level of the material attribute that is not associated with adverse events and/or inhibition of efficacy of the pharmaceutical material. For example, the specified maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects. For example, the specified maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects. The permissible level or maximum permissible level may be part of a specification at time of manufacture or lot release.
[0079] The inventors have developed machine learning techniques for determining an impact of attributes. The systems and methods disclosed herein can comprise a machine learning
model developed to a correlation status based on the modified material attribute data and the modified clinical data using a computational model for different pharmaceutical materials. The prediction can be used to support regulatory filings. The systems and methods disclosed herein can provide insights into the relationship between clinical data and material attribute data for different pharmaceutical materials. Thus, the systems and methods disclosed herein can accelerate processes such as molecule candidate selection, product development, and initiation of clinical trials, which are useful for launching the pharmaceutical material to market for earlier patient access.
[0080] Methods of assessing clinical impact of a material attribute of a pharmaceutical material
[0081] In some embodiments, a method of assessing a clinical impact of material attributes are described. Such method may further be used in a method of manufacturing a pharmaceutical material, and in a method of developing a manufacturing process for a pharmaceutical material as described herein. The method can comprise detecting a level of a material attribute of the pharmaceutical material in a formulation at one or more timepoints under storage conditions. The level of attribute exposure can be detected at two or more timepoints under the storage conditions. By way of example, the timepoints for detecting the level of material attribute may include the time of manufacture or at least one other timepoint. By way of example, the two or more timepoints for detecting the level of material attribute may include the time of manufacture and at least one other timepoint. The method can comprise determining a rate of change of the material attribute under the storage conditions. The method can comprise obtaining data on safety and/or efficacy of the pharmaceutical material in subjects that have received an administration of the pharmaceutical material. The method can comprise estimating a level of material attribute exposure received by the subjects at the time of said administration, based on (i) the rate of change of the material attribute during storage of the pharmaceutical material in the formulation and (ii) a duration that the pharmaceutical material in the formulation was under the storage conditions prior to said administration. The estimating may also take into account (iii) the level of the material attribute at the time of manufacture and/or at lot release, and/or (iv) the dose of the pharmaceutical material administered to the subject. The method can comprise determining a presence or absence of a correlation between the estimated material attribute exposure and the safety and/or efficacy of the pharmaceutical material. If the correlation is absent, it can be determined that the material attribute impacts neither clinical safety nor efficacy of the
pharmaceutical material. If the correlation is present, it can be determined that the material attribute impacts clinical safety and/or efficacy of the pharmaceutical material.
[0082] If the correlation is absent, the method may further comprise setting a specification for permissible levels of the material attribute of the pharmaceutical material, in which the permissible levels of the material attribute are based on the highest estimated levels of the material attribute received by the subjects. For example, the specification for permissible levels may be a level of the material attribute at the time of manufacture that is calculated to yield a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure in subjects. For example, the specification of permissible levels may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure in subjects. The specification may be used for manufacturing and/or release testing of the pharmaceutical material. If a lot or product of the pharmaceutical material comprises the material attribute at a level that exceeds the permissible level, the lot or product may be considered out of specification and rejected. The pharmaceutical material may be manufactured according to the specification.
[0083] If the correlation is present, the method may further comprise setting a specification for maximum permissible levels of the material attribute of the pharmaceutical material, in which the maximum permissible levels of the material attribute are based on levels of the highest estimated material attribute exposure that were not associated with adverse events and/or inhibition of efficacy of the pharmaceutical material. The specification may be used for manufacturing and/or release testing of the pharmaceutical material. For example, the specification for the maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects. For example, the specified maximum permissible level may be a level of the material attribute that is calculated to produce a level of attribute exposure at end-of-shelf that is less than or equal to 90 - 100% of the highest estimated level of material attribute exposure that is not associated with adverse events and/or an inhibition of efficacy in subjects. If a production lot or product of the pharmaceutical material comprises the material attribute at a level that exceeds the maximum permissible level, the lot or product may be considered out of specification and rejected. The pharmaceutical material may be manufactured according to the specification.
[0084] Further Aspects of Methods
[0085] Any of the methods described herein may comprise one or more further aspects. [0086] In some embodiments, for any method described herein, the estimating is further based on (iii) a dose of the pharmaceutical material in said administration, and (iv) an amount of the material attribute measured at time of manufacture and/or lot release. Regarding (iii), it is noted that a greater dose of pharmaceutical material will administer a larger quantity of a material attribute than a smaller dose having the same percent content of the material attribute. It is also contemplated that in some instances, the level of the material attribute at time of manufacture and/or the dose may not vary appreciably, and thus the estimate can be performed from (i) and (ii) alone. For example, the manufacturing process may be well- established and tightly controlled and shown to yield consistent levels of a material attribute at the time of manufacture. For example, the pharmaceutical material may be administered at a single dose without weighing the calculation by dose.
[0087] In some embodiments, for any method described herein, the level of material attribute of the pharmaceutical material in the formulation is measured at one or more timepoints at or after time of manufacture, for example, at least one, two, three, four, five, six, seven, eight, nine, or ten timepoints, including ranges between any two of the listed values, for example one-two timepoints, one-five timepoints, or one-ten timepoints. Without being limited by theory, it is contemplated that for some pharmaceutical material in formulation, for which the levels of material attribute do not change during storage, as may be the case for some pharmaceutical material in formulation that are stored frozen, it may be sufficient to measure the level of material attribute at a single timepoint. In some embodiments, for any method described herein, the level of material attribute of the pharmaceutical material in the formulation is measured at two or more timepoints, for example at least two, three, four, five, six, seven, eight, nine, or ten timepoints, including ranges between any two of the listed values, for example, two-five timepoints, two-ten timepoints, three-five timepoints, three-ten timepoints, or five-ten timepoints. The levels of the material attribute at the two or more timepoints can be used to calculate a rate of change in the levels of the material attribute under storage conditions. The timepoints may be at or after the time of manufacture. In some embodiments, for any method described herein, the two or more timepoints comprise a time of manufacture. In some embodiments, for any method described herein, the two or more timepoints comprise a time of manufacture, and at least one subsequent timepoint. In
some embodiments, for any method described herein, the two or more timepoints comprise a time of manufacture, and at least two subsequent timepoints.
[0088] Additional mathematical operations and machine learning techniques may be applied to further refine the determination of a correlation between the level of material attribute exposure, and the data on safety and efficacy. In some embodiments, for any method described herein, the correlation comprises a weighted correlation between attribute exposure and adverse event occurrences. The weighting may use the size of each groups of subjects. Without being limited by theory, if subjects experience adverse events in a multi-dose regimen of a pharmaceutical material, those subjects may stop using the pharmaceutical material. In the aggregate, there may be fewer data for subjects experiencing adverse events than for those subjects who did not experience adverse events, which may impact the determination of a correlation (or lack thereof). Accordingly, data may optionally be weighted to account for subjects who discontinued a multi-dose regimen of a pharmaceutical material before completing the regiment, for example subjects who experienced adverse events.
[0089] For different subjects, the pharmaceutical material in the formulation may be under storage conditions for different durations for different subjects. As such, there may not be a single duration for which the pharmaceutical material in the formulation was under storage conditions. Rather, different durations may be applied to different administrations of the pharmaceutical material, as appropriate. Estimating the level of the material attribute exposure may account for these different durations by selecting an appropriate time of storage in each instance. Thus, in the method of some embodiments, the pharmaceutical material was under the storage conditions for different durations in the administrations to different subjects.
[0090] For methods as described herein, the administration of the pharmaceutical material may comprise two or more administration events, for example at least two, three, four, five, or ten administration events, including ranges between any two of the listed values, such as two-three, two-five, two-ten, three-five, three-ten, or five-ten administration events. According, in some embodiments, the estimated level of attribute exposure received by the subjects at the time of administration is a maximum or an average of the two or more administration events.
[0091] For methods as described herein, the administration of the pharmaceutical material may comprise a continuous infusion. The continuous infusion may be, for example, over a
period of hours or days. In some embodiments, the administration comprises a continuous infusion, and estimating the level of attribute exposure comprises calculating an estimated level of material attribute exposure during two or more intervals of the continuous infusion, such 12-hour or 24-hour intervals.
[0092] Methods described herein may utilize clinical study data. These include, but are not limited to, the subject number, clinical outcome information (date, time and extent (severity or grade)), treatment lot number, treatment (date, time and duration). Data may be grouped according to clinical outcome (e.g., clinical endpoints and/or adverse events). Adverse event may be indicative of safety data, and clinical endpoints may be indicative of efficacy data. In clinical outcome-based grouping, subjects can be grouped according to the occurrence of the clinical outcome with the severity or extent of interest. By way of example, to analyze the impact of a material attribute on the severity of a safety adverse event, the subjects can be grouped according to the occurrence of the adverse event (AE) with the severity grade of interest (e.g. The severe fever-positive group would include the patients who had a fever with grade 3 or higher, whereas the severe fever-negative group would include the subjects who did not have a fever with grade 3 or higher). Then, the estimated level of attribute exposure for each subject can be determined as described.
[0093] By way of example, safety data for the methods described herein may comprise adverse event data. The adverse events may comprise treatment-related adverse events. For example, adverse events related to treatment with a pharmaceutical material may comprise immune adverse events, such as an immune response against the pharmaceutical material (e.g., anti-drug antibody (ADA)). Additional examples of adverse events include anemia, cytokine release syndrome (CRS), fever, infusion related reaction (IRR), lymphocytopenia, or neurological events. The adverse events can comprise one or more challenge conditions. Such challenge conditions can comprise controlled parameters including, but not limited to HCP, viral proteins, and excipients. The adverse events can comprise adverse events outlined by medically referenced adverse event guidelines such as MeDRA. The adverse events can comprise emergent adverse event phenomena (or emergent feature) such as headache-fatigue. The adverse events can be local (e.g., ocular) or systemic (throughout the body). The adverse events can be drug-related or disease-related. In this situation, the adverse events can be explored by monitoring manufacturing conditions including, but not limited to filter clogging, turbidity, and color variance. In some embodiments, for any method described herein, the
safety data comprise adverse event data. By way of example, the safety data may comprise an adverse event time course.
[0094] By way of example, efficacy data for the methods described herein may comprise clinical evidence of efficacy, for example, treatment, prevention, amelioration, or delayed onset of a disease or disorder. Clinical endpoint data may be indicative of efficacy. Accordingly, in some embodiments, efficacy data comprise clinical endpoint data.
[0095] For any of the methods herein, presence or absence of a correlation between the estimated level of material attribute exposure and the safety and/or efficacy data for the pharmaceutical material may be calculated using statistical methods. For example, a linear regression correlation between the rate of clinical event from safety and/or efficacy data and attribute exposure level may be calculated, and p-values may be determined. For example, the p values may be calculated from t-tests or F-tests. For example, Log-rank or Mantel-Cox test could be performed to calculate p-values for testing the earlier clinical safety event onsets for clinical subject groups with higher levels of the attribute exposures.
[0096] In some implementations of any of the methods described herein, a Bayesian estimation approach is used. The Bayesian estimation approach may be used for determining a presence or absence of a correlation between the estimated level of material attribute exposure and the safety and/or efficacy data for the pharmaceutical material. Due to the stepwise learning nature of a Bayesian approach, all evidences may be fully used a priori in an unbiased way. In particular, when applying a Bayesian estimation approach for determining a presence or absence of a correlation between the estimated level of material attribute exposure and the safety and/or efficacy data for the pharmaceutical material, the method may modify the probability density function of the clinical impact of each attribute to take a certain value whenever a new piece of evidence comes in, until all evidences are used. The probability of each attribute to be associated with each adverse event is derived at the end of the analysis. A computer program has been developed to harness high computation power used for the Bayesian approach. The program provides an interactive platform to visualize the outcome of Bayesian estimation approach with user-provided spreadsheets. The program also provides parameter tuning modules to define initial priors, remove outliers, and select multiple attributes. A few sets of real clinical data from previous CIA performed for mAb D as well as a few modeled data based on the mAb D CIA data were tested on the system. The Bayesian estimation method generated expected results, demonstrating the validity of the system.
[0097] FIG. 1 is a block diagram of an exemplary system 100 for determining an impact of attributes, in accordance with some embodiments of the technology described herein.
[0098] System 100 includes a computing system 110 coupled to a database 120. Computing system 110 can be configured to have software 130 execute thereon to perform various functions in connection with determining an impact of attributes. Computing system 110 can comprise a single computing device or include multiple co-located and/or distributed computing devices communicatively coupled by one or more networks. The computing system 110 can comprise one or multiple computing devices of any suitable type. For example, the computing system 110 may be a portable computing device (e.g., laptop, a smartphone) or a fixed computing device (e.g., a desktop computer, a server). When computing system 110 includes multiple computing devices, the device(s) may be physically co-located (e.g., in a single room) or distributed across multiple physical locations. In some embodiments, the computing system 110 may be part of a cloud computing infrastructure. [0099] In some embodiments, the computing system 110 may be operated by one or more user(s) 150 such as one or more researchers, health professionals, and/or other individual(s). For example, the user(s) 150 may provide material attribute data and/or clinical data associated with one or more pharmaceutical materials as input to the computing system 110 (e.g., by uploading one or more files), and/or may provide user input specifying processing or other methods to be performed on the material attribute data and/or clinical data associated with one or more pharmaceutical materials.
[00100] In the example embodiment shown in FIG. 1, computing system 110 includes a processing unit 112, a network interface 114, a display 116, a user input device 118, and a software 130. Processing unit 112 includes one or more processors, each of which may be a programmable microprocessor that executes software instructions stored in memory to execute some or all of the functions of computing system 110 as described herein. Alternatively, one, some or all of the processors in processing unit 112 may be other types of processors (e.g., application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.), and the functionality of computing system 110 as described herein may instead be implemented, in part or in whole, in hardware. Memory may include one or more physical memory devices or units containing volatile and/or non-volatile memory. Any suitable memory type or types may be used, such as read-only memory (ROM), solid-state drives (SSDs), hard disk drives (HDDs), and so on.
[00101] Network interface 114 may include any suitable hardware (e.g., front-end transmitter and receiver hardware), firmware, and/or software configured to communicate with external devices and/or systems (e.g., a client device, or one or more servers maintaining database 120) via one or more networks using one or more communication protocols. For example, network interface 114 may be or include an Ethernet interface, and/or include a wireless local area network (LAN) interface, etc.
[00102] Display 116 may use any suitable display technology (e.g., LED, OLED, LCD, etc.) to present information to a user 150, and user input device 118 may be a keyboard or other suitable input device. In some embodiments, display 116 and user input device 118 are integrated within a single device (e.g., a touchscreen display). Generally, display 116 and user input device 118 may combine to enable a user 150 to interact with user interfaces (e.g., graphical user interfaces (GUIs)) provided by computing system 110, such as those discussed in further detail below. In some embodiments, however, computing system 110 does not include display 116 and/or user input device 118, or one or both of display 116 and user input device 118 are included in another computer or system that is communicatively coupled to computing system 110 (e.g., a client device not shown in FIG. 1).
[00103] As shown in FIG. 1, software 130 includes multiple software modules for processing material attribute data and/or clinical data associated with one or more pharmaceutical materials, such as a data extraction module 132, a data processing module 134, a feature generation module 136, a model training module 138, and a correlation module 142. In the embodiment of FIG. 1, the software 130 additionally includes a user interface module 144 for obtaining user input.
[00104] In some embodiments, data extraction module 132 is generally responsible for retrieving/obtaining data (e.g., material attribute data and/or clinical data associated with one or more pharmaceutical materials) from the database 120. For example, such database 120 can comprise attribute data silos 302, clinical data silos 304, and data storage 332, as illustrated in FIG. 3. In some embodiments, data extraction module 132 retrieves material attribute data and/or clinical data associated with one or more pharmaceutical materials based on user input detected by user interface module 144. For example, user interface module 144 may generate and/or populate a GUI, and cause display 116 to present the GUI to a user. The user may then operate user input device 118 to enter one or more pharmaceutical materials via the GUI, and data extraction module 132 may retrieve material attribute data and/or clinical data associated with one or more pharmaceutical materials. In some embodiments,
the database 120 includes raw data, and data extraction module 132 constructs similar data structure(s). For example, data extraction module 132 may generate data in a more readily usable form.
[00105] In some embodiments, data processing module 134 is generally responsible for processing the data (e.g., material attribute data and/or clinical data associated with one or more pharmaceutical materials) from the database 120 extracted via the data extraction module 134. Data processing module 134 can apply one or more transformations to the data from the database 120 (e.g., raw data of material attribute data and/or clinical data associated with one or more pharmaceutical materials) using any type of computational or mathematical techniques, including machine learning techniques. The one or more transformations can include at least one of cleaning, merging, associating, selecting, or grouping to create modified material attribute data and modified clinical data. Data processing module 134 can process the data based on user input detected by user interface module 144. For example, user interface module 144 may generate and/or populate a GUI, and cause display 116 to present the GUI to a user. The user may then operate user input device 118 to enter one or more instructions related to processing the data via the GUI, and data processing module 134 may process the data based on the user input. In some embodiments, the database 120 includes raw data (e.g., data without any processing steps), and data processing module 134 can process the raw data so the data can be utilized by other modules (e.g., feature generation module 136) or processes. For example, data extraction module 132 may normalize the raw data, or may generate data in a more readily usable form (e.g., a table containing normalized material attribute data and/or clinical data associated with one or more pharmaceutical materials).
[00106] In some embodiments, the feature generation module 136 obtains processed data from the database 120, the data extraction module 132, and/or the data processing module 134, and uses the processed data to generate sets of features. Such features can be any type of material attributes including a molecular attribute, a process related impurity, or a drug substance feature. The feature generation module 136 may generate a set of features for material attributes of one pharmaceutical material at different timepoints. In some embodiments, the feature generation module 136 generates a set of features by including at least some of the obtained data in the set of features. For example, the feature generation module 136 may generate the set of features to include material attributes of one pharmaceutical material at different timepoints. For example, feature generation module 136
may generate the set of features to include a two-dimensional (2D) matrix that that stores values of material attributes of one pharmaceutical material as y dimension, and different time points as x dimension. The generated 2D matrices may be provided as input to the machine learning model. Additionally, or alternatively, the feature generation module 136 may generate a set of features including encoded data. For example, the material attributes of one pharmaceutical material at different timepoints may be one-hot encoded. The feature generation module 136 may include additional or alternative features in the set of features, as aspects of the technology described herein are not limited in this respect.
[00107] In some embodiments, the model training module 138 may be configured to train one or more models (e.g., machine learning models) to determine an impact of attributes. In some embodiments, the model training module 138 trains a machine learning model using at least one of material attribute data, clinical data, modified material attribute data, or modified clinical data of one pharmaceutical material at different timepoints. For example, the model training module 138 may obtain material attribute data, clinical data, modified material attribute data, and/or modified clinical data of one pharmaceutical material at different timepoints from the database 120. In some embodiments, the model training module 138 may provide trained machine learning model(s) to the database 120. Techniques for training a machine learning model are described elsewhere herein.
[00108] In some embodiments, the correlation module 142 obtains one or more sets of features from the feature generation module 136, obtains a trained machine learning model from the model training module 138 and database 120 (which may be a data store of any suitable type), and processes the obtained set(s) of features using the obtained machine learning model to determine a correlation status associated with a clinical impact of one or more material attributes. For example, the correlation module 136 may process the set of features generated using the trained machine learning model to obtain values of a correlation status associated with a clinical impact of one or more material attributes. Techniques for determining a correlation status using machine learning are described elsewhere herein. In some embodiments, a correlation status associated with a clinical impact of one or more material attributes may be output by the correlation module 142. For example, the correlation status may be output to user(s) 150 via user interface module 144. Additionally, or alternatively, the correlation status may be stored in memory and/or transmitted to one or more other computing devices.
[00109] As shown in FIG. 1, system 100 also includes database 120. The database 120 may store model data, material attribute data, clinical data, modified material attribute data, modified clinical data, and/or any type of raw data associated with one or more pharmaceutical materials. In some embodiments, software 130 obtains data from database 120 and/or user(s) 150 (e.g., by uploading data). The database 120 may be of any suitable type (e.g., database system, multi-file, flat file, etc.) and may store data in any suitable way and in any suitable format, as aspects of the technology described herein are not limited in this respect. The database 120 may be part of software 130 (not shown) or excluded from software 130, as shown in FIG. 1. The database 120 may be part of or external to computing system 110.
[00110] In some embodiments, the stored data may have been previously uploaded by a user (e.g., user 150), and/or from one or more public data stores and/or studies. In some embodiments, a portion of the data may be processed by the data processing module 134 to obtain processed data. In some embodiments, a portion of the data may be processed by the feature generation module 136 to generate sets of features to be provided as input to a machine learning model. In some embodiments, a portion of the data may be used to train one or more machine learning models (e.g., with the model training module 138).
[00111] User interface module 144 may be a graphical user interface (GUI), a text-based user interface, and/or any other suitable type of interface through which a user may provide input and view information generated by software 130. For example, in some embodiments, the user interface may be a webpage or web application accessible through an Internet browser. In some embodiments, the user interface may be a graphical user interface (GUI) of an app executing on the user’s mobile device. In some embodiments, the user interface may include a number of selectable elements through which a user may interact. For example, the user interface may include dropdown lists, checkboxes, text fields, or any other suitable element. [00112] FIG. 2 is a flowchart of an illustrative method 200 for determining an impact of attributes, in accordance with some embodiments of the technology described herein. One or more steps of method 200 may be performed automatically by any suitable computing system(s). For example, the step(s) may be performed by a laptop computer, a desktop computer, one or more servers, in a cloud computing environment, computer system 100, and computing device 1200 as described herein within respect to FIG. 12, and/or in any other suitable way. For example, in some embodiments, step 202 may be performed automatically
by any suitable computing system(s) and/or device(s). As another example, step 204 may be performed automatically by any suitable computing system(s) and/or device(s).
[00113] Step 202 may include obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material. The material attribute data can comprise measurement data of the one or more material attributes at one or more timepoints. The pharmaceutical material can comprise at least one of a biological therapy, a synthetic small molecule, or a nucleic acid. The biological therapy can comprise or be selected from the group consisting of: an antibody, an antigen-binding antibody fragment, an antibody protein product, a Bi-specific T cell engager (BiTE®) molecule, a bispecific antibody, a trispecific antibody, an Fc fusion protein, a recombinant protein, a recombinant virus, a recombinant T cell, a synthetic peptide, and an active fragment of a recombinant protein. The synthetic small molecule can comprise any small molecule that is artificially manufactured in a laboratory, using various chemical processes. The nucleic acid can comprise an siRNA, an mRNA or a DNA. The manufacturing the pharmaceutical material can comprise culturing a genetically engineered mammalian host cell comprising one or more nucleic acids encoding the biological therapy. The pharmaceutical material can be in a pharmaceutically acceptable formulation. Details of the pharmaceutical materials are described elsewhere herein.
[00114] The one or more material attributes can comprise at least one of a molecular attribute, a process related impurity, or a drug substance feature. The molecular attribute can comprise at least one of: acidic species, basic species, high molecular weight species, subvisible particle number, visible particles, aggregation, low molecular weight, middle molecular weight, glycosylation (such as non-glycosylated heavy chain or high mannose), glycation, deamidation, deamination, cyclization, oxidation, sulfation, hydroxylysine, isomerization, fragmentation/clipping, N-terminal and C-terminal variants, signal peptide, reduced and partial species, misfolding, disulfide scrambling , domain swapping, folded structure, surface hydrophobicity, chemical modification, glycation, covalent bond, mutations or misincorporations, a C-terminal amino acid motif PARG, a C-terminal amino acid motif PAR- Amide , drug antibody ratio (DAR), or peptide antibody ratio (PAR). The process related impurities can comprise at least one of Chinese Hamster Ovary Protein (CHOP), Host Cell Protein (HCP), residual host cell DNA, residual ProA, or a process reagent. The drug substance feature can comprise at least one of a component feature or a drug administration feature. The component feature can comprise osmolality or viscosity.
The drug administration feature can comprise syringe type or break-loose extrusion (BLE). Details of the material attributes are described elsewhere herein.
[00115] The measurement data of the one or more material attributes at the one or more timepoints can be determined by at least one of mass spectrometry, chromatography, electrophoresis, spectroscopy, light obscuration, a particle method, analytical centrifugation, imaging or imaging characterization, immunoassay, or multivariate analysis of any of these or other analytical methods yielding emergent features. In one example, a linearized peptide may have residues of interest at seemingly random positions along the peptide length, and in the folded 3D conformation the residues of interest can exhibit the feature of interest which can be an emergent phenomenon of two or more analytical methods including but not limited to mass spectrometry, x-ray crystallography, and molecular modeling. The measurement data of the one or more material attributes can show the level and/or value of the one or more material attributes. For example, if the material attribute is acidic species, the measurement data can be the percentage of acidic variants by cation-exchange chromatography (CEX) of the pharmaceutical material. The material attribute data can comprise a change of the measurement data of the one or more material attributes at the one or more time points. For instance, the material attribute data can comprise a change of pH of a pharmaceutical material between 3 months and 12 months. The material attribute data can comprise a duration that the pharmaceutical material was under storage conditions prior to the administration of the pharmaceutical material. Such duration can be at least 1 month, 2 months, 3 months, 6 months, 12 months, 18 months, 24 months or longer. Such duration can be at most 36 months, 24 months, 18 months, 12 months, 6 months, 3 months, 2 months, 1 month or shorter. The material attribute data can comprise a dose of the pharmaceutical material in the administration. The material attribute data can comprise a level of material attribute exposure received by the one or more subjects at the time of the administration. The one or more timepoints can comprise at least one of a timepoint of manufacture or a timepoint of lot release. The measurement data of the one or more material attributes can be detected at two or more timepoints under storage conditions. The one or more timepoints can comprise a timepoint of manufacture and at least two subsequent timepoints.
[00116] The number of one or more timepoints can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or more timepoints. In some embodiments, the number of one or more timepoints can be at most 90, 80, 70, 60, 50, 40, 30, 20, 10, or less timepoints. In some embodiments, the one or more timepoints span a period of months or years. In some
embodiments, one or more timepoints span a period of months or years from the beginning of storage of one pharmaceutical material. In some embodiments, the one or more timepoints span at least 1 month, 2 months, 3 months, 4 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 18 months, 24 months, 30 months, 36 months or longer. In some embodiments, the one or more timepoints span at most 36 months, 30 months, 24 months, 18 months, 12 months, 11 months, 10 months, 9 months, 8 months, 7 months, 6 months, 5 months, 4 months, 3 months or shorter.
[00117] Step 204 may include obtaining, via the one or more processors, clinical data associated with the pharmaceutical material. The clinical data can comprise subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material. The one or more clinical events can comprise one or more clinical adverse events associated with the one or more subjects that have received administration of the pharmaceutical material. The one or more clinical events can comprise any type of events associated with the one or more subjects that have received administration of the pharmaceutical material, including, but not limited to, customer complaints, faulty autoinjectors, and confusion with instruction sheets.
[00118] The subject data can comprise at least one of preexisting condition, biomarker, laboratory result, metabolomic data, or demographic information of the one or more subjects that have received the administration of the pharmaceutical material. The preexisting conditions can comprise sex, age, weight, medical conditions, and concurrently provided medications. The biomarkers can comprise preexisting anti-drug antibodies, genetic markers, metabolomic markers, and serum pharmaceutical material concentration. The laboratory results can comprise neutralizing anti-drug antibodies, metabolomic signatures, and serum pharmaceutical material concentration. The demographic information can comprise race, treatment location, life habits (such as smoking/nonsmoking), socioeconomic standing, sociopolitical data at time of treatment, and climate data at the time of treatment (e.g., long treatment windows of a treatment may overlap with annual phenomena such as “flu season” which may increase baseline rates of reported adverse events and may not be directly linked to the pharmaceutical material or some pharmaceutical material attribute).
[00119] Subjects can comprise any living or non-living organism, including but not limited to a human (e.g., a male human, female human, fetus, pregnant female, child, or the like), a nonhuman animal, a plant, a bacterium, a fungus or a protist. Any human or non-human animal can serve as a subject, including but not limited to mammal, reptile, avian, amphibian, fish,
ungulate, ruminant, bovine (e.g., cattle), equine (e.g., horse), caprine and ovine (e.g., sheep, goat), swine (e.g., pig), camelid (e.g., camel, llama, alpaca), monkey, ape (e.g., gorilla, chimpanzee), ursid (e.g., bear), poultry, dog, cat, mouse, rat, fish, dolphin, whale and shark. In some embodiments, a subject is a male or female of any stage (e.g., a man, a women or a child).
[00120] Step 206 may include applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data. The one or more transformations can include at least one of cleaning, merging, associating, selecting, or grouping. The one or more transformations can comprise cleaning the material attribute data and/or the clinical data. The cleaning can comprise normalization of the material attribute data and/or the clinical data. The cleaning can comprise adjusting the data format or the numeric representation of the material attribute data and/or the clinical data. Such cleaning can comprise filtering the material attribute data and/or the clinical data based on one or more filtering criteria. The one or more filtering criteria can comprise removing one or more orphan adverse events. The one or more filtering criteria can comprise a selection of a column of an excel file. The one or more filtering criteria can comprise relational filtering such as removing adverse events that may have a low likelihood of being related to attribute levels, removing adverse events before a treatment window start, grouping attributes or adverse events for multivariate analysis based on the outputs of clustering tools and dimensionality manipulation algorithms such as LDA, connecting referential data from an arbitrary number of data silos such as socioeconomic and climate databases, harmonizing medical terms and names of adverse events, and removing incomplete data or patients (subjects) that dropped out of the study before starting the treatment.
[00121] The one or more transformations can comprise merging the material attribute data and/or the clinical data. Such merging can comprise combining the material attribute data and the clinical data based on a level of similarity between the material attribute data and the clinical data. For instance, material attribute data and clinical data can be combined based on the similarity of the data format based on common factors such as subject, treatment location, treatment modality, and/or datetime. The merging can comprise combining different sets of material attribute data based on a level of similarity between the different sets of material attribute data. The merging can comprise combining different sets of clinical data based on a level of similarity between the different sets of clinical data. In some embodiments, such
merging can comprise linearizing the material attribute data and the clinical data based on lots and material attributes. In one example, the merging comprises joining the material attributes data with exposure data. The merging can comprise organizing and structuring data by datetime to explore the effects of time of year on baseline adverse event rates (e.g., exploring the placebo group). The merging can comprise organizing and structuring data by location to determine if a certain region, hospital, or climactic zone may have some impact on baseline adverse event rates. The merging can comprise organizing and structuring data based on treatment modality to determine the differences in adverse events between intravenous dosing and syringe-based dosing, or differences in adverse events between upper abdominal injection and gluteal injection.
[00122] The one or more transformations can comprise associating the material attribute data and/or the clinical data. Such associating can comprise correlating the material attribute data and clinical data based on one or more associating criteria. The one or more associating criteria can comprise a similar geographic location and/or a same preexisting condition of one or more subjects. The one or more associating criteria can comprise a similar administration date, a same dose, one or more similar adverse events, and an attribute level exposure. The material attribute data and/or the clinical data can be associated, aligned or combined on an individual subject basis. In this case, for every subject, the level of exposure for each attribute can be added to the excel sheet that contains the subject clinical data so that every line item for every subject can have the administration date, dose, adverse events, and attribute level exposure listed all together for every timepoint. The associating can comprise correlating different sets of material attribute data on one or more associating criteria. The associating can comprise correlating different sets of clinical data on one or more associating criteria.
[00123] The one or more transformations can select the material attribute data and/or the clinical data. Such selecting can comprise selecting the material attribute data and/or the clinical data based on one or more selection criteria. The one or more selection criteria can comprise the use of some discriminant analysis and some confidence interval, output from a method such as a classification algorithm, and data features such as ones determined by a deep learning algorithm. For example, the one or more selection criteria can comprise a numerical threshold value to output of a classification model, and input data that yield an output above the numerical threshold value can be selected.
[00124] The one or more transformations can group the material attribute data and/or the clinical data. Such grouping can comprise generating one or more subgroups based on one or
more patterns of the material attribute data and/or the clinical data. The one or more patterns of material attribute data and/or the clinical data can comprise a pattern of one or more subjects, a pattern of geographic location, a pattern of age, a pattern of adverse events, or a pattern of gender. In this case, if, for example, the subjects in one location demonstrate the same adverse events, then the subject in that location can be grouped into one subgroup. In one example, the one or more subgroups can comprise a subgroup based on subjects (320 of FIG. 3), a subgroup based on adverse event (322 of FIG. 3), and a subgroup based on arbitrary features (324 of FIG. 3), as shown in FIG. 3.
[00125] Step 208 can comprise determining, via the one or more processors, the correlation status based on the modified material attribute data and the modified clinical data using a computational model. In some embodiments, the determining the correlation status can comprise determining the correlation status based on at least one of the material attribute data, the clinical data, the modified material attribute data, or the modified clinical data using a computational model. The computational model can comprise at least one of: a logistic regression model, a support vector machine model, a multinomial logistic regression model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, a Bayesian estimation approach, a thermo-kinetic model, or a Markov model. The computational model can comprise any type of mathematical, statistical, or machine learning model, including, but not limited to, a supervised machine learning model, an unsupervised machine learning model, a reinforcement learning model, a regression model (e.g., a logistic regression model, a multinomial logistic regression model), a support vector machine model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, and/or any other suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect. In some embodiments, the machine learning model(s) may include an ensemble of machine learning models of any suitable type (the machine learning models part of the ensemble may be termed "weak learners"). The training can comprise inputting at least one of the material attribute data, the clinical data, the modified material attribute data, or the modified clinical data into the computational model.
[00126] As described above, in some embodiments, the machine learning model(s) may be implemented as a decision tree classifier. Any suitable type of decision tree classifier may be used and may be trained using any suitable supervised decision tree learning technique. For
example, the decision tree classifier may be trained by the iterative dichotomizer technique (e.g., the ID3 algorithm as described, for example, in Quinlan, J. R. 1986. Induction of Decision Trees. Mach. Learn. 1, 1 (Mar. 1986), 81-106)), the C4.5 technique (e.g., as described, for example, in Quinlan, J. R. C4.5: Programs for Machine Learning. Morgan Kaufmann Publishers, 1993), the classification and regression tree (CART) technique (e.g., as described, for example, in Breiman, Leo; Friedman, J. H.; Olshen, R. A.; Stone, C. J. (1984). Classification and regression trees. Monterey, CA: Wadsworth & Brooks/Cole Advanced Books & Software). A decision tree classifier may be trained using any other suitable training method, as aspects of the technology described herein are not limited in this respect.
[00127] In some embodiments, a gradient-boosted decision tree classifier may be used. The gradient-boosted decision tree classifier may be an ensemble of multiple decision tree classifiers (sometimes called "weak learners"). The prediction (e.g., classification) generated by the gradient-boosted decision tree classifier can be formed based on the predictions generated by the multiple decision trees part of the ensemble. The ensemble may be trained using an iterative optimization technique involving calculation of gradients of a loss function (hence the name "gradient" boosting). Any suitable supervised training algorithm may be applied to training a gradient-boosted decision tree classifier including, for example, any of the algorithms described in Hastie, T.; Tibshirani, R.; Friedman, J. H. (2009). "10. Boosting and Additive Trees". The Elements of Statistical Learning (2nd ed.). New York: Springer, pp. 337-384. In some embodiments, the gradient-boosted decision tree classifier may be implemented using any suitable publicly available gradient boosting framework such as XGBoost (e.g., as described, for example, in Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). New York, NY, USA: ACM.). The XGBoost software may be obtained from http://xgboost.ai, for example).
Another example framework that may be employed is LightGBM (e.g., as described, for example, in Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., ... Liu, T.-Y. (2017). Lightgbm: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146-3154.). The LightGBM software may be obtained from https://lightgbm.readthedocs.io/, for example).
[00128] In some embodiments, a neural network classifier may be used. The neural network classifier may be trained using any suitable neural network optimization software. The optimization software may be configured to perform neural network training by gradient
descent, stochastic gradient descent, or in any other suitable way. In some embodiments, the Adam optimizer (Kingma, D. and Ba, J. (2015) Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015)) may be used.
[00129] In some embodiments, a cluster model, described at pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York, (hereinafter “Duda 1973”), may be used. The cluster model can find natural groupings in a dataset. To identify natural groupings, measure similarity (or dissimilarity) between two samples maybe determined based on a distance function and to compute the matrix of distances between all pairs of samples in the training set. Once a method for measuring “similarity” or “dissimilarity” between points in a dataset is selected, clustering can use a criterion function that measures the clustering quality of any partition of the data. Partitions of the data set that extremize the criterion function are used to cluster the data.
[00130] In some embodiments, principal component analysis (PCA) algorithms, one type of dimensionality model, described in Jolliffe, 1986, Principal Component Analysis, Springer, New York, may be used. Principal components (PCs) can be uncorrelated and can be ordered such that the kth PC has the kth largest variance among PCs. The kth PC can be interpreted as the direction that maximizes the variation of the projections of the data points such that it is orthogonal to the first k-1 PCs. The first few PCs can capture most of the variation in a training set. In contrast, the last few PCs can often be assumed to capture only the residual ‘noise’ in the training set.
[00131] In some embodiments, a support vector machine (SVM) model, described in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, can be used. When used for classification, SVMs can separate a given set of binary labeled data training set with a hyperplane that is maximally distant from the labeled data. For cases in which no linear separation is possible, SVMs can work in combination with the technique of ‘kernels’, which automatically realizes a non-linear mapping to a feature space. The hyper-plane found by the SVM in feature space can correspond to a non-linear decision boundary in the input space. [00132] The determining the correlation status can comprise using the computational model to identify one or more patterns associated with the modified material attribute data and the modified clinical data. The correlation status can comprise a presence of correlation or an absence of correlation. In one example, the machine learning classifier can generate output as
0 or 1, where 0 represents an absence of correlation and 1 represents a presence of correlation. The determining the correlation status can comprise using clustering and/or classification algorithms to show co-variance of otherwise distinct features of the dataset, using machine learning to find emergent phenomena (e.g., headache and fatigue) in the dataset that may be treated as one emergent feature (e.g., headache-fatigue), and using modeling to predict a correlation dynamic (e.g., curve analysis using modeled curves extending beyond measured values or timepoints) that otherwise is not evident using measured datapoints. The emergent feature can comprise features of the dataset that are apparent using some structured analysis and not part of the raw data. One example of an emergent adverse event can comprise a headache-fever as determined during a CIA assessment, where a computational method determines that the adverse events "headache" and "fever" correlate and can be considered as one adverse event with potentially related associations to attributes. The determining the correlation status can comprise grouping into quartiles based on attribute level exposure and assessing if the quartiles align on a positive slopping line showing a positive correlation. The computer-implemented method can further comprise, if the correlation status comprises an absence of correlation, determining that the one or more material attributes impact neither clinical safety nor efficacy of the pharmaceutical material. The clinical safety can comprise a frequency of clinical adverse events. The efficacy of the pharmaceutical material can comprise a maximum response that can be achieved with the pharmaceutical material.
[00133] The computer-implemented method can further comprise, if the correlation status comprises a presence of correlation, determining that the one or more material attributes impact at least one of safety or efficacy of the pharmaceutical material. The computer- implemented method can further comprise, if the correlation status comprises an absence of correlation, setting a specification for permissible levels of the one or more material attributes of the pharmaceutical material. The permissible levels of the one or more material attributes are based on one or more levels of material attribute exposure received by the subjects. The permissible levels can comprise any levels prescribed by industry prior knowledge, industry/regulatory guidance, in-vitro or in-vivo measurement, or laboratory quantitation. The permissible levels of the one or more material attributes can comprise an actual level of a material attribute, such as 2% of high molecular weight or 10% high mannose. The computer- implemented method can further comprise, if the correlation status comprises a presence of correlation, setting a specification for maximum permissible levels of the one or more
material attributes of the pharmaceutical material. In this case, the maximum permissible levels of the one or more material attributes can be based on one or more levels of the one or more material attributes associated with at least one of clinical adverse events or inhibition of efficacy of the pharmaceutical material.
[00134] The computer-implemented method can further comprise, if the correlation status comprises an absence of correlation, manufacturing production lots of the pharmaceutical material comprising the one or more material attributes at or below a specified permissible level of the one or more material attributes based on one or more levels of material attribute exposure. The computer-implemented method can further comprise, if the correlation status comprises an absence of correlation, manufacturing production lots of the pharmaceutical material comprising the one or more material attributes at or below one or more specified permissible levels of the one or more material attributes based on one or more levels of material attribute exposure. In this case, each one of the one or more specified permissible levels can correspond to one of the one or more material attributes.
[00135] The computer-implemented method can further comprise, if the correlation status comprises a presence of correlation, setting a specification for levels of the one or more material attributes at the time of manufacture that do not exceed a maximum permissible level of the one or more material attributes of the pharmaceutical material. The computer- implemented method can further comprise, if the correlation status comprises a presence of correlation, setting a specification for levels of the one or more material attributes at the time of manufacture that do not exceed one or more maximum permissible levels of the one or more material attributes of the pharmaceutical material. In this case, each one of the one or more maximum permissible levels can correspond to one of the one or more material attributes. The computer-implemented method can further comprise, if the correlation status comprises an absence of correlation, establishing a manufacturing process to produce levels of the one or more material attributes at or below a permissible level based on the correlation status. The manufacturing process can comprise a roller bottle manufacturing process, a 2L bioreactor manufacturing process, a continuous manufacturing process, and a fed-batch manufacturing process.
[00136] The computer-implemented method can further comprise generating a rank of the one or more material attributes based on the correlation status using the computational model. In one example, such rank can comprise a list of material attributes that impact clinical safety and efficacy of the pharmaceutical material. In this example, the list of material attributes are
ordered based on a level of impact of each of the material attributes. The computer- implemented method can further comprise selecting a subset of the one or more material attributes based on the rank of the one or more material attributes and setting a specification for permissible levels of the subset of the one or more material attributes. In one example, the subset of the one or more material attributes can comprise top 1, top 5, or top 10 material attributes on the rank of the one or more material attributes. The computer-implemented method can further comprise generating one or more heuristics (e.g., histograms) associated with the one or more subjects based on the modified material attribute data and the modified clinical data. In some embodiments, the one or more heuristics can comprise, but not limited to, a vertical bar chart 330-1, a horizontal bar chart 330-2, a radar chart 330-3, a pie chart 330-4, a line chart 330-5, a hierarchical chart 330-6, and a venn diagram 330-7, as shown in FIG. 3.
[00137] The computer-implemented method can further comprise estimating one or more parameters associated with the administration of the pharmaceutical material. The estimating one or more parameters can comprise using molecular modeling, degradation modeling, stability modeling, in-vivo mechanisms modeling (e.g., estimating aggregate formation or dissociation in-vivo), pharmacokinetic modeling, or stochastic modeling of the drug molecule under various stressor environments to estimating one or more parameters. The estimating one or more parameters can comprise using exponential decay modeling to estimate the impact on PK or clearance of an attribute (e.g., high mannose). The one or more parameters can comprise at least one of a stability expanded time-course dosing, a pharmacokinetic expanded time-course dosing, a step dosing, or a multiple dose overlap time-course dosing. The computer-implemented method can further comprise estimating a cooperative effect of the one or more material attributes. The estimating a cooperative effect of the one or more material attributes can comprise using machine learning methods, clustering methods, and deep learning to demonstrate co-variant features of the dataset not evident to manual analysis (e.g., emergent phenomena). The co-variant features can include synergistically covariant features and inversely covariant features. Clinical data from other molecules (prior knowledge) can be used to generate a model to estimate the effect and/or behavior of a new molecule. The cooperative effect can comprise at least one of an additive effect, an inhibitory effect, or a feedback effect. HMW can comprise chemical modifications that can make it more immunogenic (additive), less immunogenic (inhibitory), or reacts with another signaling molecule.
[00138] A computer-implemented method for training a computational model for determining a correlation status associated with a clinical impact of one or more material attributes can comprise (a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (d) training, via the one or more processors, the computational model using the modified material attribute data and the modified clinical data. In some embodiments, the step (d) can comprise training, via the one or more processors, the computational model using at least one of the material attribute data, the clinical data, the modified material attribute data, or the modified clinical data.
[00139] The training can comprise inputting the modified material attribute data and the modified clinical data into the computational model. The training can comprise inputting at least one of the material attribute data, the clinical data, the modified material attribute data, or the modified clinical data into the computational model. The training can comprise inputting at least one of the material attribute data, the clinical data, the modified material attribute data, the modified clinical data, or any type of data associated with the pharmaceutical material (e.g., predictive data or modified predictive data) into the computational model. After training the computational model, a validation process can be performed. The validation process can comprise any type of validation technique, including, but not limited to, cross-validation, k-fold cross validation, leave-one-out cross-validation, bootstrapping, Monte Carlo cross-validation, holdout validation, and shuffle split. The validation process can be used to tune the hyperparameters. For example, a predefined range of value for each of the hyperparameters can be assumed, and then for each possible combination of the hyperparameter values, models can be built iteratively in such a way one of one or more pharmaceutical materials is left out for validation purposes (e.g., leave-one- out cross validation). In this case, if material attribute data and clinical data of n
pharmaceutical materials are used to train a computational model, in each iteration, a computational model can be built using material attribute data and clinical data of n-1 out of n pharmaceutical materials and with the remaining material attribute data and clinical data of 1 pharmaceutical material as validation.
[00140] This procedure can be repeated for all possible combinations of hyperparameter values in the predefined range. The combination with the lowest error can be chosen for the given dataset. Once the hyperparameters are determined, the functional form of the computational model can be set, and the model parameters can be obtained using an iterative optimization procedure. Since the results of an optimization procedure depend on the initial values, the procedure can be repeated several times to account for these differences. In some embodiments, the differences can be used to calculate a confidence interval (e.g., 95% confidence interval) of the predictions of the computational model.
[00141] FIG. 10 is a flowchart of another illustrative method 1000 for determining an impact of attributes, in accordance with some embodiments of the technology described herein. One or more steps of method 1000 may be performed automatically by any suitable computing system(s). For example, the step(s) may be performed by a laptop computer, a desktop computer, one or more servers, in a cloud computing environment, computer system 100, and computing device 1200 as described herein within respect to FIG. 12, and/or in any other suitable way. For example, in some embodiments, step 1002 may be performed automatically by any suitable computing system(s) and/or device(s). As another example, step 1004 may be performed automatically by any suitable computing system(s) and/or device(s).
[00142] Step 1002 may include obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material. The material attribute data can comprise measurement data of the one or more material attributes at one or more timepoints. The pharmaceutical material can comprise at least one of a biological therapy, a synthetic small molecule, or a nucleic acid. The biological therapy can comprise or be selected from the group consisting of: an antibody, an antigen-binding antibody fragment, an antibody protein product, a Bi-specific T cell engager (BiTE®) molecule, a bispecific antibody, a trispecific antibody, an Fc fusion protein, a recombinant protein, a recombinant virus, a recombinant T cell, a synthetic peptide, and an active fragment of a recombinant protein. The synthetic small molecule can comprise any small molecule that is artificially manufactured in a laboratory, using various chemical processes. The nucleic acid can comprise an siRNA, an mRNA or a DNA. The manufacturing the pharmaceutical material
can comprise culturing a genetically engineered mammalian host cell comprising one or more nucleic acids encoding the biological therapy. The pharmaceutical material can be in a pharmaceutically acceptable formulation. Details of the pharmaceutical materials are described elsewhere herein.
[00143] The one or more material attributes can comprise at least one of a molecular attribute, a process related impurity, or a drug substance feature. The molecular attribute can comprise at least one of: acidic species, basic species, high molecular weight species, subvisible particle number, visible particles, aggregation, low molecular weight, middle molecular weight, glycosylation (such as non-glycosylated heavy chain or high mannose), glycation, deamidation, deamination, cyclization, oxidation, sulfation, hydroxylysine, isomerization, fragmentation/clipping, N-terminal and C-terminal variants, signal peptide, reduced and partial species, misfolding, disulfide scrambling , domain swapping, folded structure, surface hydrophobicity, chemical modification, glycation, covalent bond, mutations or misincorporations, a C-terminal amino acid motif PARG, a C-terminal amino acid motif PAR- Amide , drug antibody ratio (DAR), or peptide antibody ratio (PAR). The process related impurities can comprise at least one of CHOP, HCP, residual host cell DNA, residual ProA, or a process reagent. The drug substance feature can comprise at least one of a component feature or a drug administration feature. Details of the material attributes are described elsewhere herein.
[00144] The measurement data of the one or more material attributes at the one or more timepoints can be determined by at least one of mass spectrometry, chromatography, electrophoresis, spectroscopy, light obscuration, a particle method, analytical centrifugation, imaging or imaging characterization, or immunoassay. The measurement data of the one or more material attributes can show the level and/or value of the one or more material attributes. The material attribute data can comprise a change of the measurement data of the one or more material attributes at the one or more time points. Details of the measurement data and one or more time points are described elsewhere herein.
[00145] Step 1006 may include generating, via the one or more processors, predictive data associated with the pharmaceutical material based on the material attribute data and the clinical data using a predictive model. The predictive model can comprise any type of statistical or mathematical model, including, but not limited to a pharmacokinetic/pharmacodynamic modeling (PK/PD) model, a supervised machine learning model, an unsupervised machine learning model, a reinforcement learning model, a
regression model (e.g., a logistic regression model, a multinomial logistic regression model), a support vector machine model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, and/or any other suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect. In some embodiments, the predictive model is the same as the computational model. In some other embodiments, the predictive model is different from the computational model. Details of different types of models are described elsewhere herein. The predictive data can be generated by inputting the material attribute data and/or the clinical data into the predictive model, so that the predictive data is the output of the predictive model. The predictive data can comprise any type of data generated from the predictive model using the material attribute data and/or the clinical data. The predictive data can comprise any parameters and/or value of the parameters associated with the predictive model. In some embodiments, the predictive data can comprise any type of data associated with the pharmaceutical material but not derived from the material attribute data and/or the clinical data.
[00146] Step 1008 may include applying, via the one or more processors, one or more transformations to the material attribute data, the clinical data, and the predictive data to create modified material attribute data, modified clinical data, and modified predictive data. The one or more transformations can include at least one of cleaning, merging, associating, selecting, or grouping. The one or more transformations can comprise cleaning the material attribute data, the clinical data, and the predictive data. The cleaning can comprise normalization of the material attribute data, the clinical data, and the predictive data. The cleaning can comprise adjusting the data format or the numeric representation of the material attribute data, the clinical data, and the predictive data. Such cleaning can comprise filtering the material attribute data, the clinical data, and the predictive data based on one or more filtering criteria. Details of the one or more filtering criteria are described elsewhere herein. [00147] The one or more transformations can comprise merging the material attribute data, the clinical data, and/or the predictive data. Such merging can comprise combining the material attribute data, the clinical data, and the predictive data based on a level of similarity between the material attribute data and the clinical data. For instance, the material attribute data, the clinical data, and the predictive data can be combined based on the similarity of the data format based on common factors such as subject, treatment location, treatment modality, and/or datetime. In some embodiments, such merging can comprise linearizing the material
attribute data, the clinical data, and the predictive data based on lots and material attributes. In one example, the merging comprises joining the material attributes data with exposure data. The merging can comprise organizing and structuring data by datetime to explore the effects of time of year on baseline adverse event rates (e.g., exploring the placebo group). The merging can comprise organizing and structuring data by location to determine if a certain region, hospital, or climactic zone may have some impact on baseline adverse event rates. The merging can comprise organizing and structuring data based on treatment modality to determine the differences in adverse events between intravenous dosing and syringe-based dosing, or differences in adverse events between upper abdominal injection and gluteal injection.
[00148] The one or more transformations can comprise associating the material attribute data, the clinical data, and/or the predictive data. Such associating can comprise correlating the material attribute data, the clinical data, and the predictive data based on one or more associating criteria. The one or more associating criteria can comprise a similar geographic location and/or a same preexisting condition of one or more subjects. The one or more associating criteria can comprise a similar administration date, a same dose, one or more similar adverse events, and an attribute level exposure. The material attribute data, the clinical data, and the predictive data can be associated, aligned or combined on an individual subject basis. In this case, for every subject, the level of exposure for each attribute can be added to the excel sheet that contains the subject clinical data so that every line item for every subject can have the administration date, dose, adverse events, and attribute level exposure listed all together for every timepoint (e.g., data of administration).
[00149] The one or more transformations can select the material attribute data, the clinical data, and/or the predictive data. Such selecting can comprise selecting the material attribute data, the clinical data, and/or the predictive data based on one or more selection criteria. The one or more selection criteria can comprise the use of some discriminant analysis and some confidence interval, output from a method such as a classification algorithm, and data features such as ones determined by a deep learning algorithm. For example, the one or more selection criteria can comprise a numerical threshold value to output of a classification model, and input data that yield an output above the numerical threshold value can be selected.
[00150] The one or more transformations can group the material attribute data, the clinical data, and/or the predictive data. Such grouping can comprise generating one or more
subgroups based on one or more patterns of the material attribute data, the clinical data, and the predictive data. The one or more patterns of the material attribute data, the clinical data, and/or the predictive data can comprise a pattern of one or more subjects, a pattern of geographic location, a pattern of age, a pattern of adverse events, or a pattern of gender. In this case, if, for example, the subjects in one location demonstrate the same adverse events, then the subject in that location can be grouped into one subgroup. In one example, the one or more subgroups can comprise a subgroup based on subjects (320 of FIG. 3), a subgroup based on adverse event (322 of FIG. 3), and a subgroup based on arbitrary features (324 of FIG. 3), as shown in FIG. 3
[00151] Step 1010 can comprise determining, via the one or more processors, the correlation status based on at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data using a computational model. In some embodiments, the determining the correlation status can comprise determining the correlation status based on the modified material attribute data, the modified clinical data, and the modified predictive data using a computational model. The computational model can comprise at least one of: a logistic regression model, a support vector machine model, a multinomial logistic regression model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, or a Markov model. The computational model can comprise any type of mathematical, statistical, or machine learning model, including, but not limited to, a supervised machine learning model, an unsupervised machine learning model, a reinforcement learning model, a regression model (e.g., a logistic regression model, a multinomial logistic regression model), a support vector machine model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, and/or any other suitable type of machine learning model, as aspects of the technology described herein are not limited in this respect. In some embodiments, the machine learning model(s) may include an ensemble of machine learning models of any suitable type (the machine learning models part of the ensemble may be termed "weak learners"). The training can comprise inputting at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data into the computational model. Details of the computational model are described elsewhere herein.
[00152] The determining the correlation status can comprise using the computational model to identify one or more patterns associated with at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data using a computational model. The correlation status can comprise a presence of correlation or an absence of correlation. In one example, the machine learning classifier can generate output as 0 or 1, where 0 represents an absence of correlation and 1 represents a presence of correlation. The computer-implemented method can further comprise, if the correlation status comprises an absence of correlation, determining that the material attribute impacts neither clinical safety nor efficacy of the pharmaceutical material. The clinical safety can comprise a frequency of clinical adverse events. The efficacy of the pharmaceutical material can comprise a maximum response that can be achieved with the pharmaceutical material.
[00153] A computer-implemented method for training a computational model for determining a correlation status associated with a clinical impact of one or more material attributes can comprise (a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints; (b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material; (c) generating, via the one or more processors, predictive data associated with the pharmaceutical material based on the material attribute data and the clinical data using a predictive model; (d) applying, via the one or more processors, one or more transformations to the material attribute data, the clinical data, and the predictive data to create modified material attribute data, modified clinical data, and modified predictive data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and (e) training, via the one or more processors, the computational model using at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data. In some embodiments, the step (e) can comprise training, via the one or more processors, the computational model using the material attribute data, the clinical data, and the predictive data.
[00154] FIG. 3 is a diagram depicting an illustrative technique 300 for determining an impact of attributes, in accordance with some embodiments of the technology described herein. [00155] The attribute data silos 302 can comprise one or more attribute data storage units, such as attribute data storage units 302-1, 302-2, and 302-3. The attribute data silos 302 may be within the database 120 of FIG. 1. Each of the one or more attribute data storage units can store any type of material attribute data. In some embodiments, each of the attribute data storage units can store one or more categories of material attribute data. For example, the attribute data storage unit 302-1 can store data associated with one or more molecular attributes, the attribute data storage unit 302-2 can store data associated with a process related impurity, and the attribute data storage unit 302-3 can store data associated with one or more drug substance features. Deep learning stochastic models 308-1 and/or attribute modeling 308-2 can be used to obtain material attribute data over time for a pharmaceutical material. The output of the deep learning stochastic models 308-1 and/or attribute modeling 308-2 can be combined 314 with other data from the at least one of the data cleaning 306 and PK modeling 310 (or PD modeling).
[00156] Deep learning stochastic models 308-1, attribute modeling 308-2 and/or PK modeling 310 can comprise molecular modeling, degradation modeling, stability modeling, in-vivo mechanisms modeling (e.g., estimating aggregate formation or dissociation in-vivo), pharmacokinetic modeling, stochastic modeling of the drug molecule under various stressor environments. At least one of these models can be used to obtain material attribute data for lots that may not have associated direct measurements, and to model data and features beyond the time frame of measurement, and scope of measurement (e.g., to model in-vivo dynamics without directly measure). Additionally, at least one of deep learning stochastic models 308-1, attribute modeling 308-2, or PK modeling 310 can be applied to attribute data in combination with non-attribute related factors such as subject-level biophysical, demographic, spatiotemporal, climactic, and socioeconomic data.
[00157] The clinical data silos 304 can comprise one or more clinical data storage units, such as clinical data storage units 304-1, 304-2, and 304-3. The clinical data silos 302 may be within the database 120 of FIG. 1. Each of the one or more clinical data storage units can store any type of clinical data. In some embodiments, each of the clinical data storage units can store subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material. For example, the attribute data storage unit 304-1 can store data associated with a first adverse event
associated with one or more subjects that have received an administration of the pharmaceutical material, the attribute data storage unit 304-2 can store data associated with a second adverse event associated with one or more subjects that have received an administration of the pharmaceutical material, and the attribute data storage unit 304-3 can store data associated with a third adverse event associated with one or more subjects that have received an administration of the pharmaceutical material.
[00158] Pharmacokinetic (PK) modeling 310 can be used to obtain concentration over time for a pharmaceutical material. PK modeling 310 can be used to estimate clinical data of pharmaceutical material. By associating material attribute data, levels found in a subject at any given time t as some function of a PK model can be used to model the pharmaceutical material concentration and metabolism over time based on various factors (e.g., as determined by a PK data source (e.g., literature)). In one example, a simulated system can use an exponential decay model to estimate the rate of clearance of the high mannose (HM) form of mAbs over time. In this case, the model can use the rate of clearance of the HM form, the half-life of the molecules, and the level of HM estimate the rate of clearance of the high mannose (HM) form of mAbs over time. The output of the PK modeling 310 can be combined 314 with other data from the at least one of the data cleaning 306, deep learning stochastic models 308-1, or attribute modeling 308-2.
[00159] Data cleaning 306 can clean the data from attribute data silos 302 and/or clinical data silos 304 automatically. Details of data cleaning are described elsewhere herein. In some embodiments, the cleaned data can be stored in data storage 332 as prior knowledge. The cleaned data can be combined 314 with data from at least one of the deep learning stochastic models 308-1, the attribute modeling 308-2, or PK modeling 310. The combined dataset 314 can be cleaned 316 and linearized 318. Linearization 318 can comprise linearize data based on time. Details of the data cleaning and linearization are described elsewhere herein. The linearized/combined data can be divided into one or more data frames or data formats (e.g., subgroups), including, but not limited to, data frame associated with subject 320, data frame associated with adverse events 322, and data frame associated with arbitrary feature 324. Each of the data frames can be divided into one or more subsets 326. For instance, the data frame associated with subject 320 can be divided into subset 326-1 and subset 326-2, the data frame associated with adverse event 320 can be divided into subset 326-3 and subset 326-4, and the data frame associated with arbitrary feature can be subset 326-5.
[00160] Data frame associated with subject 320 can comprise any data structure associated with a healthy subject or treatment subject enrolled into a clinical trial. Data frame associated with adverse events 322 can comprise data structure associated with a type of categorical identifier that identifies one or more adverse events, or emergent combination of conditions and other measured or observed parameters associated with the adverse events. Data frame associated with arbitrary feature 324 can include data structure associated with categorical identifier or features that can be used to identify some silo of data consistent across the data sets being analyzed (e.g., subject age, which can consistently present in all relevant component datasets being used generate a combined dataset).
[00161] The one or more subsets 326 can be input into one or more analytical pipelines 328, including, but not limited to statistical models 328-1, dimensionality reduction 328-2, machine learning models 328-3, stochastic modeling 328-4, and deep learning models 328-5. The dimensionality reduction 328-2 and machine learning models 328-3 can comprise using methods, such as dimensionality correlation analysis, to understand molecule-level associations using molecular and structural data associated with pharmaceutical materials. The dimensionality reduction 328-2 and machine learning models 328-3 can comprise using methods, such as dimensionality correlation analysis, to understand subject-level associations using subject-level data including subject-level data as well as subject-level attribute-bridged data. Subject-level data can comprise sex, gender, age, weight, race, life choices (e.g., smoking/nonsmoking), subject-level biophysical data, demographic data, spatiotemporal data, and socioeconomic data. The dimensionality reduction 328-2 and machine learning models 328-3 can comprise at least one of logistic regression, k-nearest neighbors, decision trees, support vector machine, or naive Bayes. The deep learning models 328-5 can comprise understanding interdependencies between input variables (e.g., attributes, molecular characteristics, and adverse events) so that scalar clinical adverse event responses to dose can be modeled and predicted for doses above some predetermined level (e.g., doses that are outside of the doses range in a clinical trial). The stochastic modeling 328-4 can comprise generating model to predict next clinical event or probability distribution for next clinical event given sequential data (dose to adverse event or adverse event to adverse event). Predictions from these analytical pipelines 328 can comprise combinations of precondition and predicted outcome such as given a dose of attribute or attributes, predicting the likelihood of a subject with some set of parameters and adverse event histories developing some adverse event after dose exposure. Predictions from these analytical pipelines 328 can comprise
adverse event relationships such as predicting the likelihood of a new subject developing an adverse event based on prior history of adverse event and/or other characteristics such as available subject-level data. The stochastic modeling 328-4 can comprise transformer model and Bayesian model.
[00162] The output of the analytical pipelines 328 can be visualized by one or more heuristics, including, but not limited to, a vertical bar chart 330-1, a horizontal bar chart 330- 2, a radar chart 330-3, a pie chart 330-4, a line chart 330-5, a hierarchical chart 330-6, and a venn diagram 330-7. The output of the analytical pipelines 328 can be stored in data storage 332. The data storage 332 may be within the database 120 of FIG. 1.
[00163] EXAMPLES
[00164] A. Data Selection
[00165] Clinical study data was selected for CIA determination. Criteria for selecting clinical studies for CIA were that (1) a large variability in the attribute-exposure level was present, and (2) a large number of patients was enrolled. This way, statistically/mathematically meaningful results can be obtained from analyzing the clinical impact in the context of wide range of attribute-exposure levels. So, dose-escalation studies, or studies that used multiple treatment lots, were selected, and, if larger patient number was used, patients from comparable studies were combined for the analysis.
[00166] The drug-product batches administered in the clinical studies selected for CIA were traced using batch trace record to identify the drug-product lots inspected in product quality analytical tests, so the attribute level determined at the time of lot release specific to individual treatment lots could be obtained.
[00167] The tests for lot-release analyses were SEC (size-exclusion chromatography) for high molecular weight (HMW), rCE-SDS (capillary electrophoresis in reducing, and SDS- denaturing condition) for heavy chain (HC), light chain (LC), low molecular weight (LMW) species, mid-molecular weight (MMW) species, non-glycosylated heavy chain (NGHC) species, and CEX (cation-exchange chromatography) for acidic or basic species, individually developed specifically for mAb A or mAb B. Other analytical data from HIC (hydrophobicinteraction chromatography), Hydrophilic chromatography (HILIC) for high-mannose (HM), and peptide-mapping by LCMSMS (liquid chromatography -tandem mass spectrometry) for C-terminal sequence variant species (CSV1 and CSV2) were available for limited batches of mAb A or mAb B because they were not in the panel of lot-release assays.
[00168] The data from attribute stability studies were also mined for each pharmaceutical material. The stability data represented the change in the level of individual attributes over time in the storage condition, critical to determining the attribute-exposure level at the time of treatment administration as below. Of the patients from the studies selected as described above, the ones meeting the following criteria were included in CIA: (1) Batch information was available for all treatments in the regime; (2) the attribute exposure was determined for all treatments in the regime; (3) the ADA-test results were available; and (4) no ADA test conducted prior to administering the first treatment yielded positive results.
[00169] B. Data processing
[00170] The material attribute data and/or the clinical data associated with the pharmaceutical material used for modeling was extracted from the following one or more data sources. One example of the data sources was a data table shown in FIG. 8. As shown in FIG. 8, there were a subject ID column with associated pharmaceutical material exposures (as a series of exposures), a classification column with a distribution of clinical adverse events associated with each exposure, and one or more attribute columns with scalar values associated with the level of attribute dosed to patients based on the dose of pharmaceutical material and the lot of pharmaceutical material the subject was exposed to. A sequence of automated transformations, including data extraction, cleaning, merging, associating, selecting, or grouping steps were adopted to process the material attribute data and/or the clinical data associated with the pharmaceutical material. The values of material attribute data and/or the clinical data of prior knowledge pharmaceutical materials were also added to the original data set. The data transformation steps altered some numerical values of entries and the purpose of this procedure was to ensure that all related entries were grouped under their respective categories. Any altered value was stored as a derivative data structure and did not directly alter numerical values of raw input data. The resulting data set was used for subsequent model training and validation purposes.
[00171] C Model processing
[00172] For model training and tuning, k-fold cross-validation scheme was adopted to identify the hyperparameters for the different machine learning models. The data was split into k subgroups. One subgroup of k subgroups was taken as a hold out subgroup and the remaining k-1 subgroups were taken as training set. The model was trained on the remaining k-1 subgroups and evaluated on the one subgroup. In the initial stage of model development, one or more modeling modules (e.g., different analytical pipelines) and different machine
learning algorithms like Random Forest (RF), Support Vector Regression (SVR), Partial Least Squares (PLS), Multilayer Perceptron (MLP), and Neural Networks (NN) were built and tested on the training data set.
[00173] Statistical models for CIA tool, including T-test and Bayesian inference were used to examine distributions of exposure and clinical adverse event records and scalar values to test some hypothesis either externally imposed, such as for classical statistics, or internally imposed and self-updating, such as for Bayesian statistics. Data (e.g., material attribute data and clinical data) was cleaned by modular selection of core columns such as pharmaceutical material dose and subject ID, and optional columns such as demographic information consistent across the available data. The data was linearized then grouped based on patterns or factors including subject, date-time, gender, and age group. The data was grouped into data frames of the original whole dataset. Using programmatic logic (e.g., using some API framework such as SQL, Pandas), queries were imposed on these individual subframes, and the queries included "how many male subjects experienced an adverse event during treatment,” and “after what level pharmaceutical material dose was the first incidence of the adverse event observed?" Alternative queries/questions comprised features such as numerosity of adverse events a subject had experienced during treatment, severity of adverse events, class of adverse events, analysis of adverse events using groups of adverse events, and time periods between sequential adverse events. Alternative queries/questions also comprised examining co-presenting adverse events using the above-mentioned comparisons and performing the above-mentioned comparisons using arbitrary groupings of classifier-type data including demographic data, pharmaceutical material attributes, formulation attributes, hospital conditions, and treatment geographical locations.
[00174] FIG. 5 demonstrate an exemplary statistical T-test plot showing difference between average attribute exposure level between adverse event negative (AE-) and adverse event positive (AE+) subpopulations of clinical trial subjects. The p-value shown in FIG. 5 contextualizes the difference in averages (middle dashed lines) as meaningful (e.g., conveying some statistically significant information based on some discriminant and/or some confidence interval) or not meaningful.
[00175] CIA data (e.g., the material attribute data and the clinical data) was assessed using dimensionality manipulation machine learning according to following use cases and examples. Dimensionality manipulation machine learning, such as linear discriminant analysis (LDA) and principal component analysis (PCA), feature expansion techniques,
spectral clustering, involved linearization of data based on categories such as subject ID, date-time, and dose. The dimensionality manipulation machine learning included any relevant categorical and scalar columns in the dataset that were consistent across the input datasets. Dimensionality reduction methods compared co-variance between the facets of the dataset as represented by the columns of classifier or scalar data provided to the algorithm to provide an understanding of the co-dependencies between various facets of the complex multi-faceted CIA data. LDA was used to determine co-clustering attributes, adverse events, demographic details, relative to some query variable such as some attribute, adverse event, or some demographic detail. Dimensionality manipulation methods was further harnessed to pipeline into other modular pipelines such as a statistical models pipeline 328-1, deep learning models pipeline 328-5, as shown in FIG. 3, by performing unbiased feature classification and labeling to augment the other analytical pipelines. Dimensionality manipulation/scaling algorithms were used to analyze adverse event relationship clustering as an indication of potential relatedness.
[00176] FIG. 6 depicts an exemplary dimensionality analysis of PCA. In this plot, an attribute was compared to several adverse events. The offset of the attribute point from center, and the co-clustering of adverse events at a planar position distant to the attribute point indicated that the adverse event levels were not likely to be correlated to the attribute level. FIG. 7 depicts another exemplary dimensionality analysis example of PCA. In this plot an adverse event was compared to several attributes. The offset of the adverse event point from center, and the coclustering of attribute points at a planar position distant to the adverse events points indicated that the attribute levels are not likely to be correlated to the adverse event. FIG. 11 depicts another exemplary dimensionality analysis example of 2D-LDA. In this plot different adverse events were compared. In FIG. 11, for the classes of attributes in the overall dataset, adverse event 3 behaves differently from the adverse events 1, 2, 4, and 5. In FIG. 11, the adverse event 3 has a different attribute-dependency compared to the other adverse events, while adverse events 1, 2, 4, and 5 behave similarly relative to the attribute levels but have a different scale of dependency on the attribute levels.
[00177] A neural net (NN) was utilized to determine an impact of attributes. Given a new subject with the given attribute levels, the NN produced an estimate of a subject developing clinical adverse events. The NN was modified to perform different tasks. For instance, given an adverse event, the NN was used to predict the likelihood for a subject to develop a related adverse event as a function of some attribute dose. The NN was also used to generate
predicted values for adverse events for combinations of input pharmaceutical material molecules with or across molecular families. Additionally, given some adverse event, the NN was used to find predicted values for attributes across a CIA project up to or across CIA projects ever performed.
[00178] FIG. 9 depicts two performance metrics of a trained neural network (NN). Plot 902 depicts a confusion matrix related to a trained NN, and plot 904 depicts a classification report showing the comparison between predicted values and true values related to a trained NN. The confusion matrix 902 shows predicted classes and true classes. Each number in the matrix 902 represents the number of predictions that fall in that row x column bin. The principal diagonal (top left to bottom right) in plot 902 shows each pairing of correct prediction and correct classification. The confusion matrix 902 shows that, for a first pass training, the neural network was capable of picking up co-variant trends in the data and was able to correctly classify the majority of new data given to it. The classification report 904 shows a set of values calculated using the numbers in the confusion matrix and correspond to a set of numerical quality descriptors such as accuracy, precision, sensitivity, Fl score, specificity, and false positive rate.
[00179] Histograms were generated to contextualize the distribution of variables in the material attribute data and clinical data. FIG. 4 shows adverse events associated with the material attribute data and clinical data as grouped by total instance of each adverse event to subjects. These adverse events included headache, decreased appetite, injection site reaction, pruritus, and back ache. The histogram was re-formulated by total incidences of each adverse event.
[00180] An illustrative implementation of a computer system 1200 that may be used in connection with any of the embodiments of the technology described herein (e.g., such as the methods of FIG. 2 and FIG. 10) is shown in FIG. 12. The computer system 1200 includes one or more processors 1210 and one or more articles of manufacture that comprise non- transitory computer-readable storage media (e.g., memory 1220 and one or more non-volatile storage media 1230). The processor 1210 may control writing data to and reading data from the memory 1220 and the non-volatile storage device 1230 in any suitable manner, as the aspects of the technology described herein are not limited to any particular techniques for writing or reading data. To perform any of the functionality described herein, the processor 1210 may execute one or more processor-executable instructions stored in one or more non- transitory computer-readable storage media (e.g., the memory 1220), which may serve as
non-transitory computer-readable storage media storing processor-executable instructions for execution by the processor 1210.
[00181] Computer device 1200 may also include a network input/output (I/O) interface 1240 via which the computing device may communicate with other computing devices (e.g., over a network), and may also include one or more user I/O interfaces 1250, via which the computing device may provide output to and receive input from a user. The user I/O interfaces may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or touch screen), speakers, a camera, and/or various other types of I/O devices.
[00182] The above-described embodiments can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided in a single computing device or distributed among multiple computing devices. It should be appreciated that any component or collection of components that perform the functions described above can be generically considered as one or more controllers that control the above-described functions. The one or more controllers can be implemented in numerous ways, such as with dedicated hardware, or with general purpose hardware (e.g., one or more processors) that is programmed using microcode or software to perform the functions recited above.
[00183] In this respect, it should be appreciated that one implementation of the embodiments described herein comprises at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the above-described functions of one or more embodiments. The computer-readable medium may be transportable such that the program stored thereon can be loaded onto any computing device to implement aspects of the techniques described herein. In addition, it should be appreciated that the reference to a computer program which, when executed, performs any of the above-described functions, is not limited to an application program running on a host computer. Rather, the terms computer program and software are used herein in a generic sense to reference any type
of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors to implement aspects of the techniques described herein.
[00184] The foregoing description of implementations provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above teachings or may be acquired from practice of the implementations. In other implementations the methods depicted in these figures may include fewer operations, different operations, differently ordered operations, and/or additional operations. Further, non-dependent blocks may be performed in parallel.
[00185] It will be apparent that example aspects, as described above, may be implemented in many different forms of software, firmware, and hardware in the implementations illustrated in the figures. Further, certain portions of the implementations may be implemented as a “module” that performs one or more functions. This module may include hardware, such as a processor, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA), or a combination of hardware and software.
[00186] Having thus described several aspects and embodiments of the technology set forth in the disclosure, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those of ordinary skill in the art will readily envision a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein, and each of such variations and/or modifications is deemed to be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described. In addition, any combination of two or more features, systems, articles, materials, kits, and/or methods described herein, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.
[00187] The above-described embodiments can be implemented in any of numerous ways. One or more aspects and embodiments of the present disclosure involving the performance of processes or methods may utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to perform, or control performance of, the processes or methods. In this respect, various inventive concepts may be embodied as a computer readable storage medium (or multiple computer readable storage media) (e.g., a computer memory, one or more floppy discs, compact discs, optical discs, magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement one or more of the various embodiments described above. The computer readable medium or media can be transportable, such that the program or programs stored thereon can be loaded onto one or more different computers or other processors to implement various ones of the aspects described above. In some embodiments, computer readable media may be non-transitory media.
[00188] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects as described above. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor but may be distributed in a modular fashion among a number of different computers or processors to implement various aspects of the present disclosure. [00189] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[00190] Also, data structures may be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that convey relationship between the fields. However, any suitable mechanism may be used to establish a
relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish relationship between data elements.
[00191] When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.
[00192] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats.
[00193] Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.
[00194] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[00195] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and/or ordinary meanings of the defined terms.
[00196] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.” [00197] The phrase “and/or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and/or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements
specifically identified by the “and/or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and/or B,” when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[00198] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and/or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[00199] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
[00200] The terms “approximately,” “substantially,” and “about” may be used to mean within ±20% of a target value in some embodiments, within ±10% of a target value in some embodiments, within ±5% of a target value in some embodiments, within ±2% of a target value in some embodiments. The terms “approximately,” “substantially,” and “about” may include the target value.
Claims
1. A computer-implemented method for determining a correlation status associated with a clinical impact of one or more material attributes, the method comprising:
(a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints;
(b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material;
(c) applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and
(d) determining, via the one or more processors, the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
2. The computer-implemented method of claim 1, wherein the one or more material attributes comprise at least one of a molecular attribute, a process related impurity, or a drug substance feature.
3. The computer-implemented method of claim 2, wherein the one or more material attributes comprise a molecular attribute, and wherein the molecular attribute comprises at least one of: acidic species, basic species, high molecular weight species, subvisible particle number, visible particles, aggregation, low molecular weight, middle molecular weight, glycosylation, glycation, deamidation, deamination, cyclization, oxidation, sulfation, hydroxylysine, isomerization, fragmentation/clipping, N-terminal and C-terminal variants, signal peptide, reduced and partial species, misfolding, disulfide scrambling , domain swapping, folded structure, surface hydrophobicity, chemical modification, glycation,
covalent bond, mutations or misincorporations, a C-terminal amino acid motif PARG, a C- terminal amino acid motif PAR- Amide , drug antibody ratio (DAR), or peptide antibody ratio (PAR).
4. The computer-implemented method of claim 2, wherein the one or more material attributes comprise a process related impurity, and wherein the process related impurity comprises at least one of CHOP, HCP, residual host cell DNA, residual ProA, or a process reagent.
5. The computer-implemented method of claim 2, wherein the one or more material attributes comprise a drug substance feature, and wherein the drug substance feature comprises at least one of a component feature or a drug administration feature.
6. The computer-implemented method of any one of the preceding claims, wherein the pharmaceutical material comprises at least one of a biological therapy, a synthetic small molecule, or a nucleic acid.
7. The computer-implemented method of claim 6, wherein the pharmaceutical material comprises a biological therapy, and wherein the biological therapy is selected from the group consisting of: an antibody, an antigen-binding antibody fragment, an antibody protein product, a Bi-specific T cell engager (BiTE®) molecule, a bispecific antibody, a trispecific antibody, an Fc fusion protein, a recombinant protein, a recombinant virus, a recombinant T cell, a synthetic peptide, and an active fragment of a recombinant protein.
8. The computer-implemented method of claim 6, wherein the pharmaceutical material comprises a nucleic acid, and wherein the nucleic acid comprises an siRNA, an mRNA or a DNA.
9. The computer-implemented method of claim 6, wherein the pharmaceutical material comprises a biological therapy, and wherein manufacturing the pharmaceutical material comprises culturing a genetically engineered mammalian host cell comprising one or more nucleic acids encoding the biological therapy.
10. The computer-implemented method of any one of the preceding claims, wherein the pharmaceutical material is in a pharmaceutically acceptable formulation.
11. The computer-implemented method of any one of the preceding claims, wherein the measurement data of the one or more material attributes at the one or more timepoints is determined by at least one of mass spectrometry, chromatography, electrophoresis, spectroscopy, light obscuration, a particle method, analytical centrifugation, imaging or imaging characterization, or immunoassay.
12. The computer-implemented method of any one of the preceding claims, wherein the material attribute data comprises a change of the measurement data of the one or more material attributes at the one or more time points.
13. The computer-implemented method of any one of the preceding claims, wherein the material attribute data comprises a duration that the pharmaceutical material was under storage conditions prior to the administration of the pharmaceutical material.
14. The computer-implemented method of any one of the preceding claims, wherein the material attribute data comprises a dose of the pharmaceutical material in the administration.
15. The computer-implemented method of any one of the preceding claims, wherein the material attribute data comprises a level of material attribute exposure received by the one or more subjects at the time of the administration.
16. The computer-implemented method of any one of the preceding claims, wherein the one or more timepoints comprise at least one of a timepoint of manufacture or a timepoint of lot release.
17. The computer-implemented method of any one of the preceding claims, wherein the measurement data of the one or more material attributes is detected at two or more timepoints under storage conditions.
18. The computer-implemented method of any one of the preceding claims, wherein the one or more timepoints comprise a timepoint of manufacture and at least two subsequent timepoints.
19. The computer-implemented method of any one of the preceding claims, wherein the one or more clinical events comprise one or more clinical adverse events associated with the one or more subjects that have received the administration of the pharmaceutical material.
20. The computer-implemented method of any one of the preceding claims, wherein the subject data comprises at least one of preexisting condition, biomarker, laboratory result, metabolomic data, or demographic information of the one or more subjects that have received the administration of the pharmaceutical material.
21. The computer-implemented method of any one of the preceding claims, wherein the one or more transformations comprise cleaning the material attribute data or the clinical data, and wherein the cleaning comprises filtering the material attribute data or the clinical data based on one or more filtering criteria.
22. The computer-implemented method of claim 21, wherein the one or more filtering criteria comprise removing one or more orphan adverse events.
23. The computer-implemented method of any one of the preceding claims, wherein the one or more transformations comprise merging the material attribute data and the clinical data, and wherein the merging comprises combining the material attribute data and the clinical data based on a level of similarity between the material attribute data and the clinical data.
24. The computer-implemented method of any one of the preceding claims, wherein the one or more transformations comprise associating the material attribute data and the clinical data, and wherein the associating comprises correlating the material attribute data and clinical data based on one or more associating criteria.
25. The computer-implemented method of any one of the preceding claims, wherein the one or more transformations comprise selecting the material attribute data or the clinical data, and wherein the selecting comprises selecting the material attribute data or the clinical data based on one or more selection criteria.
26. The computer-implemented method of any one of the preceding claims, wherein the one or more transformations comprise grouping the material attribute data or the clinical data, and wherein the grouping comprises generating one or more subgroups based on one or more patterns of the material attribute data and the clinical data.
27. The computer-implemented method of any one of the preceding claims, wherein the computational model comprises at least one of: a logistic regression model, a support vector machine model, a multinomial logistic regression model, a multilayer perceptron model, a random forest model, a natural language processing model, a neural network model, a cluster model, a dimensionality reduction model, or a Markov model.
28. The computer-implemented method of any one of the preceding claims, wherein determining the correlation status comprises using the computational model to identify one or more patterns associated with the modified material attribute data and the modified clinical data.
29. The computer-implemented method of any one of the preceding claims, wherein the correlation status comprises a presence of correlation or an absence of correlation.
30. The computer-implemented method of claim 29, further comprising, if the correlation status comprises an absence of correlation, determining that the one or more material attributes impact neither clinical safety nor efficacy of the pharmaceutical material.
31. The computer-implemented method of claim 29, further comprising, if the correlation status comprises a presence of correlation, determining that the one or more material attributes impact at least one of safety or efficacy of the pharmaceutical material.
32. The computer-implemented method of claim 29, further comprising, if the correlation status comprises an absence of correlation, setting a specification for permissible levels of the one or more material attributes of the pharmaceutical material, wherein the permissible levels of the one or more material attributes are based on one or more levels of material attribute exposure received by the subjects.
33. The computer-implemented method of claim 29, further comprising, if the correlation status comprises a presence of correlation, setting a specification for maximum permissible levels of the one or more material attributes of the pharmaceutical material, wherein said maximum permissible levels of the one or more material attributes are based on one or more levels of the one or more material attributes associated with at least one of clinical adverse events or inhibition of efficacy of the pharmaceutical material.
34. The computer-implemented method of claim 29, further comprising, if the correlation status comprises an absence of correlation, manufacturing production lots of the pharmaceutical material comprising the one or more material attributes at or below a specified permissible level of the one or more material attributes based on one or more levels of material attribute exposure.
35. The computer-implemented method of claim 29, further comprising, if the correlation status comprises a presence of correlation, setting a specification for levels of the one or more material attributes at a time of manufacture that do not exceed a maximum permissible level of the one or more material attributes of the pharmaceutical material.
36. The computer-implemented method of claim 29, further comprising, if the correlation status comprises an absence of correlation, establishing a manufacturing process to produce levels of the one or more material attributes at or below a permissible level based on the correlation status.
37. The computer-implemented method of any one of the preceding claims, further comprising generating a rank of the one or more material attributes based on the correlation status using the computational model.
38. The computer-implemented method of claim 37, further comprising selecting a subset of the one or more material attributes based on the rank of the one or more material attributes and setting a specification for permissible levels of the subset of the one or more material attributes.
39. The computer-implemented method of any one of the preceding claims, further comprising generating one or more heuristics associated with the one or more subjects based on the modified material attribute data and the modified clinical data.
40. The computer-implemented method of any one of the preceding claims, further comprising estimating one or more parameters associated with the administration of the pharmaceutical material, wherein the one or more parameters comprise at least one of a stability expanded time-course dosing, a pharmacokinetic expanded time-course dosing, a step dosing, or a multiple dose overlap time-course dosing.
41. The computer-implemented method of any one of the preceding claims, further comprising estimating a cooperative effect of the one or more material attributes, wherein the cooperative effect comprises at least one of an additive effect, an inhibitory effect, or a feedback effect.
42. A computer-implemented method for training a computational model for determining a correlation status associated with a clinical impact of one or more material attributes, the method comprising:
(a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints;
(b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material;
(c) applying, via the one or more processors, one or more transformations to the material attribute data and the clinical data to create modified material attribute data and
modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and
(d) training, via the one or more processors, the computational model using the modified material attribute data and the modified clinical data.
43. A computer-implemented method for determining a correlation status associated with a clinical impact of one or more material attributes, the method comprising:
(a) obtaining, via one or more processors, material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints;
(b) obtaining, via the one or more processors, clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material;
(c) generating, via the one or more processors, predictive data associated with the pharmaceutical material based on the material attribute data and the clinical data using a predictive model;
(d) applying, via the one or more processors, one or more transformations to the material attribute data, the clinical data, and the predictive data to create modified material attribute data, modified clinical data, and modified predictive data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and
(e) determining, via the one or more processors, the correlation status based on at least one of the material attribute data, the clinical data, the predictive data, the modified material attribute data, the modified clinical data, or the modified predictive data using a computational model.
44. A computer system for determining a correlation status associated with a clinical impact of one or more material attributes, comprising: a memory storing instructions; and one or more processors configured to execute the instructions to perform operations including:
-n-
(a) obtaining material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints;
(b) obtaining clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material;
(c) applying one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and
(d) determining the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
45. A non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions for performing a method of determining a correlation status associated with a clinical impact of one or more material attributes, the method comprising:
(a) obtaining material attribute data of the one or more material attributes associated with a pharmaceutical material, wherein the material attribute data comprises measurement data of the one or more material attributes at one or more timepoints;
(b) obtaining clinical data associated with the pharmaceutical material, wherein the clinical data comprises subject data including one or more clinical events associated with one or more subjects that have received an administration of the pharmaceutical material;
(c) applying one or more transformations to the material attribute data and the clinical data to create modified material attribute data and modified clinical data, wherein the one or more transformations comprise at least one of cleaning, merging, associating, selecting, or grouping; and
(d) determining the correlation status based on the modified material attribute data and the modified clinical data using a computational model.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363522493P | 2023-06-22 | 2023-06-22 | |
| PCT/US2024/035108 WO2024263980A1 (en) | 2023-06-22 | 2024-06-21 | Methods and systems for determining an impact of attributes |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4732294A1 true EP4732294A1 (en) | 2026-04-29 |
Family
ID=91950088
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24743145.5A Pending EP4732294A1 (en) | 2023-06-22 | 2024-06-21 | Methods and systems for determining an impact of attributes |
Country Status (6)
| Country | Link |
|---|---|
| EP (1) | EP4732294A1 (en) |
| KR (1) | KR20260026492A (en) |
| CN (1) | CN121359213A (en) |
| AU (1) | AU2024312108A1 (en) |
| MX (1) | MX2025015279A (en) |
| WO (1) | WO2024263980A1 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20230121850A (en) * | 2020-12-16 | 2023-08-21 | 암젠 인크 | Methods of making biological therapies |
-
2024
- 2024-06-21 KR KR1020257042151A patent/KR20260026492A/en active Pending
- 2024-06-21 WO PCT/US2024/035108 patent/WO2024263980A1/en not_active Ceased
- 2024-06-21 CN CN202480040823.9A patent/CN121359213A/en active Pending
- 2024-06-21 AU AU2024312108A patent/AU2024312108A1/en active Pending
- 2024-06-21 EP EP24743145.5A patent/EP4732294A1/en active Pending
-
2025
- 2025-12-16 MX MX2025015279A patent/MX2025015279A/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024263980A1 (en) | 2024-12-26 |
| AU2024312108A1 (en) | 2025-12-04 |
| KR20260026492A (en) | 2026-02-26 |
| MX2025015279A (en) | 2026-02-03 |
| CN121359213A (en) | 2026-01-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12372470B2 (en) | Use of raman spectroscopy to monitor culture medium | |
| AU2018309027B2 (en) | Systems and methods for real time preparation of a polypeptide sample for analysis with mass spectrometry | |
| US20200392447A1 (en) | Cross-scale modeling of bioreactor cultures using raman spectroscopy | |
| US20240197922A1 (en) | Methods of Manufacturing Biological Therapies | |
| US20180291329A1 (en) | Cell culture methods and systems | |
| US11697670B2 (en) | Methods for purifying antibodies having reduced high molecular weight aggregates | |
| EP4732294A1 (en) | Methods and systems for determining an impact of attributes | |
| KR20210007958A (en) | Systems and methods for quantification and modification of protein viscosity | |
| CN116615232A (en) | Methods of making biotherapeutics | |
| WO2024263774A1 (en) | Product quality attribute target ranges | |
| CA3071676C (en) | Systems and methods for real time preparation of a polypeptide sample for analysis with mass spectrometry | |
| CN121311767A (en) | Near real-time sialic acid quantification of glycoproteins |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20260121 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |