WO2024256360A1 - Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis - Google Patents
Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis Download PDFInfo
- Publication number
- WO2024256360A1 WO2024256360A1 PCT/EP2024/066006 EP2024066006W WO2024256360A1 WO 2024256360 A1 WO2024256360 A1 WO 2024256360A1 EP 2024066006 W EP2024066006 W EP 2024066006W WO 2024256360 A1 WO2024256360 A1 WO 2024256360A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- experimental
- ids
- metrics
- metric
- aggregate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/10—Analysis or design of chemical reactions, syntheses or processes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/04—Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/06—Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/10—Office automation; Time management
- G06Q10/103—Workflow collaboration or project management
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q10/00—Administration; Management
- G06Q10/30—Administration of product recycling or disposal
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q50/00—Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
- G06Q50/04—Manufacturing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q2220/00—Business processing using cryptography
Definitions
- the present disclosure relates to chemical synthesis. More particularly, the present disclosure relates to a method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis.
- PMI Process Mass Intensity
- the PMI can be used as an indicator of both the cost-effectiveness and environmental compatibility of a process.
- the PMI provides information on how efficiently the process mass is used compared to the mass of the target material or product.
- the process mass refers to all chemicals, organic solvents, water, auxiliaries such as catalysts or pH buffers, rinsing media, etc. used in the production process. This mass-based resource consumption is measured in relation to the mass of a produced chemical compound (usually kg per kg).
- a low (small) PMI value indicates that the process is comparatively efficient and that the process mass is used optimally, while a high (large) PMI indicates that the process is inefficient and the process mass is not being used optimally, which possibly leads to a larger amount of waste and thus a high environmental impact.
- Yet another use of PMI for the evaluation of chemical processes was pioneered by Jimenez-Gonzales et al. (Org. Process Res. Dev. 2011, 15, 4, 912- 917), who showed a direct correlation between the PMI and the potential amount of eCO2 as a measure of global warming potential (GWP).
- SI Solvent Intensity
- WI Water Intensity
- DOZN TM Tool An online application that allows for a semi-quantitative evaluation of a single process step based on the 12 principles of green chemistry. It is not suitable for complex syntheses, like linear or branched multi-step syntheses, and requires detailed manual inputs. Mass-based resource consumption and Global Warming potential cannot be calculated. [0012] American Chemical Society (ACS) Green Chemistry Institute, Process Mass Intensity Calculator: An Excel -based tool used for single or multi-step reaction PMI calculation by manually entering reactants, reagents, solvents and aqueous systems, as well as the interdependencies of the above steps.
- ACS Chemical Society
- Process Mass Intensity Calculator An Excel -based tool used for single or multi-step reaction PMI calculation by manually entering reactants, reagents, solvents and aqueous systems, as well as the interdependencies of the above steps.
- Chem Pager / Roche A module connected to an ERP system that allows PMI calculation for established production processes. It is unsuitable for process development as the data source is limited to ERP system recipes.
- LCAs Life Cycle Assessments
- PCFs Product Carbon Footprints
- the present disclosure relates to a computer-implemented method for calculating mass flow and estimating sustainability metrics in chemical processes.
- the method may be initiated by selecting one or more experimental IDs from an electronic laboratory notebook (ELN), where each experimental ID corresponds to a subprocess with input and output materials.
- ESN electronic laboratory notebook
- the selection process may be performed through a Graphical User Interface.
- Experimental data corresponding to one or more previously performed experiments may be associated with the selected experimental IDs.
- Further steps of the method may include linking the selected experimental IDs to form an integrated process, which synthesizes a target material.
- the integrated process can include multiple corresponding subprocesses, and a synthesis tree corresponding to the one or more experimental IDs is provided by the software to support the process.
- the synthesis is adjustable by the user in some embodiments.
- the method estimates a plurality of process metrics, such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or environmental metric, wherein each of the plurality of process metrics corresponds to the respective experimental IDs of the one or more experimental IDs.
- the process metrics may also be a ratio, such as a PMI over Yield ratio.
- the process metrics may be provided for the whole process, parts of the process, or each of the subprocess steps.
- the Process metrics may also be compared to one or more process metrics from other processes.
- the process metrics can also be intuitively visualized by graphs, like a Sankey-Diagram.
- An aggregate process metric may be calculated, which is a function of the estimated plurality of process metrics.
- the aggregate process metrics may be, for example, a cost, a yield, an environmental metric, or the aggregation of any process metrics.
- the aggregate process metric may be a summation of the plurality of process metrics.
- the metrics may be on a unit basis, a mass basis, a volume basis, a batch basis, or in any unit known to one of ordinary skill in the relevant art.
- the method may include searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each of which corresponds to one or more respective experiments of the one or more experimental IDs.
- the search may look for reactions with the same input and output molecules and, in some embodiment, may ignore waste or support molecules such as catalysts, for example.
- the search may also look for recycled mass streams, that means, waste streams that are recycled to the beginning of a subprocess or a process. Based on this search, at least one alternative experimental ID may be selected for replacing one or more of the selected experimental IDs.
- a prophetic experiment can also be determined by querying a prediction engine and replacing at least one of the one or more experimental IDs with the prophetic experiment.
- the method may then estimate a prophetic process metric corresponding to the prophetic experiment and update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
- the prophetic process metric may include a range of values.
- the range of values may be a confidence interval, a confidence region, a credible interval, and/or a credible region. Those cited ranges of values may also include multi-step processes.
- the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
- the act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material.
- the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
- this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs.
- the plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
- Some embodiments include a data processing system comprising means for the implementation of any of the acts described above.
- the disclosure also includes a computer program and a computer-readable medium that can execute and perform the steps described in any one of the acts described above.
- some embodiments provide an intelligent and efficient computer- implemented method for calculating mass flow and estimating sustainability metrics in chemical processes.
- the method employs experimental data to select and link experimental IDs to form an integrated process, estimates process metrics, and calculates an aggregate process metric.
- Some embodiments allow for the selection of alternative experimental IDs, the adjustment of a process parameter, and querying of a prediction engine to identify a prophetic experiment.
- the method involves selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material.
- the method then links the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses, wherein the integrated process synthesizes a target material.
- the method further estimates a plurality of process metrics, wherein each of the process metrics corresponds to a respective experimental ID of the one or more experimental IDs.
- the method estimates an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics.
- the method also provides a synthesis tree corresponding to the one or more experimental IDs.
- the method searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, with each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs.
- the method selects an alternative experimental ID for at least one of the one or more experimental IDs and updates the aggregate process metric in accordance with the alternative experimental ID.
- the method may involve displaying the process metrics that are estimated.
- the plurality of process metrics estimated such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other environmental metrics, can be visually displayed through the graphical user interface.
- This display of the estimated process metrics can provide users with a convenient overview and analysis of the efficiency, sustainability, and potential environmental impact of the integrated process.
- the visualization of process metrics can aid in interpreting the estimation results, identifying areas for potential improvement, and guiding decisions during chemical process development.
- advanced data visualization tools may be leveraged to provide intuitive charts, graphs, Sankey diagrams or other graphical representations of the estimated process metrics.
- each experimental ID corresponds to a subprocess and includes experimental data from one or more previously performed experiments or historic experiments.
- the method is able to incorporate experimental data from previously conducted experiments into the analysis of the overall integrated process.
- This previously conducted experimental data provides the necessary information to estimate process metrics for the subprocesses associated with each experimental ID in the integrated process, as well as an aggregate process metric for the overall integrated process.
- the experimental data may encompass details regarding factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other parameters relevant to calculating process metrics like process mass intensity, yield, cost, and environmental impact indicators. Overall, the consideration of previously conducted experimental data lends greater accuracy and reliability to the method's ability to analyze the efficiency and sustainability of chemical synthesis processes.
- the method further comprises comparing the aggregate process metric with the alternative experimental ID to the aggregate process metric before updating the aggregate process metric based on the alternative experimental ID. More specifically, the aggregate process metric calculated using the initially selected experimental ID(s) is compared to what the aggregate process metric would be if a particular alternative experimental ID was used instead. This comparison step allows for assessment of the impact of selecting a different experimental ID on the overall aggregate process metric prior to updating the metric. By enabling this comparison, users can make informed decisions about whether to replace an initial experimental ID with an alternative option, based on how that replacement would affect key aggregate metrics like process mass intensity, cost, yield, etc. Only after comparing the aggregate metrics is the aggregate process formally updated to incorporate the selected alternative experimental ID in place of the initial experimental ID. This embodies an optimization approach that empowers users to explore multiple experimental options and select the one that optimizes the aggregate process according to user-defined criteria.
- the aggregate process metric with the alternative experimental ID may be compared to the aggregate process metric prior to updating the aggregate process metric. This allows for the comparison of the process metrics before and after the replacement of an experimental ID. The results of this comparison may additionally be displayed through the graphical user interface or other visualization methods. By showing the comparison, users can evaluate the impact of selecting a particular alternative experimental ID on the overall aggregate process metric.
- the aggregate process metric estimated by the process metrics estimator component may be a process mass intensity (PMI).
- the PMI provides a quantitative indicator of the efficiency of a chemical process by measuring the total mass of materials used in the process per unit mass of product generated. A lower PMI value generally indicates a more efficient process with less waste.
- the aggregate process metric calculated as a function of the individual process metrics for each experimental ID may be this PMI value, representing the overall mass efficiency of the integrated chemical process under analysis.
- PMI as the aggregate process metric enables effective assessment and comparison of the sustainability of alternative integrated processes in terms of their mass utilization.
- the system provides the flexibility to use PMI as the key aggregate metric for optimizing mass flow through the integrated synthesis process.
- the aggregate process metric estimated by the method may represent one of several sustainability indicators, including solvent intensity, water intensity, or global warming potential. More specifically, in some embodiments, the aggregate process metric calculated as a function of the plurality of estimated process metrics for the individual experimental IDs may be the overall solvent intensity, water intensity, or global warming potential for the integrated process.
- the aggregate process metric estimated for the integrated process can include cost metrics, yield estimates, and various environmental metrics.
- the aggregate process metric calculated by the process metrics estimator component may represent overall cost parameters such as total raw material costs, energy costs, and overall process economics. It may also incorporate chemical yield projections and optimizations based on the experimental data.
- the aggregate process metric can encompass composite environmental impact measures that account for factors like greenhouse gas emissions, wastewater generation, solid waste production, and other sustainability indicators. By consolidating different parameters into one overarching metric, the system aims to provide users with a high-level quantification of the performance, economics, and environmental profile associated with the integrated process. The flexibility to compute aggregate metrics spanning cost, yield, and sustainability domains allows for multi-objective analysis and aids in the identification of optimal process configurations.
- the method may involve selecting a role for a molecule that is associated with an experimental ID of the one or more experimental IDs. For example, a user may utilize the GUI to select a specific role, such as product, starting material, reagent, solvent, catalyst, etc., for a given molecule linked to an experimental ID.
- a specific role such as product, starting material, reagent, solvent, catalyst, etc.
- Associating molecules with particular roles can help provide additional context and clarity regarding how various compounds are being utilized within the integrated process. Defining these molecule roles enables the system to make appropriate assumptions and calculations when estimating the process metrics. The ability to select molecule roles provides an added level of customization, allowing users to tailor the analysis to their specific needs.
- a process parameter may be adjusted.
- This adjustment provides users with the capability to modify specific inputs or conditions associated with the integrated process in order to analyze the impacts on the overall process metrics. For example, users can tweak reaction temperatures, reagent quantities, solvent volumes, equipment settings, or other relevant process parameters.
- at least one of the plurality of estimated process metrics can then be updated accordingly. This update allows users to immediately see how modifications to the process parameters influence metrics like process mass intensity, yield, cost, global warming potential, etc.
- the method enables iterative optimization of the integrated chemical synthesis process. Users can repeatedly modify parameters and analyze the effects on mass flow, cost, sustainability indicators, and other output values of interest. Overall, the adjustment of process parameters combined with live updating of estimated metrics facilitates customizable scenario analysis and supports data- driven decision making during chemical process development.
- the method may involve adjusting a process parameter.
- Some of the experimental IDs may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based on the input parameters. Modifying the input parameters can update these output values.
- the method may further comprise updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
- the selection of the one or more experimental IDs from the electronic notebook is performed by a user utilizing a Graphical User Interface (GUI).
- GUI Graphical User Interface
- the GUI provides an interface which enables the user to browse, search, or otherwise access the experimental IDs stored within the electronic notebook.
- the user can then manually select the desired experimental IDs through interactions with GUI elements such as checkboxes, dropdown menus, or search filters.
- This manual selection initiates the process of linking together the selected experimental IDs into an integrated process for the synthesis of a target material.
- the GUI therefore facilitates convenient user control over the initial experimental ID selection, while the subsequent linking and analysis steps can be automated by the system based on this user input. Overall, this GUI-enabled selection process allows for customizable construction of integrated processes according to the specific needs and priorities of each user.
- the method involves querying a prediction engine to determine a prophetic experiment.
- a prophetic experiment refers to a hypothetical or simulated experiment that has not yet been physically performed.
- the prediction engine can utilize various techniques to generate predictions or suggestions for potential chemical reactions or processes to be explored.
- the method then replaces at least one of the initially selected one or more experimental IDs with the prophetic experiment identified by the prediction engine.
- the method estimates a prophetic process metric corresponding to the prophetic experiment.
- This prophetic process metric is then incorporated into the aggregate process metric calculation by updating the aggregate process metric to include the prophetic experiment in place of the replaced experimental ID(s).
- the prediction engine allows hypothetical experiments to be simulated within the overall integrated process, enabling users to explore different process variations and scenarios.
- the aggregate process metric reflecting the entire integrated process can then be re-estimated based on the incorporation of these prophetic experiments.
- the method may include a prophetic process metric corresponding to the prophetic experiment determined by a prediction engine.
- the prophetic process metric includes a range of values, such as a confidence interval, a confidence region, a credible interval, or a credible region. This range provides an indication of the uncertainty or variability associated with the predicted experiment. Accounting for inherent uncertainty when incorporating prophetic experiments can help assess the reliability and robustness of the predictions.
- the prophetic process metric estimated for the prophetic experiment suggested by the prediction engine may comprise a range of values indicating uncertainty.
- This range could take various standard statistical forms for representing uncertainty, including a confidence interval, a confidence region, a credible interval, or a credible region.
- Using these types of value ranges allows users to assess the reliability and variability inherent in the predicted experiments from the prediction engine when evaluating the potential impact on the overall process metrics.
- the prediction engine can quote a confidence interval, credible region, etc. to quantify the uncertainty in its estimations. This provides greater transparency into the accuracy of the predictive models used by the engine.
- the plurality of process metrics estimated may be represented as a ratio.
- the ratio may be process mass intensity (PMI) over yield.
- PMI process mass intensity
- Expressing the process metrics as a ratio can provide additional insights into the efficiency and sustainability of the integrated process by examining the relative values between different metrics. A lower ratio value may indicate improved performance, while a higher ratio value may highlight areas needing further optimization.
- Using customizable ratio metrics allows users to focus the analysis on parameters of greatest relevance to their specific needs and objectives. The ability to calculate and compare ratio-based process metrics adds an extra dimension of flexibility and customization to the overall method.
- the plurality of process metrics may also be represented as a ratio.
- this ratio is process mass intensity (PMI) over yield.
- PMI process mass intensity
- Using the ratio of PMI over yield as a process metric enables evaluating both the mass efficiency and productivity of the integrated process in a simple, intuitive metric.
- a lower PMVyield ratio indicates greater efficiency and yield for the overall chemical synthesis pathway. Tracking how modifications to the process impact both PMI and yield through their ratio can further optimization efforts by balancing mass utilization and target output.
- the use of PMI over yield as a ratio metric provides a sustainability indicator for the integrated process.
- the one or more experimental IDs may be selected out of order by the user.
- the act of linking the plurality of experimental IDs includes the act of ordering the selected experimental IDs by the system to properly form the complete integrated process. This integrated process includes all of the corresponding subprocesses associated with the selected experimental IDs, arranged in the correct sequence. Automatically ordering the experimental IDs ensures that the full sequence of subprocesses is configured in the proper order to synthesize the target material. By rearranging the selected experimental IDs into the right process sequence, the system can link subprocesses that may have been originally selected out of order and still generate an integrated process that connects all associated subprocesses to produce the desired target output.
- the method may involve subprocesses including chemical reactions, purifications, or a combination thereof. More specifically, in some embodiments each experimental ID selected from the electronic notebook corresponds to either a chemical reaction subprocess, a purification subprocess, or a subprocess involving both a chemical reaction and a purification. Therefore, when linking the experimental IDs to form the integrated process for synthesizing the target material, the resulting integrated process may feature multiple subprocesses including chemical reactions, purifications, or a combination of both chemical reactions and purifications. The ability to link together experimental IDs representing diverse types of chemical subprocesses provides flexibility in constructing integrated processes to synthesize desired target materials.
- the method may involve generating a report that summarizes the estimated process metrics and the aggregate process metric. More specifically, in some embodiments, the method includes an additional step of producing a report that provides an overview of the plurality of process metrics calculated for each experimental ID as well as the aggregate process metric estimated for the integrated process. This report condenses the key information and metrics into a concise summary that allows users to easily review the overall efficiency, sustainability, cost, yield, and environmental impact indicators for the chemical process. By gathering the estimated metrics into a single report, users can assess the process performance holistically rather than examining individual metrics in isolation. The report provides a tool for convenient evaluation and comparison of different process configurations or alternatives explored using the system.
- the report may present the metrics in various graphical visualizations, such as bar charts, line plots, or Sankey diagrams, to enable intuitive interpretation of the data.
- the summary report equips users with a quick yet comprehensive perspective of the integrated process and its characterized performance based on the computed process metrics.
- the method may involve verifying the compatibility of input and output materials when linking the one or more experimental IDs.
- This verification step acts to ensure continuity in the resulting integrated process that synthesizes the target material.
- the input materials required for each experimental ID or subprocess are checked against the output materials produced by any preceding experimental IDs or subprocesses in the integrated process flow. Any mismatches in output versus required input materials can indicate possible breaks in continuity of the integrated process, which can then be addressed.
- an embodiment aims to confirm that the linkage of experimental IDs successfully creates an end-to-end integrated process with no gaps that would prevent the target synthesis.
- the plurality of process metrics estimated by the process metrics estimator component may include energy consumption metrics.
- These energy consumption metrics can be calculated based on the input and output materials and the subprocesses utilized in each experimental ID of the integrated process. Specifically, factors such as the quantities and types of input materials, chemical transformations involved in the subprocesses, output materials produced, reaction conditions like temperature and pressure, and other relevant parameters can be used to estimate the energy consumption associated with each subprocess and experimental ID. Advanced techniques like computational fluid dynamics simulations, thermodynamic calculations, or empirical correlations based on historical data may be leveraged to quantify the energy consumption. By aggregating the energy consumption across all experimental IDs, an overall energy usage estimate can be obtained for the integrated process. This facilitates analysis of the environmental impact and cost-effectiveness of the overall synthesis route. The energy consumption metrics, along with the other estimated process metrics, allow users to make informed decisions when optimizing the integrated process.
- the method may include a step of validating the estimated process metrics against predetermined criteria before updating the aggregate process metric.
- the estimated process metrics such as process mass intensity, solvent intensity, water intensity, and other metrics
- the validation ensures accuracy and reliability of the metrics prior to using them to update the overall aggregate process metric. If the estimated process metrics for the individual experimental IDs meet the predetermined validation criteria, then they can be applied to update the combined aggregate metric for the full integrated process. However, if one or more process metrics fail to satisfy the validation checks, the method may require troubleshooting, metric re-estimation, or other corrective measures before the aggregate metric is updated.
- This validation step may act as a quality check on the estimated metrics, enhancing the reliability and robustness of the overall mass flow calculations and sustainability analysis enabled by the method.
- machine learning models are utilized for the searching for similar reaction subprocesses.
- the method involves searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, as previously described.
- the search can make use of machine learning algorithms and models that are trained on historical experimental data to identify patterns and similarities between experiments. These models enable the search to suggest relevant alternative experiments based on the provided experimental IDs.
- the utilization of machine learning models can expand the capabilities of the search and provide more comprehensive results, optimizing the process metrics and sustainability indicators.
- the method may include the step of alerting a user when the aggregate process metric exceeds a predetermined environmental impact threshold. More specifically, the aggregate process metric estimated as a function of the plurality of process metrics can be compared against a defined threshold representing the maximum allowable environmental impact. If the aggregate process metric is calculated to exceed this predetermined threshold, the system may generate and display an alert to notify the user.
- This alert feature enables users to recognize when certain sustainability or environmental impact constraints are violated due to excessive resource consumption or emissions associated with the integrated chemical synthesis process under analysis. The ability to define environmental impact limits and automatically trigger notifications when these constraints are surpassed can aid in adhering to sustainability guidelines and minimizing ecological footprints.
- each experimental ID may be associated with specific equipment used in the corresponding subprocess.
- a reaction subprocess may utilize a particular reactor or purification equipment.
- the method may further comprise adjusting the settings or parameters of this equipment based on the requirements of that subprocess. For instance, if a reaction subprocess operates at a certain temperature and pressure, the reactor settings can be automatically adjusted to match those conditions. This automatic adjustment of equipment settings streamlines the experimental workflow and ensures alignment between the subprocess details contained within the experimental ID and the actual equipment configuration used to carry out that subprocess. By linking the experimental IDs to the relevant equipment in this manner, the method enables improved standardization, repeatability, and optimization of the subprocesses.
- the method may further involve automatically ordering chemicals and materials needed for the subprocesses.
- This automatic ordering is based on the input materials listed in the experimental IDs that were selected from the electronic notebook.
- the system can determine the required reagents, solvents, catalysts, and other chemicals needed to carry out each subprocess. It can then automatically generate purchase orders or materials requests to obtain the necessary supplies, ensuring analysts and researchers have the required ingredients on hand before commencing the chemical reactions or purifications encompassed within that subprocess.
- This just-in-time materials ordering facilitated by the automated system helps minimize inventory and procurement overhead for organizations frequently synthesizing new target compounds or materials.
- the synthesis tree provided to the user may include multiple alternative synthesis pathways that could be used to synthesize the target material.
- the method may provide the capability for the user to select one of these alternative synthesis pathways based on user-defined criteria or preferences. For example, the criteria could be related to optimizing particular process metrics like cost, environmental impact, or yield.
- the user interface allows the user to input these criteria and priorities.
- the method then facilitates the selection of the optimal synthesis pathway that best meets the specified criteria out of the alternatives included in the synthesis tree visualization. Enabling users to explore alternative synthesis routes and choose based on customizable metrics provides flexibility and aids in process optimization.
- the method may further involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs.
- This energy consumption metric can provide an indication of the overall energy requirements associated with the synthesis of the target material via the linked subprocesses.
- empirical correlations, thermodynamic calculations, or other techniques may be applied using parameters from the experimental data corresponding to each experimental ID. Such parameters can include temperature, pressure, flow rates, batch sizes, and other relevant factors that influence energy usage.
- parameters can include temperature, pressure, flow rates, batch sizes, and other relevant factors that influence energy usage.
- the method may involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs.
- Estimating the energy consumption metric can comprise applying empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs. For example, thermodynamic calculations and empirically gained correlations may be used to estimate the energy uptake of individual process steps, such as heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction.
- This approach leverages established techniques to assess the energy requirements and associated environmental impacts of the subprocesses linked to form the integrated process.
- the method may involve estimating a carbon footprint metric for the integrated process.
- the carbon footprint metric may be estimated based on the one or more experimental IDs that were selected from the electronic notebook and linked to form the integrated process.
- the carbon footprint estimation may utilize life cycle inventory data for the various input materials involved in the subprocesses and integrated process.
- life cycle inventory data for the various input materials involved in the subprocesses and integrated process.
- the method further comprises providing a visualization of mass flow through the integrated process based on the one or more experimental IDs that were selected from the electronic notebook.
- This visualization may utilize graphical tools and diagrams that allow users to see how mass flows through each subprocess and the overall integrated process used to synthesize the target material. For example, Sankey diagrams or other types of flow charts could be generated to intuitively display quantities and mass balances across different stages of the chemical synthesis process defined by linking the experimental IDs. Enabling intuitive visualization of mass flow helps users identify areas of inefficient material usage and opportunities for reducing waste or environmental impact.
- the visualization component allows users to visually explore the mass flow impacts of selecting alternative experimental IDs or adjusting process parameters within the integrated process.
- the method includes a step of visualizing the mass flow through the integrated process based on the one or more experimental IDs.
- This visualization of the mass flow may employ a Sankey diagram, which allows an intuitive graphical representation of flows and their quantitative values within a system.
- the Sankey diagram can illustrate the sequential flow of materials, energy transfers, or waste through each subprocess and throughout the overall integrated process.
- the thickness of the arrows in the Sankey diagram represents the magnitude or amount of mass flow.
- This Sankey diagram visualization provides users with an easily interpretable overview of how mass flows through and is transformed within the integrated chemical synthesis process under analysis.
- the visualization supports identification of mass intensive steps, recycling opportunities, yield losses, and other relevant mass flow characteristics, aiding in process assessment, optimization, and improvement.
- the system may allow for the synthesis tree, which corresponds to the one or more experimental IDs, to be edited by the user through a graphical user interface. More specifically, in some embodiments the system includes functionality enabling user modification of the visualization of the overall synthesis route. This can facilitate optimization, process adjustments, or exploring alternative synthesis pathways.
- the act of searching for similar reaction subprocesses includes identifying subprocesses that have the same input and output molecular structures as a given experimental ID from among the one or more experimental IDs. Specifically, when searching for similar subprocesses, those processes that match the molecular structure of both the input material and the output material for a particular experimental ID can be recognized. By finding reaction subprocesses with identical input and output chemicals, alternative synthetic routes or optimizations for specific reaction steps within the integrated process may be determined. This approach focuses the search on processes that are highly comparable in terms of the chemical transformations being carried out.
- the method involves estimating a range of values for at least one of the plurality of process metrics based on the plurality of similar experiments identified in the search. More specifically, after searching for and determining similar reaction subprocesses and experiments for each experimental ID, a range of values may be calculated for process metrics like process mass intensity, solvent intensity, yield, cost, or other metrics. This range represents the variability in the metric across the multiple similar experiments found that correspond to the experimental ID(s) selected by the user. Providing such a range gives an indication of the uncertainty and potential fluctuation in the process metric, allowing for a more comprehensive analysis when evaluating and optimizing the integrated chemical synthesis process. The range could take the form of a confidence interval, confidence region, credible interval, credible region, or other representation of variability. This enhanced uncertainty quantification through metric value ranges enables more informed decision making during chemical process development.
- Some embodiments of the method may involve suggesting an alternative solvent or an alternative reagent for at least one of the experimental IDs based on the plurality of similar experiments found by the search component.
- the search component identifies experiments with similar reactions and can determine if there may be better solvents or reagents that could be used for an experimental ID by analyzing the reagents and solvents used in those similar experiments.
- the system can recommend testing alternate solvents or reagents for an experimental ID that may improve yield, lower cost, reduce waste, or provide other benefits over the original solvent or reagent selected. This allows researchers to easily get suggestions on alternate reaction conditions to try that could optimize their process.
- the method allows for the selection of one or more experimental IDs from an electronic notebook, wherein the selection is based on a target molecule specified by a user.
- a user may first specify or input a desired target molecule they wish to synthesize. Based on this target molecule, relevant experimental IDs can then be retrieved and selected from the electronic notebook that correspond to subprocesses involved in the production or synthesis of the specified target material.
- the target molecule provides a means for the user to indicate what final product they intend to make, while the system identifies and selects the necessary reaction steps and corresponding experimental IDs that can lead to the target molecule.
- This target molecule-based selection of experimental IDs facilitates the process of linking subprocesses into an overall integrated process for synthesizing the user- specified target material. Overall, enabling the selection of experimental IDs based on a user- defined target molecule allows the system to automatically identify the building blocks needed for a user's molecule of interest based on available experimental data.
- Some embodiments provide the capability of tracking modifications made to the integrated process over time, as well as the resulting impacts on the aggregate process metric. Specifically, as changes or optimizations are made to the linked experimental IDs that comprise the integrated process, the system can monitor and log these modifications. Concurrently, the impacts of the changes on process metrics like the overall process mass intensity, solvent intensity, yield, cost, or other aggregate sustainability indicators can be reestimated and recorded. By tracking edits to the process flow alongside corresponding shifts in the process metrics, users can evaluate how impactful or beneficial particular process tweaks have been historically. This logging functionality enables data-driven analysis of how the integrated process has evolved regarding sustainability and can guide future optimization efforts. Overall, the ability to trace modifications and quantify associated effects on the aggregate process metric can facilitate systematic improvements over multiple iterations.
- the method may further involve estimating the impacts of solvent recycling or waste stream recycling on the plurality of process metrics.
- the process metrics estimator component estimates how recycling solvents or waste streams back into the integrated process influences metrics such as process mass intensity, solvent intensity, water intensity, and global warming potential. For example, solvent recycling can reduce the amount of fresh solvent utilized in the process, thereby lowering the solvent intensity. Similarly, recycling waste streams may decrease the material inputs needed, which can positively impact sustainability indicators like process mass intensity. Quantifying these recycling impacts provides additional insights into optimization opportunities for improving the environmental performance of the integrated process.
- the method may involve integrating the estimated process metrics and aggregate process metric with an enterprise resource planning (ERP) system.
- ERP enterprise resource planning
- the ERP system can provide comprehensive data management capabilities to track materials, processes, inventory, orders, accounting, and other operational data across the enterprise. Integrating the sustainability metrics estimated by the system with the ERP allows for holistic tracking of environmental factors alongside traditional business metrics. This integration enables seamless monitoring and optimization of both process efficiency and environmental impact through a centralized platform. Overall, incorporating the calculated process metrics and aggregate sustainability indicators into an organization's broader ERP infrastructure can facilitate comprehensive data analysis to inform sustainable decision-making across chemical research and manufacturing activities.
- the method involves storing the one or more experimental IDs, the plurality of process metrics, and the aggregate process metric in a database.
- This database storage allows the experimental data, process metrics, and aggregate metrics to be saved for later retrieval and analysis. By storing this information, users can access a historical record of prior experiments, process metrics calculations, and overall sustainability indicators.
- the database also enables data sharing, collaboration between multiple users, integration with other systems, and long-term tracking of process improvements over time.
- the stored data may be retrieved at a later point from the database for additional analysis, comparison to alternative processes, or to demonstrate progress in process optimization efforts. Overall, the database storage provides persistence and easy access to valuable experimental records and sustainability metrics.
- the method may further involve retrieving the one or more experimental IDs that were stored in the database, along with the associated plurality of process metrics and the aggregate process metric. This allows a user to revisit and analyze previous experimental data that had been saved, enabling the comparison of different process options or tracking changes in process metrics over time.
- the storage and retrieval capabilities facilitate data management and reuse, avoiding duplication of effort and taking full advantage of experimental data from past synthesis routes or process configurations.
- users can conveniently restore experimental IDs and process metrics that pertain to earlier experiments or process alternatives.
- the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
- the act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material.
- the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
- this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs.
- the plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
- the method may further involve updating the plurality of process metrics and the aggregate process metric based on edited input data associated with at least one of the selected experimental IDs.
- the system enables users to manually edit the input data corresponding to the experimental IDs through the graphical user interface. For example, a user could modify the quantities of reagents used in a reaction or change the reaction conditions like temperature and pressure.
- the process metrics estimator component automatically recalculates and updates the plurality of process metrics to reflect the changes in input data. This would update metrics like process mass intensity, solvent intensity, yield percentage, etc. in accordance with the manual edits.
- the aggregate process metric is a function of the individual process metrics, it is also updated accordingly when changes are made to the input data. This dynamic update allows users to immediately see the impact of any input data modifications on the overall process metrics and sustainability indicators. It facilitates rapid scenario testing and process optimization during chemical process development.
- the method includes displaying changes between the edited input data that the user manually adjusted and the original input data that was initially associated with at least one experimental ID.
- Visually highlighting the changes between the original and edited input data or providing a side-by-side comparison can make it easier for users to see where modifications have been made and understand the impacts on the estimated process metrics and aggregate process metric. This capability enhances transparency and facilitates iterative optimization of the integrated process through assessment of different input data scenarios.
- the method includes enabling a user to manually edit input data associated with at least one of the one or more experimental IDs.
- the plurality of process metrics and aggregate process metrics are updated based on the edited input data.
- Changes between the edited input data and the original input data associated with the at least one experimental ID may be displayed. Displaying the changes may involve visually highlighting the differences between the edited and original input data within the user interface. Alternatively, displaying the changes may involve providing a side-by-side comparison of the original and edited input data, allowing the user to clearly see how the data has been modified.
- the plurality of estimated process metrics may be provided for the integrated process as a whole, for parts of the integrated process, or for each individual subprocess step corresponding to the respective experimental IDs. That is, the process metric estimates generated by the process metrics estimator component can be presented at different levels of granularity depending on the specific requirements. For example, a cumulative or aggregate PMI value may be calculated and displayed for the complete multi-step integrated process to synthesize the target material. Alternatively, separate PMI values may be estimated and shown for each distinct subprocess, giving insights into the PMI contributions of individual steps. As another option, PMI metrics may be calculated and visualized for logical subsections of the integrated process, such as a sequence of reactions or a particular purification train. This flexibility in process metric estimation and visualization at multiple levels enables detailed analysis and comparison of different process options.
- the process metrics estimated for the integrated process may be compared to one or more process metrics from other processes.
- this comparison to external process metrics can be performed in order to facilitate evaluation of the relative performance of the integrated process under consideration.
- the comparison can provide additional context and enable easier assessment of the efficiency, cost-effectiveness, sustainability, or other relevant parameters associated with the integrated process synthesized from the selected experimental IDs. The capability to draw such comparisons against suitable benchmarks expands the utility of the estimated process metrics and aggregate process metrics calculated by the system.
- the method may further involve providing the plurality of process metrics for each individual subprocess within the integrated process formed by linking the experimental IDs.
- the process metrics estimator component can estimate process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
- process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
- process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
- the method includes comparing at least one process metric of the plurality of process metrics to a corresponding process metric from a different integrated process. For example, a process mass intensity or solvent intensity value estimated for one of the linked experimental IDs may be compared to process mass intensity or solvent intensity values obtained from a separate, distinct integrated process. This enables assessing the relative efficiency or sustainability of the process under consideration compared to alternative processes for producing the same or similar target compounds.
- the ability to benchmark against historical data or industry standards supports evaluating opportunities for improving the integrated process through suitable modifications to reaction conditions, workup procedures, or selection of starting materials and reagents.
- At least one process metric of the plurality of process metrics may be visualized using a Sankey diagram.
- the Sankey diagram provides an intuitive way to represent the flow of mass, energy, cost, or other metrics through the integrated process. This visualization can help highlight inefficiencies, mass imbalances, and opportunities for process improvement in a graphical manner.
- users can gain insight into how changes in one subprocess may propagate through the integrated process to impact other metrics of interest.
- Some implementations may allow users to interact with the Sankey diagram representation of the process metrics, for example by selecting specific pathways or modifying process parameters to observe the effect on the overall mass flow in real time.
- the Sankey diagram enables an interactive, graphical analysis that complements the detailed numeric process metrics estimated by the system.
- the method includes identifying recycled mass streams that are recycled to the beginning of a subprocess or the integrated process.
- the act of searching for similar reaction subprocesses by the search component may involve recognizing mass streams from waste or byproducts that are recycled and fed back into an earlier subprocess or the start of the overall integrated process. By detecting these recycled mass streams, the system can analyze the impact of recycling on the estimated mass flow and process metrics. This recycling functionality provides another tool to explore optimization opportunities and improve the efficiency of the chemical synthesis process under analysis.
- the search component is configured to identify these recycled streams within the historical data and suggest integration opportunities accordingly.
- the range of values estimated for the prophetic process metric may correspond to a multi-step process rather than just a single reaction step.
- the predicted process metrics for this overall pathway may include confidence intervals or credible regions. These uncertainty ranges can account for the inherent variability when predicting outcomes across several chemical transformations or over an extended process with several distinct steps. By propagating uncertainties across the various stages, reliable estimates for the overall reliability of the predicted process metrics and sustainability indicators can be provided even for complex, integrated processes. This allows users to realistically assess the accuracy of recommendations from in silico tools when dealing with intricate, multi-stage synthetic routes.
- the method involves automatically linking the one or more experimental IDs to form the integrated process synthesizing the target material.
- This automatic linkage is performed based on the input and output chemical structures specified in each of the experimental IDs. That is, the software analyzes the chemical structures entering and exiting each subprocess to determine compatibility and continuity of materials flow. It then automatically connects the experimental IDs in the proper order to construct the overall integrated process including all necessary subprocesses, without requiring extensive manual input or oversight from the user.
- the linking process can be streamlined and expedited. This automation enables more rapid set up and estimation of mass flow and sustainability metrics for the chemical synthesis route.
- the plurality of process metrics are represented as a ratio.
- the ratio can be process mass intensity over yield.
- the process mass intensity reflects the mass-based resource consumption per unit mass of product, while the yield represents the amount of product obtained from the process.
- Taking the ratio of these two metrics provides a useful indicator of the efficiency and environmental impact of the integrated process, accounting for both the resource usage and productivity.
- a lower ratio indicates a more efficient and sustainable process.
- users can assess the impacts of modifications or alternate pathways on the overall process performance.
- representing the plurality of process metrics as a process mass intensity over yield ratio enables straightforward yet comprehensive evaluation of chemical synthesis processes.
- the system may include both hardware and software components configured to execute the computer-implemented method for calculating mass flow and estimating sustainability metrics.
- the system may include one or more processors, memory, storage, network interfaces, databases, and other computing resources capable of carrying out the various steps involved in selecting experimental IDs, linking them to form an integrated process, estimating process metrics, providing a synthesis tree, searching for similar reactions, selecting alternative experimental IDs, and updating aggregate metrics.
- the system is designed to implement the full functionality of the method efficiently through specialized algorithms, predictive models, data structures, or other technological means.
- the modular, customizable architecture of the system also allows for easy extension or updating of capabilities in accordance with advancements in the field. Overall, embodiments provide an intelligent data processing system leveraging automation and advanced analytics to revolutionize mass flow calculations, process optimization, and sustainability assessment during chemical process development.
- the disclosure includes a computer program containing instructions that, when executed by a computer, cause the computer to perform the method described herein.
- This computer program allows the computer to select one or more experimental IDs, link the IDs to form an integrated process, estimate process metrics for each ID, calculate an aggregate process metric, provide a synthesis tree, search for similar reactions, select alternative IDs, and update metrics accordingly.
- the computer can fully implement the method for calculating mass flow and estimating sustainability metrics in chemical processes.
- the computer program enables automated calculation of metrics, exploration of alternatives, and optimization of chemical synthesis procedures in a computerized environment.
- the method may be implemented as a computer program with instructions stored on a computer-readable medium.
- these instructions When these instructions are executed by a computer, they cause the computer to carry out the method comprising: selecting one or more experimental IDs from an electronic notebook, linking the IDs to form an integrated process synthesizing a target material, estimating process metrics for each ID, calculating an aggregate metric, providing a synthesis tree, searching for similar reactions, selecting alternative IDs, and updating the aggregate metric accordingly.
- the computer-readable medium allows the computational implementation of the method, enabling the computer to perform the selection, linking, estimation, calculation, provision, searching, alternative selection, and updating involved in analyzing mass flow and sustainability metrics for chemical processes.
- a computer can carry out the full method by accessing and executing the appropriate program from the computer-readable medium on which it resides.
- FIG. 1 shows a block diagram illustration of a cloud-based system to calculate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure
- FIG. 2 show a block diagram illustration of a computing device to estimate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure
- Fig. 3 is a flow-chart diagram of a method of estimating mass flow and estimate sustainability metrics in accordance with an embodiment of the present disclosure
- Fig. 4 shows a GUI of linked experimental IDs in accordance with an embodiment of the present disclosure
- FIG. 5 shows a summary of the output of the of linked experimental IDs of Fig. 4 in accordance with an embodiment of the present disclosure
- FIG. 6 shows a synthesis tree in accordance with an embodiment of the present disclosure
- Fig. 7 shows an example result of the search component, in which the input data can be modified in accordance with an embodiment of the present disclosure
- FIGs. 8A and 8B show an example result of the search component in which the mass flow is visualized in accordance with an embodiment of the present disclosure
- Fig. 9 shows an example result of the search component in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 6 in accordance with an embodiment of the present disclosure
- Fig. 10 shows PMI vs. Yield values for each alternative experimental ID that the search component identifies in accordance with an embodiment of the present disclosure
- Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow is estimated in accordance with an embodiment of the present disclosure
- Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure
- Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations in accordance with an embodiment of the present disclosure
- Fig. 14 shows model compounds for raw materials to synthesize APIs in accordance with an embodiment of the present disclosure
- Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure
- Fig. 16 shows an average PMI and chemical yield for similar reactions of one synthesis step that was suggested by a retro-synthesis tool using data is obtained using the similarity search function in accordance with an embodiment of the present disclosure
- Fig.17 shows the Input metric for the energy calculation module in accordance with an embodiment of the present disclosure
- Fig.18 shows the output metric for the energy calculation module in accordance with an embodiment of the present disclosure.
- Fig. 19 shows the GUI displaying weight based carbon footprints received from an external database, which are assigned based on data of individual raw materials, on a general role, or manually edited data in accordance with an embodiment of the present disclosure.
- Fig. 1 shows a block diagram illustration of a cloud-based system 100 to estimate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure.
- the system 100 automatically calculates metrics without significant manual interaction or input.
- the system 100 includes a cloud service provider 102, one or more personal computers 104, and a mobile device 106.
- the system 100 also includes a mass Flow-estimator component 112.
- the mass flow-estimator component 112 can estimate mass flow in chemical synthesis and enable a user to modify the chemical synthesis process to change or reduce the mass flow as described herein.
- the mass flow estimator component 112 can be used in chemical process development to optimize certain values, such as costs, yield or environmental compatibility.
- these may be calculated early in the development and updated regularly.
- These data may be stored in a data pool, e.g., in subprocess metrics 144 of a database 132.
- the database 132 may be accessed via standard interfaces, such as Oracle DB interfaces.
- the system 100 can be used for estimating mass flow in a chemical synthesis or an industrial process.
- the system 100 may be used in the pharmaceutical industry to estimate the mass flow in drug manufacturing processes, may be used in the chemical industry to estimate the mass flow in the production of various chemical compounds or formulations, and/or may be used in the biotechnology industry to estimate the mass flow in the production of biologies.
- the cloud service provider 102 may be configured to provide remote capabilities to estimate mass flow in a chemical synthesis or industrial process by providing remote access to the mass-flow estimator component 112.
- the cloud service provider 102 may be a hosted service such as a company that offers cloud computing services to businesses and individuals such that the cloud service provider 102 provides the infrastructure, software, and platforms required to host, manage, and deliver cloud-based services.
- the cloud service provider 102 may provide infrastructure as a service, platform as a service, software as a service, and/or may be an interface into a blockchain infrastructure that may or may not be hosted by the cloud service provider 102.
- the cloud service provider 102 may be configured to scale up or down its computing resources based upon demand from users at a given moment.
- the cloud service provider 102 may be implemented on a block chain.
- the cloud service provider 102 may utilize a distributed ledger to store and verify mass flow estimates generated by the mass-flow estimator component 112. Users may be authenticated and/or authorized by a secure digital certificate, encryption key, single-sign on, or other secure mechanism.
- the data may be stored and calculated in a secure and tamper proof manner to provide transparency and accountability to all users.
- the smart contracts may include executable code that defines a manufacturing process in terms of one or more or the components within the mass-flow estimator component 112 in a manner consistent with transparency and security settings.
- the personal computers 104 and the mobile device 106 communicate with each other via a network 108.
- the network 108 may be Wi-Fi, ethernet, Bluetooth, etc. and may utilize the internet and associated protocols, such as TCP/IP.
- the network 108 may be a local area network, a wide-area network, a physical bus (such as a Universal Serial Bus), the internet, or some combination thereof.
- the personal computer 104 and mobile device 106 may interface with the cloud service provider 102 to the calculate mass flow in a chemical or industrial process as determined by the subprocess metrics 144.
- a specialized application for interfacing with the mass-flow estimator component 112 may be used, such as a mobile application on the mobile device 106 or a desktop application on the personal computer 104.
- the communications may include transmitting data in HTML, XML, JSON, YAML, or any data format.
- the mass-flow estimator component 112 may provide user-level accounts to individuals through a typical login mechanism.
- the mass-flow estimator component 112 may be a web application, a webserver, a web service, etc. and may utilize one or more protocols to communicate data.
- the cloud service provider 102 may provide the mass-flow estimator component 112 as a webpage, a webapp, a program for download and execution on the computer 104 or the mobile device 106.
- the mass-flow estimator component 112 includes various sub-components, such as an electronic notebook component 150, a component for linking experimental IDs 142, a process metrics estimator component 152, a synthesis tree component 156, a search component 158, a GUI component 160, and a prediction engine 164.
- the mass-flow estimator component 112 of the system 100 may be implemented as a cloud-based software application accessible via a web interface.
- the system 100 may also be implemented as a standalone software application installed on a local computer or server.
- the system 100 may include specialized hardware such as a high-performance computer or server, or other types of computing and networking equipment.
- the mass flow calculator component 112 of the system 100 may be implemented on a variety of computing platforms, including desktop, laptop, server, or cloud-based configurations.
- the software may be written in various programming languages, such as Python, Julia, Java, C++, or other languages.
- the system 100 may also incorporate various data visualization tools, such as d3.js, Plotly.js, or Matplotlib, to allow users to visualize the experimental data and the process metrics in an intuitive manner.
- the electronic notebook component 150 enables the selection of a one or more experimental IDs 142 that are stored within a database 132.
- the electronic notebook component 150 may use a REST API and may be accessed via the Mass flow calculator component.
- Each experimental ID 142 corresponds to a subprocess having an input material and an output material.
- Information about each experimental ID 142 may be found in the subprocess metrics 144 also included in the database 132.
- the subprocess metrics 144 may include historical data such that each experimental ID 142 can be associated within one or more historical experiments as found within the subprocess metrics 144 to estimate or predict a metric.
- the linking component 154 links the one or more experimental IDs 142 to form an integrated process including all of the corresponding subprocesses, thereby synthesizing a target material.
- the process metrics estimator component 152 estimates a plurality of process metrics (e.g., using the subprocess metrics 144), each of which corresponds to a respective experimental ID of the one or more experimental IDs 142. Fig. 4 and Fig 5.
- the process metrics estimator component 152 also estimates an aggregate process metric that is a function of the estimated plurality of process metrics (which can be column summations).
- the process metrics estimator component 152 can estimate various process metrics related to the chemical synthesis process. This component can be used for evaluating the efficiency, sustainability, and environmental impact of the integrated process formed by linking the selected experimental IDs.
- the process metrics estimator component 152 estimates a plurality of process metrics, where each of these metrics corresponds to a respective experimental ID from the one or more experimental IDs selected by the user.
- process metrics can encompass a wide range of parameters, including but not limited to process mass intensity (PMI), solvent intensity (SI), water intensity (WI), global warming potential (GWP), cost, yield, and various environmental metrics.
- the process metrics estimator component 152 may calculate these metrics based on the experimental data associated with each experimental ID, which may include information from one or more previously performed experiments. This data can be retrieved from the subprocess metrics 144 stored in the database 132 or obtained from other relevant sources.
- the process metrics estimator component 152 may employ advanced algorithms, machine learning techniques, or empirical models to estimate the process metrics accurately. It may consider factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other relevant parameters to derive these metrics.
- the process metrics estimator component 152 may be used to estimate an aggregate process metric corresponding to the integrated process.
- This aggregate process metric is a function of the estimated plurality of process metrics for the individual experimental IDs.
- the function used to calculate the aggregate process metric can vary depending on the specific requirements and the nature of the process being analyzed.
- the aggregate process metric may be a simple summation of the individual process metrics. For example, if the process metrics being considered are PMI values for each subprocess, the aggregate process metric could be the sum of these PMI values, representing the overall PMI for the integrated process.
- the aggregate process metric may be a weighted combination of the individual process metrics, where different weights are assigned to different metrics based on their relative importance or impact on the overall process.
- the aggregate process metric could be a weighted sum of PMI, SI, WI, and GWP, reflecting the combined environmental impact of the integrated process.
- the aggregate process metric may be a more complex function that incorporates additional factors or constraints.
- the aggregate process metric could be a multi-objective optimization function that considers not only the process metrics but also factors such as cost, yield, or specific environmental targets.
- the process metrics estimator component 152 may also provide the capability to compare the estimated process metrics and the aggregate process metric with corresponding metrics from other processes or industry benchmarks. This comparison can aid in evaluating the relative performance and sustainability of the integrated process under consideration.
- the process metrics estimator component 152 may interact with other components of the system 100, such as the search component 158 and the prediction engine 164, to explore alternative experimental IDs or prophetic experiments. When an alternative experimental ID or a prophetic experiment is selected, the process metrics estimator component 152 can update the aggregate process metric accordingly, reflecting the impact of the proposed change on the overall process metrics.
- the process metrics estimator component 152 may provide visualizations or graphical representations of the estimated process metrics and the aggregate process metric. These visualizations can aid in interpreting the data and identifying areas for potential improvement or optimization.
- the process metrics estimator component 152 can be designed with a modular architecture, allowing for the integration of various modules or sub-components responsible for estimating specific process metrics. This modular approach enables flexibility, scalability, and customization, as different modules can be added, removed, or updated independently based on the specific requirements or advancements in the field.
- one module within the process metrics estimator component 152 could be dedicated to estimating process mass intensity (PMI) and solvent intensity (SI).
- This module may leverage advanced algorithms and machine learning techniques to analyze the input and output materials, reaction conditions, and solvent usage data associated with each experimental ID. It may also incorporate industry-specific heuristics or empirical models to improve the accuracy of PMI and SI estimations.
- Another module could focus on estimating water intensity (WI) and global warming potential (GWP).
- This module may integrate with external databases or life cycle assessment (LCA) tools to obtain relevant data on the environmental impact of various materials and processes involved in the chemical synthesis. It could also employ computational fluid dynamics (CFD) simulations or thermodynamic calculations to estimate the energy consumption and associated GWP contributions.
- CFD computational fluid dynamics
- Yet another module within the process metrics estimator component 152 could be responsible for estimating cost-related metrics, such as raw material costs, energy costs, and overall process costs.
- This module may interface with enterprise resource planning (ERP) systems, supply chain management systems, or market data feeds to obtain up-to-date information on material prices, energy costs, and other relevant cost factors.
- ERP enterprise resource planning
- the process metrics estimator component 152 may incorporate advanced uncertainty quantification techniques to provide confidence intervals, credible regions, or probability distributions for the estimated process metrics. These uncertainty estimates can be particularly valuable when dealing with incomplete or uncertain input data, or when accounting for inherent variability in the chemical synthesis processes.
- the process metrics estimator component 152 may also be integrated with optimization algorithms or decision support systems.
- the estimated process metrics could serve as objective functions or constraints in the optimization process, enabling the identification of optimal process parameters, reaction conditions, or material selections that minimize undesirable metrics (e.g., PMI, GWP) while maximizing desirable ones (e.g., yield, cost-effectiveness).
- undesirable metrics e.g., PMI, GWP
- desirable ones e.g., yield, cost-effectiveness
- process metrics estimator component 152 could be designed to leverage distributed computing resources or cloud-based platforms. This would allow for parallel processing of multiple experimental IDs or the execution of computationally intensive simulations or calculations required for estimating certain process metrics. Cloud-based deployment could also facilitate collaborative efforts, where multiple users or research teams can contribute experimental data and access the estimated process metrics from various locations.
- the process metrics estimator component 152 may incorporate machine learning capabilities to continuously improve its estimation accuracy. As more experimental data becomes available, the component could retrain its models or update its algorithms, leveraging techniques such as transfer learning or active learning to enhance the estimation performance over time.
- the process metrics estimator component 152 could also be integrated with advanced visualization tools or dashboards, providing intuitive and interactive representations of the estimated process metrics. These visualizations could include interactive charts, graphs, Sankey diagrams, or even virtual reality (VR) or augmented reality (AR) environments, allowing users to explore the data from different perspectives and gain deeper insights into the chemical synthesis process.
- advanced visualization tools or dashboards providing intuitive and interactive representations of the estimated process metrics.
- These visualizations could include interactive charts, graphs, Sankey diagrams, or even virtual reality (VR) or augmented reality (AR) environments, allowing users to explore the data from different perspectives and gain deeper insights into the chemical synthesis process.
- the process metrics estimator component 152 may support the integration of user-defined metrics or custom calculations. This would enable users to incorporate domain-specific knowledge, proprietary algorithms, or specialized requirements into the estimation process, tailoring the component to their specific needs or industry standards.
- Fig. 4 may be values for a particular batch run or for a per kilogram of target material or may use any other unit of measurement.
- Fig. 5 is the GUI showing a summary of the outputs of the linked experimental IDs of Fig. 4.
- the synthesis tree component 156 provides a synthesis (illustrated in Fig. 6) as a synthesis tree 500 (as shown in the GUI in Figs. 7) corresponding to the one or more experimental IDs 142.
- the search component 158 is configured for searching for structurally similar reaction subprocesses for each of the experimental IDs 142 to determine a plurality of similar experiments.
- the search component 158 may be configured to search for similar reaction subprocesses for each of the experimental IDs 142 stored in the database 132. This search can help determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the selected one or more experimental IDs 142.
- the search component 158 may identify similar experiments by looking for experiments with the same primary input molecule and target molecule as one or more of the selected experimental IDs 142.
- the search component 158 may ignore or consider helper molecules, such as buffers or catalysts, when identifying similar experiments.
- the search component 158 can facilitate the selection of an alternative experimental ID 142 for at least one of the initially selected one or more experimental IDs 142. This selection of an alternative experimental ID 142 can be made through the GUI component 160, which may display the search results from the search component 158.
- the process metrics estimator component 152 can then update the aggregate process metric in accordance with the selected alternative experimental ID 142.
- the search component 158 may identify recycled mass streams that are recycled to the beginning of a subprocess or the integrated process. This information can be used to estimate the effect of solvent recycling on the mass flow and process metrics, as illustrated in Fig. 11 of the patent application. [0039]
- the search component 158 may utilize various techniques to identify similar experiments, such as chemical structure matching, reaction type matching, or machine learning models trained on historical experimental data. The search component 158 can provide a tool for process optimization by allowing users to explore alternative experimental IDs 142 and their corresponding process metrics, enabling informed decision-making during chemical process development.
- the search component 158 may work in conjunction with other components of the mass-flow estimator component 112, such as the synthesis tree component 156, which provides a visual representation of the linked experimental IDs 142, and the prediction engine 164, which can suggest prophetic experiments based on predictive models or retrosynthesis tools.
- the mass-flow estimator component 112 such as the synthesis tree component 156, which provides a visual representation of the linked experimental IDs 142, and the prediction engine 164, which can suggest prophetic experiments based on predictive models or retrosynthesis tools.
- the search component 158 may employ chemical structure matching algorithms to identify experimental IDs 142 that involve the same input and output molecular structures. This approach can be particularly useful when the goal is to find alternative synthetic routes or optimize specific reaction steps within the integrated process.
- the search component 158 may maintain a database or index of chemical structures associated with each experimental ID 142, enabling efficient querying and matching based on structural similarity metrics. These metrics can consider factors such as functional groups, bond connectivity, stereochemistry, and other relevant structural features.
- the search component 158 can utilize well-established chemical structure representation formats, such as SMILES, InChi, or molecular fingerprints, to facilitate structure matching and comparison.
- the search component 158 may employ reaction type matching, where it identifies similar experiments based on the types of chemical reactions involved. This approach can be beneficial when optimizing specific reaction classes or exploring alternative reagents or conditions for a particular transformation.
- the search component 158 may leverage reaction classification algorithms or databases that categorize reactions based on their mechanistic patterns, functional group transformations, or other relevant criteria.
- the search component 158 can incorporate machine learning techniques to identify similar experiments.
- the search component 158 may train machine learning models on historical experimental data, including reaction conditions, reagents, yields, and other relevant features. These models can learn patterns and similarities between experiments, enabling the search component 158 to suggest relevant alternatives based on the provided experimental IDs 142.
- the machine learning models employed by the search component 158 can range from simple similarity measures, such as k-nearest neighbors or cosine similarity, to more complex models like neural networks or decision trees. The choice of model can depend on factors such as the complexity of the chemical space, the availability of training data, and the desired trade-off between accuracy and computational efficiency.
- the search component 158 may combine multiple search strategies, such as chemical structure matching and reaction type matching, to provide more comprehensive and relevant search results. This hybrid approach can leverage the strengths of different search techniques and provide users with a diverse set of alternative experimental IDs 142 to explore.
- the search component 158 may incorporate user feedback and preferences to refine and enhance its search capabilities over time. For example, users may be able to provide ratings or rankings for the search results, which can be used to update the underlying search algorithms or models. This adaptive approach can ensure that the search component 158 remains relevant and aligned with the specific needs and objectives of the users.
- the search component 158 may also integrate with external databases or resources to expand its search capabilities. For instance, it may access public or proprietary reaction databases, literature repositories, or patent databases to identify additional relevant experiments or synthetic routes. This integration can provide users with a broader perspective and facilitate the exploration of alternative experimental IDs 142 beyond the scope of the local database 132.
- the search component 158 may provide additional information or analysis to support the decision-making process. This information can include statistical summaries of the identified similar experiments, such as distributions of yields, process metrics, or reaction conditions. The search component 158 may also highlight key differences or similarities between the alternative experimental IDs 142 and the initially selected ones, enabling users to make informed choices based on their specific requirements or constraints.
- the searching may be activated by the button 502 on the webpage as is shown in Fig. 7. Also shown in the GUI on Fig. 7, is that the molecular role 504 is selectable as well.
- the GUI of Fig. 7 When the GUI of Fig. 7 is presented to a user, it shows the synthesis steps with to thereby give the user the ability to modify the reactants. Although Fig. 7 only shows 2 steps, additional steps of the process may be shown by allowing the user to scroll down.
- Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure where there is a drop-down menu to change experiments. [00106] Referring again to Fig. 1, each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs 142.
- the GUI component 160 is configured for selecting an alternative experimental ID 142 for at least one of the one or more experimental IDs 142.
- Figs 8A and 8B show Sankey Diagrams representing the mass flow from the selected experiment, e.g., the experiment of Fig. 6.
- Fig. 9 shows an example result of the search component 158 in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 7.
- Fig. 9 shows alternative experimental IDs and in some embodiments, it can be matched up to the ones that it replaces via another column, for example.
- the resulting PMI vs. Yield values for each alternative experimental ID may be shown as in Fig. 10. Referring yet again to Fig.
- the process metrics estimator component 152 updates the aggregate process metric in accordance with the alternative experimental ID 142 when a user selects the alternative experimental ID 142 to replace one of the initially entered in experimental ID 142.
- Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow starting materials’, e.g., reactants, catalysts, etc. is estimated for the experiment. Thus, Fig. 11 shows another view off the GUI.
- the synthesis tree component 156 of the system 100 may be used to identify a set of possible synthesis routes leading to a target material.
- the search component 158 may be used to identify a set of similar experiments to each experimental ID 142 of the experimental IDs 142 and determine an optimal set of reactions to synthesize the target material.
- the process metrics estimator component 152 can estimate the process metrics for each of the optimal set of reactions, and the aggregate process metric for the integrated processes.
- the GUI component 160 can be used to visualize the synthesis route and adjust the experimental IDs 142, as well as adjust the process parameters to optimize the process metrics.
- the experimental data corresponding to each experimental ID 142 includes information about one or more previously performed experiments which are stored in the subprocess metrics 144.
- the aggregate process metric can be a process mass intensity, a solvent intensity, a water intensity, a global warming potential, cost, yield, an environmental metric, etc.
- the GUI component 160 of the system 100 may include various user interface elements, such as dropdown menus, checkboxes, sliders, and text input fields, to enable users to select experimental IDs 142, adjust process parameters, and visualize the synthesis route.
- the GUI component 160 can also be designed to be accessible via a web interface, allowing users to access the system 100 from anywhere with an internet connection.
- the GUI component 160 is configured for selecting a role for a molecule such that the molecule is associates with an experimental ID 142 of the one or more experimental IDs 142.
- the GUI component 160 can also be configured for adjusting a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
- some of the experimental IDs 142 may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based upon the input parameters. Modifying the input parameters can update these output values.
- the act of selecting the one or more experimental IDs 142 is performed by a user utilizing the GUI component 160.
- the GUI component 160 can also be used to query a prediction engine 164 to determine a prophetic experiment. It can be used to replace at least one of the one or more experimental IDs 142 with the prophetic experiment and estimate a prophetic process metric corresponding to the prophetic experiment.
- the system 100 can also update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs 142.
- the prophetic process metric can include a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region.
- the plurality of process metrics can be a ratio, such as PMI over Yield.
- the one or more experimental IDs 142 can be selected out of order, and the act of linking the plurality of experiment IDs can include the act of ordering the one or more experimental IDs 142 to form the integrated process including all of the corresponding subprocesses, in order to synthesize the target material.
- the system 100 is designed to efficiently estimate mass flow and enable users to adjust and optimize various process metrics to improve the efficiency and sustainability of chemical or industrial processes.
- the mass-flow estimator component 112 may assign a confidence score. If the mass-flow estimator component 112 bases some or all of the estimates on data, models generated from data, or Monte Carlo simulation data, a confidence score can be assigned to the estimates of mass flow values to indicate the quality of the estimate. This could be included directly in the output or be derived from the standard deviation or variance in the sample data used to make the estimation, the min-max of mass flow data points, etc.
- a frequentist confidence store may be derived using frequentist statistics.
- a confidence score may use sample data of a distribution, hypothesis testing, p-values, significance testing, confidence intervals etc.
- a confidence score is calculated for each (or a set of) sample values using posterior probabilities in a Bayesian estimate, which represent the updated belief about the mass flow values.
- the confidence score for example, may be a credible interval of a posterior distribution or of a Bayesian estimator.
- the mass-flow estimator component 112 also includes the communications component 162.
- the communications component 162 may facilitate seamless communication and data exchange between multiple software applications, devices, and systems. That is, the communications component 162 may include protocol handling, message formatting, data serializing, encryption, and authentication to facilitate the communication with the computers 104 and/or the mobile device 106.
- the communications component 162 may utilize a message formatting mechanism to format the messages into formats, such as XML, JSON, binary formats, and/or proprietary message formats.
- the communications component 162 may utilize various encryption algorithms, such as RSA, AES, ECC, symmetric encryption, asymmetric encryption etc. to enable secure communications between the mass-flow estimator component 112 and the computers 104 and/or the mobile device 106.
- the GUI component 160 can render a display for use by the computer 104 and/or the mobile device 106.
- the GUI component 160 may be a webpage-based provider, such as flask, an HTML server, a web framework, etc.
- the GUI component 160 may provide widgets, information, buttons, options, and menus to thereby facilitate a user’s interaction with the mass-flow estimator component 112.
- the GUI component 160 can be used to log into user accounts 146 so that a user can create, save, or retrieve the experimental IDs 142 and/or the subprocess metrics 144, or otherwise interface with any account features. Additionally or alternatively, the GUI component 160 can save favorites, select default parameters, or adjust default values.
- the GUI component 160 can direct other components to execute instructions based upon a workflow initiated by a user. That is, the GUI component 160 may receive events, such as a mouse click, button press, or GUI widget interaction to initiate a routine, series of steps, or series of acts. For example, the GUI component 160 may guide a user step-by-step on how to set up and work with the subprocess metrics 144 within the database 132.
- the GUI component 160 may also be used to visualize the results of the process models and the waste estimates, in aggregate, in simulation, and/or may provide various visualization tools to analyze the data.
- the data may be stored in the database 132 or the Electronic Notebook Component 150.
- each user can log into a user account 146 to visual the results of their processes, the results of modifications to their processes on the entire synthesis chain, obtain a direct comparison and/or historical accuracy of their process mass flow values.
- a resource dispatcher 110 may dispatch requests to perform an action to one or more virtual servers 122, each of which has a virtual processor 124, a virtual memory 126, and a virtual disk space 128.
- the virtual servers 122 can be executed on one or more servers 121 on a server farm 119 as dispatched and activated by the resource dispatcher 110.
- Fig. 2 show a block diagram illustration of a computing device 200 to calculate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure.
- the computing device 200 of Fig. 2 may be the computer 104 or mobile device 106 of Fig. 1.
- the computing device 200 includes an VO interface 210 to communicate therewithin.
- the computing device 200 includes a data store 204, a processor 206, a network interface 208, a memory 225, and user I/O devices 226.
- the data store 204 stores data and may be a hard drive, flash drive, thumb drive, volatile memory, non-volatile memory, semi-volatile memory etc.
- the processor 206 can execute one or more processor-executable instructions 212, which may be stored in the data store 204 and/or the memory 225.
- the processor 206 can execute processor-executable instructions 212 stored in memory 225 that was retrieved from the data store 204.
- the memory 225 also includes program data 214 that may include information related to the processor-executable instructions 212.
- the computing device 200 may include user I/O devices 226, such as a cursor device 230 (e.g., touchscreen or mouse), a keyboard 232 (virtual or physical), and/or a monitor 228 (which may be a touchscreen).
- the computing device 200 communicates with the network 202 via a network interface 208.
- the mass flow calculation functionality may reside wholly within the computing device 200 of Fig. 2.
- the mass-flow estimator component 112 of Fig. 1 may reside within the processor-executable instructions 212 of Fig. 2 as mass flow calculation and mass-flow estimator component 242.
- the electronic notebook component 150 may be the same or similar to the electronic notebook component 150, the process metrics estimator component 152, the linking component 154, the synthesis tree component 156, the search component 158, the GUI Component 160, the communications component 162, and the prediction engine 164 of Fig. 1, respectively.
- the database 244 may be similar to the database 132 of Fig. 1.
- the database 244 may, for example, be an Oracle, SQLite, postgresql, MariaDB, MySQL, or any other database embedded on the computing device 200.
- the experimental IDs 246, the subprocess metrics 248, and the user accounts 250 of Figs. 2 may be similar or identical to the experimental IDs 142, the subprocess metrics 144, and the user accounts 146 of Fig. 1, respectively.
- the mass-flow estimator component 242 may reside wholly on a local device (such as on the computers 104, the mobile device 106, etc.) may be partially within a cloud service provider 102, and/or may be organized in a hybrid local and cloud configuration.
- the mass-flow estimator component 242 may be an application, may be executed on the computers 104, the mobile device 106, the cloud service provider 102, the computing device 200, etc. or some combination thereof.
- Fig. 3 is a flowchart of an example process 300.
- one or more process blocks of Fig. 3 may be performed by a device, such as a computing device.
- the process involves selecting one or more experimental IDs from an electronic notebook.
- Each experimental ID corresponds to a subprocess that has at least one input material and at least one output material. Those input and output materials might be the same or different.
- the experimental ID may include experimental data corresponding to one or more previously performed experiments.
- the experimental IDs may be selected by using a GUI, such as that shown in Fig. 4. Or, Fig. 4 may show the listed of experimental IDs after a list is inputted into a GUI and may be after linking.
- a role is selected for a molecule associated with an experimental ID from the One or more experimental IDs.
- Fig. 7 shows a drop-down box where a role can be selected.
- Act 306 the one or more experimental IDs are linked to form an integrated process that includes all of the corresponding subprocesses, which together synthesize a target material. This is done automatically or with user interaction.
- Act 308 involves estimating a plurality of process metrics, where each of the metrics corresponds to a respective experimental ID from the one or more experimental IDs.
- an aggregate process metric is estimated, which corresponds to the integrated process, and is a function of the estimated plurality of process metrics.
- the aggregate process metrics may be, for example, a process mass intensity, a solvent intensity, a water intensity, a global warming potential, a cost, an energy uptake, a waste score, a yield, and any other environmental metric.
- the process then moves to Act 312, which involves providing a synthesis tree (e.g., as shown in Fig. 5) that corresponds to the one or more experimental IDs.
- Act 314 searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs.
- similar experiments are experiments with a primary input molecule and a target molecule that is the same as one or more experimental IDs. Helper molecules, such as buffers, may be different.
- Act 316 selects an alternative experimental ID for at least one of the one or more experimental IDs. Then, in Act 318, the aggregate process metric is updated in accordance with the alternative experimental ID. In Act 320, a process parameter is updated. Act 322 involves updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. Act 324 queries a prediction engine to determine a prophetic experiment. Act 326 then replaces at least one of the one or more experimental IDs with the prophetic experiment.
- a prophetic process metric is estimated, which corresponds to the prophetic experiment, and the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
- the prophetic process metrics can have a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region, and the multiple process metrics can be expressed as a ratio, such as PMI over Yield.
- the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
- Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations so that different so that different experiments may be compared.
- the process allows for the consideration and comparison of experimental data from one or more previously conducted experiments within each experimental ID.
- the aggregate process metric may be based on process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield or environmental impact, either alone or in combination with other customizable variations.
- a molecule associated with an experimental ID can be selected to play a specific role (e.g. product, starting material, reagent, solvent, catalyst), while process parameters can be adjusted to achieve optimal results.
- process metrics can be updated to achieve target parameters.
- the user can select experimental IDs using a graphical user interface. Additionally, the one or more experimental IDs can be selected out of order, but must be properly ordered to form an integrated process that includes all corresponding subprocesses that synthesize the target material.
- the customizable variations can be implemented separately or in combination with one another, and the process may include more, fewer, or differently arranged blocks than those illustrated in Figure 3.
- the blocks in the process can be performed simultaneously when necessary.
- the process can provide a versatile and customizable process for optimizing chemical reactions through the consideration of various experimental and process parameters.
- process 300 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 3. Additionally, or alternatively, two or more of the blocks of process 300 may be performed in parallel.
- Figs. 14, 15 and 16 are described as follows to illustrate the operation of the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2
- the predictive estimation of sustainability metrics (solvent intensity, water intensity, product carbon footprint) with an electronic notebook-data using the system 100 of Fig. 1 can utilize a prediction engine, e.g., retro-synthesis tools
- the production of highly complex molecules such as APIs (active pharmaceutical ingredients) is a multi-step synthesis process that starts from already complex raw materials.
- building blocks are for example the model compounds shown in Fig. 14.
- the prediction engine 164 is a component of the mass-flow estimator component 112 in the system 100 for calculating mass flow and estimating sustainability metrics in chemical synthesis. This prediction engine 164 can serve various purposes and functionalities within the overall system.
- the prediction engine 164 may be queried to determine a prophetic experiment.
- the term "prophetic experiment” can refer to a hypothetical or simulated experiment that has not yet been physically performed.
- the prediction engine 164 can utilize various techniques, such as machine learning models, computational chemistry methods, or rule-based algorithms, to generate predictions or suggestions for potential chemical reactions or processes that could be explored. These predicted or prophetic experiments can then be incorporated into the overall integrated process being analyzed by the system.
- the prediction engine 164 may be used to replace at least one of the one or more experimental IDs 142 with a prophetic experiment suggested by the prediction engine 164.
- the process metrics estimator component 152 can then estimate a prophetic process metric corresponding to this prophetic experiment.
- the aggregate process metric which is a function of the estimated plurality of process metrics, can be updated to include the prophetic experiment in place of the replaced experimental ID(s) 142.
- the prophetic process metric estimated by the process metrics estimator component 152 for the prophetic experiment may include a range of values.
- This range of values can take various forms, such as a confidence interval, a confidence region, a credible interval, or a credible region. These ranges can provide an indication of the uncertainty or variability associated with the predicted or simulated experiment, allowing users to assess the reliability or robustness of the predictions made by the prediction engine 164.
- the prediction engine 164 may employ retrosynthesis tools or algorithms to suggest potential synthesis routes or pathways for a target molecule. These retrosynthesis tools can analyze the target molecule and propose a series of chemical reactions or transformations that could be used to synthesize the target from simpler starting materials or building blocks. The prediction engine 164 can then use this suggested synthesis route, along with data from the electronic notebook component 150 and the subprocess metrics 144, to estimate process metrics and sustainability indicators for the proposed synthetic pathway.
- the prediction engine 164 may be integrated with other components of the system, such as the search component 158, to identify similar reactions or processes from the database 132 that could inform or refine the predictions made by the prediction engine 164.
- the search component 158 may identify similar experimental IDs 142 or subprocesses that have been previously performed, and the prediction engine 164 can use this information to improve the accuracy or reliability of its predictions.
- the prediction engine 164 can also be used in conjunction with the synthesis tree component 156 and the GUI component 160.
- the synthesis tree component 156 provides a visual representation of the integrated process, including the one or more experimental IDs 142 and their corresponding subprocesses.
- the prediction engine 164 may suggest modifications or alternatives to this synthesis tree, which can be visualized and edited through the GUI component 160, allowing users to explore different scenarios and optimize the integrated process based on the predictions made by the prediction engine 164.
- the prediction engine 164 may incorporate various types of data and models to make its predictions. This can include experimental data from the electronic notebook component 150, thermodynamic calculations, empirical correlations, or machine learning models trained on relevant chemical data.
- the prediction engine 164 may also integrate with external databases or resources to obtain additional information or data relevant to the chemical processes being analyzed.
- Model compound A is a halogenated aromatic with an additional nitrile group
- compound B is a halogenated heteroaromatic molecule.
- Such are so-called value- added complex intermediates in the chemical industry, which means that they are produced from base chemicals via one or more synthesis steps.
- Such molecules usually carry more than one functional group (e.g. halogenic substituent, nitrile group, amine group, alcohol group, ester from boronic acid, etc.).
- Starting materials should be base chemicals, available in bulk quantities, that are available in a life cycle inventory database (e.g. ecoinvent)
- SynthiaTM is a retro-synthesis tool that suggests production routes for a certain molecule and the above-mentioned information.
- An alternative to SynthiaTM is the tool ASKCOS from MIT (Massachusetts Institute of Technology), which is available free of charge.
- SynthiaTM suggests a two-step synthesis from the base chemical maleic acid and hydrazine (Fig. 15).
- Fig. 14 shows model compounds for raw materials to synthesize active pharmacal ingredients while Fig. 15 shows how it may be be presented in the GUI.
- Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure.
- Both chemicals are available in life cycle inventory database such as ecoinvent. Additionally, the type of reaction is shown and basic information on reaction conditions (e.g. solvent, catalyst, temperature during reaction) are mentioned.
- the disclosed method does offer the automated generation of combined PMI / yield data of recorded chemical conversions and related processes, thereby preparing data sets, which can be used to allow a faster and more qualified plausibility testing of 'greener' retro-synthetic planning by systematic analysis of similar structures and chemical conversions.
- This information can be used to search for similar reactions by the search component 158 to obtain the average amount of solvent, water, catalyst, etc. consumed in each step by looking at the average PMI for the specific reaction type that was suggested by the retro-synthesis tool.
- the prediction engine 164 can also provide an average chemical yield for the respective synthesis steps as shown in Fig. 16. Based on the comparison of PMI and chemical yield for dozens of similar reactions documented in the ELN-system, an average amount of solvent, water, catalyst, etc. can be applied.
- an uncertainty estimate may be made in the environmental footprint evaluation.
- the deviation of the PMI data points to the average PMI as well as the deviation of data points for the chemical yield to the average yield can provide an indication on the accuracy on the environmental footprint evaluation.
- Product carbon footprints can be estimated based on life cycle inventory database entries for raw materials, solvents, water, catalysts, etc.
- Figs. 11 and 12 describe the function of the recycling functionality incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2. In this module the effect of solvent recycling on sustainability metrics is estimated.
- Fig. 17 describes the function of the energy module incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2.
- Fig. 17 may be a separate module used for energy calculation at each step using, e.g., input metrics.
- energy consumption of the selected processes or subprocesses can be estimated based on basic physical phenomena and/or empirical relationships. Individual unit-operations are selected within the tool and important parameters are adjusted by the user. Examples for unit operations are, but are not limited to, heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction.
- a summary of the energy contribution the unit-operations of the unit operations in the synthesis tree are visualized in the software, like shown in Fig. 18. The software allows adaption to various scales. The energy contribution might be included into the estimation of the final product carbon footprint.
- Fig. 19 describes the possibility to include weight-based carbon footprints for each individual staring material. These can be assigned manually, can be retrieved from external date vendors (e.g. Ecoinvent), or can be assigned based on the individual role of each starting material. The output presented in the GUI can also be a weight-based carbon footprint estimation.
- weight-based carbon footprints for each individual staring material. These can be assigned manually, can be retrieved from external date vendors (e.g. Ecoinvent), or can be assigned based on the individual role of each starting material.
- the output presented in the GUI can also be a weight-based carbon footprint estimation.
- a computer-implemented method for calculating mass flow and estimating sustainability metrics comprising: selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material; linking the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses wherein the integrated process synthesizes a target material; estimating a plurality of process metrics wherein each of the plurality of process metrics corresponds to a respective experimental IDs of the one or more experimental IDs; estimating an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics; providing a synthesis tree corresponding to the one or more experimental IDs; searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs; selecting an alternative experimental ID for at least one of the one or more experimental IDs; and updating the aggregate process
- each experimental ID includes experimental data corresponding to one or more previously performed experiments.
- the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
- the subprocesses include at least one of chemical reactions or purifications.
- each experimental ID is associated with specific equipment used in the subprocess, and the method further comprises adjusting equipment settings based on the subprocess requirements.
- estimating the energy consumption metric comprises applying at least one of empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs.
- estimating the energy consumption metric comprises applying at least one of empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs.
- a data processing system comprising means for carrying out the method of any one of aspects 1 to 56.
- a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56.
- a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56.
Landscapes
- Business, Economics & Management (AREA)
- Engineering & Computer Science (AREA)
- Human Resources & Organizations (AREA)
- Strategic Management (AREA)
- Economics (AREA)
- Theoretical Computer Science (AREA)
- Entrepreneurship & Innovation (AREA)
- General Physics & Mathematics (AREA)
- Marketing (AREA)
- Tourism & Hospitality (AREA)
- Physics & Mathematics (AREA)
- General Business, Economics & Management (AREA)
- Operations Research (AREA)
- Quality & Reliability (AREA)
- Chemical & Material Sciences (AREA)
- Game Theory and Decision Science (AREA)
- Development Economics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Primary Health Care (AREA)
- Analytical Chemistry (AREA)
- Manufacturing & Machinery (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Sustainable Development (AREA)
- Data Mining & Analysis (AREA)
- Educational Administration (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Crystallography & Structural Chemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Computing Systems (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Stored Programmes (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24732622.6A EP4725021A1 (en) | 2023-06-12 | 2024-06-11 | Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23178759.9 | 2023-06-12 | ||
| EP23178759 | 2023-06-12 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024256360A1 true WO2024256360A1 (en) | 2024-12-19 |
Family
ID=86760182
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2024/066006 Ceased WO2024256360A1 (en) | 2023-06-12 | 2024-06-11 | Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4725021A1 (en) |
| TW (1) | TW202520282A (en) |
| WO (1) | WO2024256360A1 (en) |
-
2024
- 2024-06-11 EP EP24732622.6A patent/EP4725021A1/en active Pending
- 2024-06-11 TW TW113121459A patent/TW202520282A/en unknown
- 2024-06-11 WO PCT/EP2024/066006 patent/WO2024256360A1/en not_active Ceased
Non-Patent Citations (8)
| Title |
|---|
| GREEN CHEM., vol. 17, no. 3111, 2015 |
| GSK STUDIE: ORG. PROCESS RES. DEV., vol. 15, 2011, pages 912 - 917 |
| JIMENEZ-GONZALES ET AL., ORG. PROCESS RES. DEV, vol. 15, no. 4, 2011, pages 912 - 917 |
| LOUREIRO HUGO ET AL: "ChemPager: Now Expanded for Even Greener Chemistry", CHIMIA INTERNATIONAL JOURNAL FOR CHEMISTRY, vol. 73, no. 9, 18 September 2019 (2019-09-18), CH, pages 724, XP093208077, ISSN: 0009-4293, Retrieved from the Internet <URL:https://chimia.ch/chimia/article/download/2019_724/621> DOI: 10.2533/chimia.2019.724 * |
| PARVATKER ET AL., ACS SUSTAINABLE CHEM. ENG., vol. 7, 2019, pages 6580 - 6591 |
| PROCESSES, vol. 10, 2022, pages 1274 |
| SHARMA PANKAJ ET AL: "DOZNTM 2.0: A quantitative green chemistry evaluator for a sustainable future", JOURNAL OF ORGANOMETALLIC CHEMISTRY, ELSEVIER, AMSTERDAM, NL, vol. 970, 29 April 2022 (2022-04-29), XP087076775, ISSN: 0022-328X, [retrieved on 20220429], DOI: 10.1016/J.JORGANCHEM.2022.122367 * |
| SHERER EDWARD C. ET AL: "Driving Aspirational Process Mass Intensity Using Simple Structure-Based Prediction", ORGANIC PROCESS RESEARCH & DEVELOPMENT, vol. 26, no. 5, 18 April 2022 (2022-04-18), US, pages 1405 - 1410, XP093208123, ISSN: 1083-6160, Retrieved from the Internet <URL:https://pubs.acs.org/doi/pdf/10.1021/acs.oprd.1c00477> DOI: 10.1021/acs.oprd.1c00477 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4725021A1 (en) | 2026-04-15 |
| TW202520282A (en) | 2025-05-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10636007B2 (en) | Method and system for data-based optimization of performance indicators in process and manufacturing industries | |
| US11120347B2 (en) | Optimizing data-to-learning-to-action | |
| US20250094841A1 (en) | Hybrid Machine Learning | |
| US20240420026A1 (en) | Systems and methods for advanced prediction using machine-learning and statistical models | |
| Pu et al. | The analysis of strategic management decisions and corporate competitiveness based on artificial intelligence | |
| Machireddy et al. | Enhancing predictive analytics with AI-powered RPA in cloud data warehousing: A comparative study of traditional and modern approaches | |
| Sakhrawi et al. | Support vector regression for enhancement effort prediction of Scrum projects from COSMIC functional size | |
| Gangadharan et al. | Metaheuristic approaches in biopharmaceutical process development data analysis | |
| Mostofi et al. | Performance‐driven contractor recommendation system using a weighted activity–contractor network | |
| Hou et al. | A novel technology life cycle analysis method based on LSTM and CRF | |
| JP2008171171A (en) | Demand forecast method, demand forecast analysis server, and demand forecast program | |
| WO2024256360A1 (en) | Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis | |
| Walton et al. | Automated resonance fitting for nuclear data evaluation | |
| US20220067628A1 (en) | Directional stream value analysis system and server | |
| CN120543180A (en) | An automated management system for watch after-sales service | |
| Costantini et al. | On the use of mean square error and directional forecast accuracy for model selection: a simulation study | |
| Mhaskey | Unlocking business potential: The transformative impact of ERP analytics | |
| Sokolovas | Investigation of process automation with large language models | |
| JP4738898B2 (en) | Demand forecast method, demand forecast analysis server, and demand forecast program | |
| Fekete et al. | A comprehensive causal AI framework for analysing factors affecting energy consumption and costs in customised manufacturing | |
| US20250013632A1 (en) | Smart selection of data fields during data analysis | |
| Islam et al. | PREDICTIVE ANALYTICS IN SUPPLY CHAIN MANAGEMENT A REVIEW OF BUSINESS ANALYST-LED OPTIMIZATION TOOLS | |
| Chain | Review of Applied Science and Technology | |
| JORDAN | AI-POWERED PORTFOLIO MANAGEMENT IN PHARMACEUTICAL R&D5. AI-POWERED PORTFOLIO MANAGEMENT IN PHARMACEUTICAL R&D | |
| Lingqa et al. | User Interface Design of Safety Stock Prediction System Using Demand Response-ARMA Method Case Study: Jaya Lestari |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24732622 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2024732622 Country of ref document: EP Effective date: 20260112 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024732622 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024732622 Country of ref document: EP Effective date: 20260112 |
|
| ENP | Entry into the national phase |
Ref document number: 2024732622 Country of ref document: EP Effective date: 20260112 |
|
| ENP | Entry into the national phase |
Ref document number: 2024732622 Country of ref document: EP Effective date: 20260112 |
|
| ENP | Entry into the national phase |
Ref document number: 2024732622 Country of ref document: EP Effective date: 20260112 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024732622 Country of ref document: EP |