WO2024256360A1 - Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis - Google Patents

Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis Download PDF

Info

Publication number
WO2024256360A1
WO2024256360A1 PCT/EP2024/066006 EP2024066006W WO2024256360A1 WO 2024256360 A1 WO2024256360 A1 WO 2024256360A1 EP 2024066006 W EP2024066006 W EP 2024066006W WO 2024256360 A1 WO2024256360 A1 WO 2024256360A1
Authority
WO
WIPO (PCT)
Prior art keywords
experimental
ids
metrics
metric
aggregate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/EP2024/066006
Other languages
French (fr)
Inventor
Helmut Haensel
Sara WIRSING
Andreas Bathe
Alexander Dauth
Benjamin KUEHNE
Heinrich Becker
Kerstin Hell
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Merck Patent GmbH
Original Assignee
Merck Patent GmbH
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Merck Patent GmbH filed Critical Merck Patent GmbH
Priority to EP24732622.6A priority Critical patent/EP4725021A1/en
Publication of WO2024256360A1 publication Critical patent/WO2024256360A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/10Analysis or design of chemical reactions, syntheses or processes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/04Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/06Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/10Office automation; Time management
    • G06Q10/103Workflow collaboration or project management
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/30Administration of product recycling or disposal
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q50/00Information and communication technology [ICT] specially adapted for implementation of business processes of specific business sectors, e.g. utilities or tourism
    • G06Q50/04Manufacturing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q2220/00Business processing using cryptography

Definitions

  • the present disclosure relates to chemical synthesis. More particularly, the present disclosure relates to a method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis.
  • PMI Process Mass Intensity
  • the PMI can be used as an indicator of both the cost-effectiveness and environmental compatibility of a process.
  • the PMI provides information on how efficiently the process mass is used compared to the mass of the target material or product.
  • the process mass refers to all chemicals, organic solvents, water, auxiliaries such as catalysts or pH buffers, rinsing media, etc. used in the production process. This mass-based resource consumption is measured in relation to the mass of a produced chemical compound (usually kg per kg).
  • a low (small) PMI value indicates that the process is comparatively efficient and that the process mass is used optimally, while a high (large) PMI indicates that the process is inefficient and the process mass is not being used optimally, which possibly leads to a larger amount of waste and thus a high environmental impact.
  • Yet another use of PMI for the evaluation of chemical processes was pioneered by Jimenez-Gonzales et al. (Org. Process Res. Dev. 2011, 15, 4, 912- 917), who showed a direct correlation between the PMI and the potential amount of eCO2 as a measure of global warming potential (GWP).
  • SI Solvent Intensity
  • WI Water Intensity
  • DOZN TM Tool An online application that allows for a semi-quantitative evaluation of a single process step based on the 12 principles of green chemistry. It is not suitable for complex syntheses, like linear or branched multi-step syntheses, and requires detailed manual inputs. Mass-based resource consumption and Global Warming potential cannot be calculated. [0012] American Chemical Society (ACS) Green Chemistry Institute, Process Mass Intensity Calculator: An Excel -based tool used for single or multi-step reaction PMI calculation by manually entering reactants, reagents, solvents and aqueous systems, as well as the interdependencies of the above steps.
  • ACS Chemical Society
  • Process Mass Intensity Calculator An Excel -based tool used for single or multi-step reaction PMI calculation by manually entering reactants, reagents, solvents and aqueous systems, as well as the interdependencies of the above steps.
  • Chem Pager / Roche A module connected to an ERP system that allows PMI calculation for established production processes. It is unsuitable for process development as the data source is limited to ERP system recipes.
  • LCAs Life Cycle Assessments
  • PCFs Product Carbon Footprints
  • the present disclosure relates to a computer-implemented method for calculating mass flow and estimating sustainability metrics in chemical processes.
  • the method may be initiated by selecting one or more experimental IDs from an electronic laboratory notebook (ELN), where each experimental ID corresponds to a subprocess with input and output materials.
  • ESN electronic laboratory notebook
  • the selection process may be performed through a Graphical User Interface.
  • Experimental data corresponding to one or more previously performed experiments may be associated with the selected experimental IDs.
  • Further steps of the method may include linking the selected experimental IDs to form an integrated process, which synthesizes a target material.
  • the integrated process can include multiple corresponding subprocesses, and a synthesis tree corresponding to the one or more experimental IDs is provided by the software to support the process.
  • the synthesis is adjustable by the user in some embodiments.
  • the method estimates a plurality of process metrics, such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or environmental metric, wherein each of the plurality of process metrics corresponds to the respective experimental IDs of the one or more experimental IDs.
  • the process metrics may also be a ratio, such as a PMI over Yield ratio.
  • the process metrics may be provided for the whole process, parts of the process, or each of the subprocess steps.
  • the Process metrics may also be compared to one or more process metrics from other processes.
  • the process metrics can also be intuitively visualized by graphs, like a Sankey-Diagram.
  • An aggregate process metric may be calculated, which is a function of the estimated plurality of process metrics.
  • the aggregate process metrics may be, for example, a cost, a yield, an environmental metric, or the aggregation of any process metrics.
  • the aggregate process metric may be a summation of the plurality of process metrics.
  • the metrics may be on a unit basis, a mass basis, a volume basis, a batch basis, or in any unit known to one of ordinary skill in the relevant art.
  • the method may include searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each of which corresponds to one or more respective experiments of the one or more experimental IDs.
  • the search may look for reactions with the same input and output molecules and, in some embodiment, may ignore waste or support molecules such as catalysts, for example.
  • the search may also look for recycled mass streams, that means, waste streams that are recycled to the beginning of a subprocess or a process. Based on this search, at least one alternative experimental ID may be selected for replacing one or more of the selected experimental IDs.
  • a prophetic experiment can also be determined by querying a prediction engine and replacing at least one of the one or more experimental IDs with the prophetic experiment.
  • the method may then estimate a prophetic process metric corresponding to the prophetic experiment and update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
  • the prophetic process metric may include a range of values.
  • the range of values may be a confidence interval, a confidence region, a credible interval, and/or a credible region. Those cited ranges of values may also include multi-step processes.
  • the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
  • the act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material.
  • the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
  • this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs.
  • the plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
  • Some embodiments include a data processing system comprising means for the implementation of any of the acts described above.
  • the disclosure also includes a computer program and a computer-readable medium that can execute and perform the steps described in any one of the acts described above.
  • some embodiments provide an intelligent and efficient computer- implemented method for calculating mass flow and estimating sustainability metrics in chemical processes.
  • the method employs experimental data to select and link experimental IDs to form an integrated process, estimates process metrics, and calculates an aggregate process metric.
  • Some embodiments allow for the selection of alternative experimental IDs, the adjustment of a process parameter, and querying of a prediction engine to identify a prophetic experiment.
  • the method involves selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material.
  • the method then links the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses, wherein the integrated process synthesizes a target material.
  • the method further estimates a plurality of process metrics, wherein each of the process metrics corresponds to a respective experimental ID of the one or more experimental IDs.
  • the method estimates an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics.
  • the method also provides a synthesis tree corresponding to the one or more experimental IDs.
  • the method searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, with each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs.
  • the method selects an alternative experimental ID for at least one of the one or more experimental IDs and updates the aggregate process metric in accordance with the alternative experimental ID.
  • the method may involve displaying the process metrics that are estimated.
  • the plurality of process metrics estimated such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other environmental metrics, can be visually displayed through the graphical user interface.
  • This display of the estimated process metrics can provide users with a convenient overview and analysis of the efficiency, sustainability, and potential environmental impact of the integrated process.
  • the visualization of process metrics can aid in interpreting the estimation results, identifying areas for potential improvement, and guiding decisions during chemical process development.
  • advanced data visualization tools may be leveraged to provide intuitive charts, graphs, Sankey diagrams or other graphical representations of the estimated process metrics.
  • each experimental ID corresponds to a subprocess and includes experimental data from one or more previously performed experiments or historic experiments.
  • the method is able to incorporate experimental data from previously conducted experiments into the analysis of the overall integrated process.
  • This previously conducted experimental data provides the necessary information to estimate process metrics for the subprocesses associated with each experimental ID in the integrated process, as well as an aggregate process metric for the overall integrated process.
  • the experimental data may encompass details regarding factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other parameters relevant to calculating process metrics like process mass intensity, yield, cost, and environmental impact indicators. Overall, the consideration of previously conducted experimental data lends greater accuracy and reliability to the method's ability to analyze the efficiency and sustainability of chemical synthesis processes.
  • the method further comprises comparing the aggregate process metric with the alternative experimental ID to the aggregate process metric before updating the aggregate process metric based on the alternative experimental ID. More specifically, the aggregate process metric calculated using the initially selected experimental ID(s) is compared to what the aggregate process metric would be if a particular alternative experimental ID was used instead. This comparison step allows for assessment of the impact of selecting a different experimental ID on the overall aggregate process metric prior to updating the metric. By enabling this comparison, users can make informed decisions about whether to replace an initial experimental ID with an alternative option, based on how that replacement would affect key aggregate metrics like process mass intensity, cost, yield, etc. Only after comparing the aggregate metrics is the aggregate process formally updated to incorporate the selected alternative experimental ID in place of the initial experimental ID. This embodies an optimization approach that empowers users to explore multiple experimental options and select the one that optimizes the aggregate process according to user-defined criteria.
  • the aggregate process metric with the alternative experimental ID may be compared to the aggregate process metric prior to updating the aggregate process metric. This allows for the comparison of the process metrics before and after the replacement of an experimental ID. The results of this comparison may additionally be displayed through the graphical user interface or other visualization methods. By showing the comparison, users can evaluate the impact of selecting a particular alternative experimental ID on the overall aggregate process metric.
  • the aggregate process metric estimated by the process metrics estimator component may be a process mass intensity (PMI).
  • the PMI provides a quantitative indicator of the efficiency of a chemical process by measuring the total mass of materials used in the process per unit mass of product generated. A lower PMI value generally indicates a more efficient process with less waste.
  • the aggregate process metric calculated as a function of the individual process metrics for each experimental ID may be this PMI value, representing the overall mass efficiency of the integrated chemical process under analysis.
  • PMI as the aggregate process metric enables effective assessment and comparison of the sustainability of alternative integrated processes in terms of their mass utilization.
  • the system provides the flexibility to use PMI as the key aggregate metric for optimizing mass flow through the integrated synthesis process.
  • the aggregate process metric estimated by the method may represent one of several sustainability indicators, including solvent intensity, water intensity, or global warming potential. More specifically, in some embodiments, the aggregate process metric calculated as a function of the plurality of estimated process metrics for the individual experimental IDs may be the overall solvent intensity, water intensity, or global warming potential for the integrated process.
  • the aggregate process metric estimated for the integrated process can include cost metrics, yield estimates, and various environmental metrics.
  • the aggregate process metric calculated by the process metrics estimator component may represent overall cost parameters such as total raw material costs, energy costs, and overall process economics. It may also incorporate chemical yield projections and optimizations based on the experimental data.
  • the aggregate process metric can encompass composite environmental impact measures that account for factors like greenhouse gas emissions, wastewater generation, solid waste production, and other sustainability indicators. By consolidating different parameters into one overarching metric, the system aims to provide users with a high-level quantification of the performance, economics, and environmental profile associated with the integrated process. The flexibility to compute aggregate metrics spanning cost, yield, and sustainability domains allows for multi-objective analysis and aids in the identification of optimal process configurations.
  • the method may involve selecting a role for a molecule that is associated with an experimental ID of the one or more experimental IDs. For example, a user may utilize the GUI to select a specific role, such as product, starting material, reagent, solvent, catalyst, etc., for a given molecule linked to an experimental ID.
  • a specific role such as product, starting material, reagent, solvent, catalyst, etc.
  • Associating molecules with particular roles can help provide additional context and clarity regarding how various compounds are being utilized within the integrated process. Defining these molecule roles enables the system to make appropriate assumptions and calculations when estimating the process metrics. The ability to select molecule roles provides an added level of customization, allowing users to tailor the analysis to their specific needs.
  • a process parameter may be adjusted.
  • This adjustment provides users with the capability to modify specific inputs or conditions associated with the integrated process in order to analyze the impacts on the overall process metrics. For example, users can tweak reaction temperatures, reagent quantities, solvent volumes, equipment settings, or other relevant process parameters.
  • at least one of the plurality of estimated process metrics can then be updated accordingly. This update allows users to immediately see how modifications to the process parameters influence metrics like process mass intensity, yield, cost, global warming potential, etc.
  • the method enables iterative optimization of the integrated chemical synthesis process. Users can repeatedly modify parameters and analyze the effects on mass flow, cost, sustainability indicators, and other output values of interest. Overall, the adjustment of process parameters combined with live updating of estimated metrics facilitates customizable scenario analysis and supports data- driven decision making during chemical process development.
  • the method may involve adjusting a process parameter.
  • Some of the experimental IDs may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based on the input parameters. Modifying the input parameters can update these output values.
  • the method may further comprise updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
  • the selection of the one or more experimental IDs from the electronic notebook is performed by a user utilizing a Graphical User Interface (GUI).
  • GUI Graphical User Interface
  • the GUI provides an interface which enables the user to browse, search, or otherwise access the experimental IDs stored within the electronic notebook.
  • the user can then manually select the desired experimental IDs through interactions with GUI elements such as checkboxes, dropdown menus, or search filters.
  • This manual selection initiates the process of linking together the selected experimental IDs into an integrated process for the synthesis of a target material.
  • the GUI therefore facilitates convenient user control over the initial experimental ID selection, while the subsequent linking and analysis steps can be automated by the system based on this user input. Overall, this GUI-enabled selection process allows for customizable construction of integrated processes according to the specific needs and priorities of each user.
  • the method involves querying a prediction engine to determine a prophetic experiment.
  • a prophetic experiment refers to a hypothetical or simulated experiment that has not yet been physically performed.
  • the prediction engine can utilize various techniques to generate predictions or suggestions for potential chemical reactions or processes to be explored.
  • the method then replaces at least one of the initially selected one or more experimental IDs with the prophetic experiment identified by the prediction engine.
  • the method estimates a prophetic process metric corresponding to the prophetic experiment.
  • This prophetic process metric is then incorporated into the aggregate process metric calculation by updating the aggregate process metric to include the prophetic experiment in place of the replaced experimental ID(s).
  • the prediction engine allows hypothetical experiments to be simulated within the overall integrated process, enabling users to explore different process variations and scenarios.
  • the aggregate process metric reflecting the entire integrated process can then be re-estimated based on the incorporation of these prophetic experiments.
  • the method may include a prophetic process metric corresponding to the prophetic experiment determined by a prediction engine.
  • the prophetic process metric includes a range of values, such as a confidence interval, a confidence region, a credible interval, or a credible region. This range provides an indication of the uncertainty or variability associated with the predicted experiment. Accounting for inherent uncertainty when incorporating prophetic experiments can help assess the reliability and robustness of the predictions.
  • the prophetic process metric estimated for the prophetic experiment suggested by the prediction engine may comprise a range of values indicating uncertainty.
  • This range could take various standard statistical forms for representing uncertainty, including a confidence interval, a confidence region, a credible interval, or a credible region.
  • Using these types of value ranges allows users to assess the reliability and variability inherent in the predicted experiments from the prediction engine when evaluating the potential impact on the overall process metrics.
  • the prediction engine can quote a confidence interval, credible region, etc. to quantify the uncertainty in its estimations. This provides greater transparency into the accuracy of the predictive models used by the engine.
  • the plurality of process metrics estimated may be represented as a ratio.
  • the ratio may be process mass intensity (PMI) over yield.
  • PMI process mass intensity
  • Expressing the process metrics as a ratio can provide additional insights into the efficiency and sustainability of the integrated process by examining the relative values between different metrics. A lower ratio value may indicate improved performance, while a higher ratio value may highlight areas needing further optimization.
  • Using customizable ratio metrics allows users to focus the analysis on parameters of greatest relevance to their specific needs and objectives. The ability to calculate and compare ratio-based process metrics adds an extra dimension of flexibility and customization to the overall method.
  • the plurality of process metrics may also be represented as a ratio.
  • this ratio is process mass intensity (PMI) over yield.
  • PMI process mass intensity
  • Using the ratio of PMI over yield as a process metric enables evaluating both the mass efficiency and productivity of the integrated process in a simple, intuitive metric.
  • a lower PMVyield ratio indicates greater efficiency and yield for the overall chemical synthesis pathway. Tracking how modifications to the process impact both PMI and yield through their ratio can further optimization efforts by balancing mass utilization and target output.
  • the use of PMI over yield as a ratio metric provides a sustainability indicator for the integrated process.
  • the one or more experimental IDs may be selected out of order by the user.
  • the act of linking the plurality of experimental IDs includes the act of ordering the selected experimental IDs by the system to properly form the complete integrated process. This integrated process includes all of the corresponding subprocesses associated with the selected experimental IDs, arranged in the correct sequence. Automatically ordering the experimental IDs ensures that the full sequence of subprocesses is configured in the proper order to synthesize the target material. By rearranging the selected experimental IDs into the right process sequence, the system can link subprocesses that may have been originally selected out of order and still generate an integrated process that connects all associated subprocesses to produce the desired target output.
  • the method may involve subprocesses including chemical reactions, purifications, or a combination thereof. More specifically, in some embodiments each experimental ID selected from the electronic notebook corresponds to either a chemical reaction subprocess, a purification subprocess, or a subprocess involving both a chemical reaction and a purification. Therefore, when linking the experimental IDs to form the integrated process for synthesizing the target material, the resulting integrated process may feature multiple subprocesses including chemical reactions, purifications, or a combination of both chemical reactions and purifications. The ability to link together experimental IDs representing diverse types of chemical subprocesses provides flexibility in constructing integrated processes to synthesize desired target materials.
  • the method may involve generating a report that summarizes the estimated process metrics and the aggregate process metric. More specifically, in some embodiments, the method includes an additional step of producing a report that provides an overview of the plurality of process metrics calculated for each experimental ID as well as the aggregate process metric estimated for the integrated process. This report condenses the key information and metrics into a concise summary that allows users to easily review the overall efficiency, sustainability, cost, yield, and environmental impact indicators for the chemical process. By gathering the estimated metrics into a single report, users can assess the process performance holistically rather than examining individual metrics in isolation. The report provides a tool for convenient evaluation and comparison of different process configurations or alternatives explored using the system.
  • the report may present the metrics in various graphical visualizations, such as bar charts, line plots, or Sankey diagrams, to enable intuitive interpretation of the data.
  • the summary report equips users with a quick yet comprehensive perspective of the integrated process and its characterized performance based on the computed process metrics.
  • the method may involve verifying the compatibility of input and output materials when linking the one or more experimental IDs.
  • This verification step acts to ensure continuity in the resulting integrated process that synthesizes the target material.
  • the input materials required for each experimental ID or subprocess are checked against the output materials produced by any preceding experimental IDs or subprocesses in the integrated process flow. Any mismatches in output versus required input materials can indicate possible breaks in continuity of the integrated process, which can then be addressed.
  • an embodiment aims to confirm that the linkage of experimental IDs successfully creates an end-to-end integrated process with no gaps that would prevent the target synthesis.
  • the plurality of process metrics estimated by the process metrics estimator component may include energy consumption metrics.
  • These energy consumption metrics can be calculated based on the input and output materials and the subprocesses utilized in each experimental ID of the integrated process. Specifically, factors such as the quantities and types of input materials, chemical transformations involved in the subprocesses, output materials produced, reaction conditions like temperature and pressure, and other relevant parameters can be used to estimate the energy consumption associated with each subprocess and experimental ID. Advanced techniques like computational fluid dynamics simulations, thermodynamic calculations, or empirical correlations based on historical data may be leveraged to quantify the energy consumption. By aggregating the energy consumption across all experimental IDs, an overall energy usage estimate can be obtained for the integrated process. This facilitates analysis of the environmental impact and cost-effectiveness of the overall synthesis route. The energy consumption metrics, along with the other estimated process metrics, allow users to make informed decisions when optimizing the integrated process.
  • the method may include a step of validating the estimated process metrics against predetermined criteria before updating the aggregate process metric.
  • the estimated process metrics such as process mass intensity, solvent intensity, water intensity, and other metrics
  • the validation ensures accuracy and reliability of the metrics prior to using them to update the overall aggregate process metric. If the estimated process metrics for the individual experimental IDs meet the predetermined validation criteria, then they can be applied to update the combined aggregate metric for the full integrated process. However, if one or more process metrics fail to satisfy the validation checks, the method may require troubleshooting, metric re-estimation, or other corrective measures before the aggregate metric is updated.
  • This validation step may act as a quality check on the estimated metrics, enhancing the reliability and robustness of the overall mass flow calculations and sustainability analysis enabled by the method.
  • machine learning models are utilized for the searching for similar reaction subprocesses.
  • the method involves searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, as previously described.
  • the search can make use of machine learning algorithms and models that are trained on historical experimental data to identify patterns and similarities between experiments. These models enable the search to suggest relevant alternative experiments based on the provided experimental IDs.
  • the utilization of machine learning models can expand the capabilities of the search and provide more comprehensive results, optimizing the process metrics and sustainability indicators.
  • the method may include the step of alerting a user when the aggregate process metric exceeds a predetermined environmental impact threshold. More specifically, the aggregate process metric estimated as a function of the plurality of process metrics can be compared against a defined threshold representing the maximum allowable environmental impact. If the aggregate process metric is calculated to exceed this predetermined threshold, the system may generate and display an alert to notify the user.
  • This alert feature enables users to recognize when certain sustainability or environmental impact constraints are violated due to excessive resource consumption or emissions associated with the integrated chemical synthesis process under analysis. The ability to define environmental impact limits and automatically trigger notifications when these constraints are surpassed can aid in adhering to sustainability guidelines and minimizing ecological footprints.
  • each experimental ID may be associated with specific equipment used in the corresponding subprocess.
  • a reaction subprocess may utilize a particular reactor or purification equipment.
  • the method may further comprise adjusting the settings or parameters of this equipment based on the requirements of that subprocess. For instance, if a reaction subprocess operates at a certain temperature and pressure, the reactor settings can be automatically adjusted to match those conditions. This automatic adjustment of equipment settings streamlines the experimental workflow and ensures alignment between the subprocess details contained within the experimental ID and the actual equipment configuration used to carry out that subprocess. By linking the experimental IDs to the relevant equipment in this manner, the method enables improved standardization, repeatability, and optimization of the subprocesses.
  • the method may further involve automatically ordering chemicals and materials needed for the subprocesses.
  • This automatic ordering is based on the input materials listed in the experimental IDs that were selected from the electronic notebook.
  • the system can determine the required reagents, solvents, catalysts, and other chemicals needed to carry out each subprocess. It can then automatically generate purchase orders or materials requests to obtain the necessary supplies, ensuring analysts and researchers have the required ingredients on hand before commencing the chemical reactions or purifications encompassed within that subprocess.
  • This just-in-time materials ordering facilitated by the automated system helps minimize inventory and procurement overhead for organizations frequently synthesizing new target compounds or materials.
  • the synthesis tree provided to the user may include multiple alternative synthesis pathways that could be used to synthesize the target material.
  • the method may provide the capability for the user to select one of these alternative synthesis pathways based on user-defined criteria or preferences. For example, the criteria could be related to optimizing particular process metrics like cost, environmental impact, or yield.
  • the user interface allows the user to input these criteria and priorities.
  • the method then facilitates the selection of the optimal synthesis pathway that best meets the specified criteria out of the alternatives included in the synthesis tree visualization. Enabling users to explore alternative synthesis routes and choose based on customizable metrics provides flexibility and aids in process optimization.
  • the method may further involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs.
  • This energy consumption metric can provide an indication of the overall energy requirements associated with the synthesis of the target material via the linked subprocesses.
  • empirical correlations, thermodynamic calculations, or other techniques may be applied using parameters from the experimental data corresponding to each experimental ID. Such parameters can include temperature, pressure, flow rates, batch sizes, and other relevant factors that influence energy usage.
  • parameters can include temperature, pressure, flow rates, batch sizes, and other relevant factors that influence energy usage.
  • the method may involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs.
  • Estimating the energy consumption metric can comprise applying empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs. For example, thermodynamic calculations and empirically gained correlations may be used to estimate the energy uptake of individual process steps, such as heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction.
  • This approach leverages established techniques to assess the energy requirements and associated environmental impacts of the subprocesses linked to form the integrated process.
  • the method may involve estimating a carbon footprint metric for the integrated process.
  • the carbon footprint metric may be estimated based on the one or more experimental IDs that were selected from the electronic notebook and linked to form the integrated process.
  • the carbon footprint estimation may utilize life cycle inventory data for the various input materials involved in the subprocesses and integrated process.
  • life cycle inventory data for the various input materials involved in the subprocesses and integrated process.
  • the method further comprises providing a visualization of mass flow through the integrated process based on the one or more experimental IDs that were selected from the electronic notebook.
  • This visualization may utilize graphical tools and diagrams that allow users to see how mass flows through each subprocess and the overall integrated process used to synthesize the target material. For example, Sankey diagrams or other types of flow charts could be generated to intuitively display quantities and mass balances across different stages of the chemical synthesis process defined by linking the experimental IDs. Enabling intuitive visualization of mass flow helps users identify areas of inefficient material usage and opportunities for reducing waste or environmental impact.
  • the visualization component allows users to visually explore the mass flow impacts of selecting alternative experimental IDs or adjusting process parameters within the integrated process.
  • the method includes a step of visualizing the mass flow through the integrated process based on the one or more experimental IDs.
  • This visualization of the mass flow may employ a Sankey diagram, which allows an intuitive graphical representation of flows and their quantitative values within a system.
  • the Sankey diagram can illustrate the sequential flow of materials, energy transfers, or waste through each subprocess and throughout the overall integrated process.
  • the thickness of the arrows in the Sankey diagram represents the magnitude or amount of mass flow.
  • This Sankey diagram visualization provides users with an easily interpretable overview of how mass flows through and is transformed within the integrated chemical synthesis process under analysis.
  • the visualization supports identification of mass intensive steps, recycling opportunities, yield losses, and other relevant mass flow characteristics, aiding in process assessment, optimization, and improvement.
  • the system may allow for the synthesis tree, which corresponds to the one or more experimental IDs, to be edited by the user through a graphical user interface. More specifically, in some embodiments the system includes functionality enabling user modification of the visualization of the overall synthesis route. This can facilitate optimization, process adjustments, or exploring alternative synthesis pathways.
  • the act of searching for similar reaction subprocesses includes identifying subprocesses that have the same input and output molecular structures as a given experimental ID from among the one or more experimental IDs. Specifically, when searching for similar subprocesses, those processes that match the molecular structure of both the input material and the output material for a particular experimental ID can be recognized. By finding reaction subprocesses with identical input and output chemicals, alternative synthetic routes or optimizations for specific reaction steps within the integrated process may be determined. This approach focuses the search on processes that are highly comparable in terms of the chemical transformations being carried out.
  • the method involves estimating a range of values for at least one of the plurality of process metrics based on the plurality of similar experiments identified in the search. More specifically, after searching for and determining similar reaction subprocesses and experiments for each experimental ID, a range of values may be calculated for process metrics like process mass intensity, solvent intensity, yield, cost, or other metrics. This range represents the variability in the metric across the multiple similar experiments found that correspond to the experimental ID(s) selected by the user. Providing such a range gives an indication of the uncertainty and potential fluctuation in the process metric, allowing for a more comprehensive analysis when evaluating and optimizing the integrated chemical synthesis process. The range could take the form of a confidence interval, confidence region, credible interval, credible region, or other representation of variability. This enhanced uncertainty quantification through metric value ranges enables more informed decision making during chemical process development.
  • Some embodiments of the method may involve suggesting an alternative solvent or an alternative reagent for at least one of the experimental IDs based on the plurality of similar experiments found by the search component.
  • the search component identifies experiments with similar reactions and can determine if there may be better solvents or reagents that could be used for an experimental ID by analyzing the reagents and solvents used in those similar experiments.
  • the system can recommend testing alternate solvents or reagents for an experimental ID that may improve yield, lower cost, reduce waste, or provide other benefits over the original solvent or reagent selected. This allows researchers to easily get suggestions on alternate reaction conditions to try that could optimize their process.
  • the method allows for the selection of one or more experimental IDs from an electronic notebook, wherein the selection is based on a target molecule specified by a user.
  • a user may first specify or input a desired target molecule they wish to synthesize. Based on this target molecule, relevant experimental IDs can then be retrieved and selected from the electronic notebook that correspond to subprocesses involved in the production or synthesis of the specified target material.
  • the target molecule provides a means for the user to indicate what final product they intend to make, while the system identifies and selects the necessary reaction steps and corresponding experimental IDs that can lead to the target molecule.
  • This target molecule-based selection of experimental IDs facilitates the process of linking subprocesses into an overall integrated process for synthesizing the user- specified target material. Overall, enabling the selection of experimental IDs based on a user- defined target molecule allows the system to automatically identify the building blocks needed for a user's molecule of interest based on available experimental data.
  • Some embodiments provide the capability of tracking modifications made to the integrated process over time, as well as the resulting impacts on the aggregate process metric. Specifically, as changes or optimizations are made to the linked experimental IDs that comprise the integrated process, the system can monitor and log these modifications. Concurrently, the impacts of the changes on process metrics like the overall process mass intensity, solvent intensity, yield, cost, or other aggregate sustainability indicators can be reestimated and recorded. By tracking edits to the process flow alongside corresponding shifts in the process metrics, users can evaluate how impactful or beneficial particular process tweaks have been historically. This logging functionality enables data-driven analysis of how the integrated process has evolved regarding sustainability and can guide future optimization efforts. Overall, the ability to trace modifications and quantify associated effects on the aggregate process metric can facilitate systematic improvements over multiple iterations.
  • the method may further involve estimating the impacts of solvent recycling or waste stream recycling on the plurality of process metrics.
  • the process metrics estimator component estimates how recycling solvents or waste streams back into the integrated process influences metrics such as process mass intensity, solvent intensity, water intensity, and global warming potential. For example, solvent recycling can reduce the amount of fresh solvent utilized in the process, thereby lowering the solvent intensity. Similarly, recycling waste streams may decrease the material inputs needed, which can positively impact sustainability indicators like process mass intensity. Quantifying these recycling impacts provides additional insights into optimization opportunities for improving the environmental performance of the integrated process.
  • the method may involve integrating the estimated process metrics and aggregate process metric with an enterprise resource planning (ERP) system.
  • ERP enterprise resource planning
  • the ERP system can provide comprehensive data management capabilities to track materials, processes, inventory, orders, accounting, and other operational data across the enterprise. Integrating the sustainability metrics estimated by the system with the ERP allows for holistic tracking of environmental factors alongside traditional business metrics. This integration enables seamless monitoring and optimization of both process efficiency and environmental impact through a centralized platform. Overall, incorporating the calculated process metrics and aggregate sustainability indicators into an organization's broader ERP infrastructure can facilitate comprehensive data analysis to inform sustainable decision-making across chemical research and manufacturing activities.
  • the method involves storing the one or more experimental IDs, the plurality of process metrics, and the aggregate process metric in a database.
  • This database storage allows the experimental data, process metrics, and aggregate metrics to be saved for later retrieval and analysis. By storing this information, users can access a historical record of prior experiments, process metrics calculations, and overall sustainability indicators.
  • the database also enables data sharing, collaboration between multiple users, integration with other systems, and long-term tracking of process improvements over time.
  • the stored data may be retrieved at a later point from the database for additional analysis, comparison to alternative processes, or to demonstrate progress in process optimization efforts. Overall, the database storage provides persistence and easy access to valuable experimental records and sustainability metrics.
  • the method may further involve retrieving the one or more experimental IDs that were stored in the database, along with the associated plurality of process metrics and the aggregate process metric. This allows a user to revisit and analyze previous experimental data that had been saved, enabling the comparison of different process options or tracking changes in process metrics over time.
  • the storage and retrieval capabilities facilitate data management and reuse, avoiding duplication of effort and taking full advantage of experimental data from past synthesis routes or process configurations.
  • users can conveniently restore experimental IDs and process metrics that pertain to earlier experiments or process alternatives.
  • the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
  • the act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material.
  • the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
  • this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs.
  • the plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
  • the method may further involve updating the plurality of process metrics and the aggregate process metric based on edited input data associated with at least one of the selected experimental IDs.
  • the system enables users to manually edit the input data corresponding to the experimental IDs through the graphical user interface. For example, a user could modify the quantities of reagents used in a reaction or change the reaction conditions like temperature and pressure.
  • the process metrics estimator component automatically recalculates and updates the plurality of process metrics to reflect the changes in input data. This would update metrics like process mass intensity, solvent intensity, yield percentage, etc. in accordance with the manual edits.
  • the aggregate process metric is a function of the individual process metrics, it is also updated accordingly when changes are made to the input data. This dynamic update allows users to immediately see the impact of any input data modifications on the overall process metrics and sustainability indicators. It facilitates rapid scenario testing and process optimization during chemical process development.
  • the method includes displaying changes between the edited input data that the user manually adjusted and the original input data that was initially associated with at least one experimental ID.
  • Visually highlighting the changes between the original and edited input data or providing a side-by-side comparison can make it easier for users to see where modifications have been made and understand the impacts on the estimated process metrics and aggregate process metric. This capability enhances transparency and facilitates iterative optimization of the integrated process through assessment of different input data scenarios.
  • the method includes enabling a user to manually edit input data associated with at least one of the one or more experimental IDs.
  • the plurality of process metrics and aggregate process metrics are updated based on the edited input data.
  • Changes between the edited input data and the original input data associated with the at least one experimental ID may be displayed. Displaying the changes may involve visually highlighting the differences between the edited and original input data within the user interface. Alternatively, displaying the changes may involve providing a side-by-side comparison of the original and edited input data, allowing the user to clearly see how the data has been modified.
  • the plurality of estimated process metrics may be provided for the integrated process as a whole, for parts of the integrated process, or for each individual subprocess step corresponding to the respective experimental IDs. That is, the process metric estimates generated by the process metrics estimator component can be presented at different levels of granularity depending on the specific requirements. For example, a cumulative or aggregate PMI value may be calculated and displayed for the complete multi-step integrated process to synthesize the target material. Alternatively, separate PMI values may be estimated and shown for each distinct subprocess, giving insights into the PMI contributions of individual steps. As another option, PMI metrics may be calculated and visualized for logical subsections of the integrated process, such as a sequence of reactions or a particular purification train. This flexibility in process metric estimation and visualization at multiple levels enables detailed analysis and comparison of different process options.
  • the process metrics estimated for the integrated process may be compared to one or more process metrics from other processes.
  • this comparison to external process metrics can be performed in order to facilitate evaluation of the relative performance of the integrated process under consideration.
  • the comparison can provide additional context and enable easier assessment of the efficiency, cost-effectiveness, sustainability, or other relevant parameters associated with the integrated process synthesized from the selected experimental IDs. The capability to draw such comparisons against suitable benchmarks expands the utility of the estimated process metrics and aggregate process metrics calculated by the system.
  • the method may further involve providing the plurality of process metrics for each individual subprocess within the integrated process formed by linking the experimental IDs.
  • the process metrics estimator component can estimate process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
  • process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
  • process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user.
  • the method includes comparing at least one process metric of the plurality of process metrics to a corresponding process metric from a different integrated process. For example, a process mass intensity or solvent intensity value estimated for one of the linked experimental IDs may be compared to process mass intensity or solvent intensity values obtained from a separate, distinct integrated process. This enables assessing the relative efficiency or sustainability of the process under consideration compared to alternative processes for producing the same or similar target compounds.
  • the ability to benchmark against historical data or industry standards supports evaluating opportunities for improving the integrated process through suitable modifications to reaction conditions, workup procedures, or selection of starting materials and reagents.
  • At least one process metric of the plurality of process metrics may be visualized using a Sankey diagram.
  • the Sankey diagram provides an intuitive way to represent the flow of mass, energy, cost, or other metrics through the integrated process. This visualization can help highlight inefficiencies, mass imbalances, and opportunities for process improvement in a graphical manner.
  • users can gain insight into how changes in one subprocess may propagate through the integrated process to impact other metrics of interest.
  • Some implementations may allow users to interact with the Sankey diagram representation of the process metrics, for example by selecting specific pathways or modifying process parameters to observe the effect on the overall mass flow in real time.
  • the Sankey diagram enables an interactive, graphical analysis that complements the detailed numeric process metrics estimated by the system.
  • the method includes identifying recycled mass streams that are recycled to the beginning of a subprocess or the integrated process.
  • the act of searching for similar reaction subprocesses by the search component may involve recognizing mass streams from waste or byproducts that are recycled and fed back into an earlier subprocess or the start of the overall integrated process. By detecting these recycled mass streams, the system can analyze the impact of recycling on the estimated mass flow and process metrics. This recycling functionality provides another tool to explore optimization opportunities and improve the efficiency of the chemical synthesis process under analysis.
  • the search component is configured to identify these recycled streams within the historical data and suggest integration opportunities accordingly.
  • the range of values estimated for the prophetic process metric may correspond to a multi-step process rather than just a single reaction step.
  • the predicted process metrics for this overall pathway may include confidence intervals or credible regions. These uncertainty ranges can account for the inherent variability when predicting outcomes across several chemical transformations or over an extended process with several distinct steps. By propagating uncertainties across the various stages, reliable estimates for the overall reliability of the predicted process metrics and sustainability indicators can be provided even for complex, integrated processes. This allows users to realistically assess the accuracy of recommendations from in silico tools when dealing with intricate, multi-stage synthetic routes.
  • the method involves automatically linking the one or more experimental IDs to form the integrated process synthesizing the target material.
  • This automatic linkage is performed based on the input and output chemical structures specified in each of the experimental IDs. That is, the software analyzes the chemical structures entering and exiting each subprocess to determine compatibility and continuity of materials flow. It then automatically connects the experimental IDs in the proper order to construct the overall integrated process including all necessary subprocesses, without requiring extensive manual input or oversight from the user.
  • the linking process can be streamlined and expedited. This automation enables more rapid set up and estimation of mass flow and sustainability metrics for the chemical synthesis route.
  • the plurality of process metrics are represented as a ratio.
  • the ratio can be process mass intensity over yield.
  • the process mass intensity reflects the mass-based resource consumption per unit mass of product, while the yield represents the amount of product obtained from the process.
  • Taking the ratio of these two metrics provides a useful indicator of the efficiency and environmental impact of the integrated process, accounting for both the resource usage and productivity.
  • a lower ratio indicates a more efficient and sustainable process.
  • users can assess the impacts of modifications or alternate pathways on the overall process performance.
  • representing the plurality of process metrics as a process mass intensity over yield ratio enables straightforward yet comprehensive evaluation of chemical synthesis processes.
  • the system may include both hardware and software components configured to execute the computer-implemented method for calculating mass flow and estimating sustainability metrics.
  • the system may include one or more processors, memory, storage, network interfaces, databases, and other computing resources capable of carrying out the various steps involved in selecting experimental IDs, linking them to form an integrated process, estimating process metrics, providing a synthesis tree, searching for similar reactions, selecting alternative experimental IDs, and updating aggregate metrics.
  • the system is designed to implement the full functionality of the method efficiently through specialized algorithms, predictive models, data structures, or other technological means.
  • the modular, customizable architecture of the system also allows for easy extension or updating of capabilities in accordance with advancements in the field. Overall, embodiments provide an intelligent data processing system leveraging automation and advanced analytics to revolutionize mass flow calculations, process optimization, and sustainability assessment during chemical process development.
  • the disclosure includes a computer program containing instructions that, when executed by a computer, cause the computer to perform the method described herein.
  • This computer program allows the computer to select one or more experimental IDs, link the IDs to form an integrated process, estimate process metrics for each ID, calculate an aggregate process metric, provide a synthesis tree, search for similar reactions, select alternative IDs, and update metrics accordingly.
  • the computer can fully implement the method for calculating mass flow and estimating sustainability metrics in chemical processes.
  • the computer program enables automated calculation of metrics, exploration of alternatives, and optimization of chemical synthesis procedures in a computerized environment.
  • the method may be implemented as a computer program with instructions stored on a computer-readable medium.
  • these instructions When these instructions are executed by a computer, they cause the computer to carry out the method comprising: selecting one or more experimental IDs from an electronic notebook, linking the IDs to form an integrated process synthesizing a target material, estimating process metrics for each ID, calculating an aggregate metric, providing a synthesis tree, searching for similar reactions, selecting alternative IDs, and updating the aggregate metric accordingly.
  • the computer-readable medium allows the computational implementation of the method, enabling the computer to perform the selection, linking, estimation, calculation, provision, searching, alternative selection, and updating involved in analyzing mass flow and sustainability metrics for chemical processes.
  • a computer can carry out the full method by accessing and executing the appropriate program from the computer-readable medium on which it resides.
  • FIG. 1 shows a block diagram illustration of a cloud-based system to calculate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure
  • FIG. 2 show a block diagram illustration of a computing device to estimate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure
  • Fig. 3 is a flow-chart diagram of a method of estimating mass flow and estimate sustainability metrics in accordance with an embodiment of the present disclosure
  • Fig. 4 shows a GUI of linked experimental IDs in accordance with an embodiment of the present disclosure
  • FIG. 5 shows a summary of the output of the of linked experimental IDs of Fig. 4 in accordance with an embodiment of the present disclosure
  • FIG. 6 shows a synthesis tree in accordance with an embodiment of the present disclosure
  • Fig. 7 shows an example result of the search component, in which the input data can be modified in accordance with an embodiment of the present disclosure
  • FIGs. 8A and 8B show an example result of the search component in which the mass flow is visualized in accordance with an embodiment of the present disclosure
  • Fig. 9 shows an example result of the search component in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 6 in accordance with an embodiment of the present disclosure
  • Fig. 10 shows PMI vs. Yield values for each alternative experimental ID that the search component identifies in accordance with an embodiment of the present disclosure
  • Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow is estimated in accordance with an embodiment of the present disclosure
  • Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure
  • Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations in accordance with an embodiment of the present disclosure
  • Fig. 14 shows model compounds for raw materials to synthesize APIs in accordance with an embodiment of the present disclosure
  • Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure
  • Fig. 16 shows an average PMI and chemical yield for similar reactions of one synthesis step that was suggested by a retro-synthesis tool using data is obtained using the similarity search function in accordance with an embodiment of the present disclosure
  • Fig.17 shows the Input metric for the energy calculation module in accordance with an embodiment of the present disclosure
  • Fig.18 shows the output metric for the energy calculation module in accordance with an embodiment of the present disclosure.
  • Fig. 19 shows the GUI displaying weight based carbon footprints received from an external database, which are assigned based on data of individual raw materials, on a general role, or manually edited data in accordance with an embodiment of the present disclosure.
  • Fig. 1 shows a block diagram illustration of a cloud-based system 100 to estimate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure.
  • the system 100 automatically calculates metrics without significant manual interaction or input.
  • the system 100 includes a cloud service provider 102, one or more personal computers 104, and a mobile device 106.
  • the system 100 also includes a mass Flow-estimator component 112.
  • the mass flow-estimator component 112 can estimate mass flow in chemical synthesis and enable a user to modify the chemical synthesis process to change or reduce the mass flow as described herein.
  • the mass flow estimator component 112 can be used in chemical process development to optimize certain values, such as costs, yield or environmental compatibility.
  • these may be calculated early in the development and updated regularly.
  • These data may be stored in a data pool, e.g., in subprocess metrics 144 of a database 132.
  • the database 132 may be accessed via standard interfaces, such as Oracle DB interfaces.
  • the system 100 can be used for estimating mass flow in a chemical synthesis or an industrial process.
  • the system 100 may be used in the pharmaceutical industry to estimate the mass flow in drug manufacturing processes, may be used in the chemical industry to estimate the mass flow in the production of various chemical compounds or formulations, and/or may be used in the biotechnology industry to estimate the mass flow in the production of biologies.
  • the cloud service provider 102 may be configured to provide remote capabilities to estimate mass flow in a chemical synthesis or industrial process by providing remote access to the mass-flow estimator component 112.
  • the cloud service provider 102 may be a hosted service such as a company that offers cloud computing services to businesses and individuals such that the cloud service provider 102 provides the infrastructure, software, and platforms required to host, manage, and deliver cloud-based services.
  • the cloud service provider 102 may provide infrastructure as a service, platform as a service, software as a service, and/or may be an interface into a blockchain infrastructure that may or may not be hosted by the cloud service provider 102.
  • the cloud service provider 102 may be configured to scale up or down its computing resources based upon demand from users at a given moment.
  • the cloud service provider 102 may be implemented on a block chain.
  • the cloud service provider 102 may utilize a distributed ledger to store and verify mass flow estimates generated by the mass-flow estimator component 112. Users may be authenticated and/or authorized by a secure digital certificate, encryption key, single-sign on, or other secure mechanism.
  • the data may be stored and calculated in a secure and tamper proof manner to provide transparency and accountability to all users.
  • the smart contracts may include executable code that defines a manufacturing process in terms of one or more or the components within the mass-flow estimator component 112 in a manner consistent with transparency and security settings.
  • the personal computers 104 and the mobile device 106 communicate with each other via a network 108.
  • the network 108 may be Wi-Fi, ethernet, Bluetooth, etc. and may utilize the internet and associated protocols, such as TCP/IP.
  • the network 108 may be a local area network, a wide-area network, a physical bus (such as a Universal Serial Bus), the internet, or some combination thereof.
  • the personal computer 104 and mobile device 106 may interface with the cloud service provider 102 to the calculate mass flow in a chemical or industrial process as determined by the subprocess metrics 144.
  • a specialized application for interfacing with the mass-flow estimator component 112 may be used, such as a mobile application on the mobile device 106 or a desktop application on the personal computer 104.
  • the communications may include transmitting data in HTML, XML, JSON, YAML, or any data format.
  • the mass-flow estimator component 112 may provide user-level accounts to individuals through a typical login mechanism.
  • the mass-flow estimator component 112 may be a web application, a webserver, a web service, etc. and may utilize one or more protocols to communicate data.
  • the cloud service provider 102 may provide the mass-flow estimator component 112 as a webpage, a webapp, a program for download and execution on the computer 104 or the mobile device 106.
  • the mass-flow estimator component 112 includes various sub-components, such as an electronic notebook component 150, a component for linking experimental IDs 142, a process metrics estimator component 152, a synthesis tree component 156, a search component 158, a GUI component 160, and a prediction engine 164.
  • the mass-flow estimator component 112 of the system 100 may be implemented as a cloud-based software application accessible via a web interface.
  • the system 100 may also be implemented as a standalone software application installed on a local computer or server.
  • the system 100 may include specialized hardware such as a high-performance computer or server, or other types of computing and networking equipment.
  • the mass flow calculator component 112 of the system 100 may be implemented on a variety of computing platforms, including desktop, laptop, server, or cloud-based configurations.
  • the software may be written in various programming languages, such as Python, Julia, Java, C++, or other languages.
  • the system 100 may also incorporate various data visualization tools, such as d3.js, Plotly.js, or Matplotlib, to allow users to visualize the experimental data and the process metrics in an intuitive manner.
  • the electronic notebook component 150 enables the selection of a one or more experimental IDs 142 that are stored within a database 132.
  • the electronic notebook component 150 may use a REST API and may be accessed via the Mass flow calculator component.
  • Each experimental ID 142 corresponds to a subprocess having an input material and an output material.
  • Information about each experimental ID 142 may be found in the subprocess metrics 144 also included in the database 132.
  • the subprocess metrics 144 may include historical data such that each experimental ID 142 can be associated within one or more historical experiments as found within the subprocess metrics 144 to estimate or predict a metric.
  • the linking component 154 links the one or more experimental IDs 142 to form an integrated process including all of the corresponding subprocesses, thereby synthesizing a target material.
  • the process metrics estimator component 152 estimates a plurality of process metrics (e.g., using the subprocess metrics 144), each of which corresponds to a respective experimental ID of the one or more experimental IDs 142. Fig. 4 and Fig 5.
  • the process metrics estimator component 152 also estimates an aggregate process metric that is a function of the estimated plurality of process metrics (which can be column summations).
  • the process metrics estimator component 152 can estimate various process metrics related to the chemical synthesis process. This component can be used for evaluating the efficiency, sustainability, and environmental impact of the integrated process formed by linking the selected experimental IDs.
  • the process metrics estimator component 152 estimates a plurality of process metrics, where each of these metrics corresponds to a respective experimental ID from the one or more experimental IDs selected by the user.
  • process metrics can encompass a wide range of parameters, including but not limited to process mass intensity (PMI), solvent intensity (SI), water intensity (WI), global warming potential (GWP), cost, yield, and various environmental metrics.
  • the process metrics estimator component 152 may calculate these metrics based on the experimental data associated with each experimental ID, which may include information from one or more previously performed experiments. This data can be retrieved from the subprocess metrics 144 stored in the database 132 or obtained from other relevant sources.
  • the process metrics estimator component 152 may employ advanced algorithms, machine learning techniques, or empirical models to estimate the process metrics accurately. It may consider factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other relevant parameters to derive these metrics.
  • the process metrics estimator component 152 may be used to estimate an aggregate process metric corresponding to the integrated process.
  • This aggregate process metric is a function of the estimated plurality of process metrics for the individual experimental IDs.
  • the function used to calculate the aggregate process metric can vary depending on the specific requirements and the nature of the process being analyzed.
  • the aggregate process metric may be a simple summation of the individual process metrics. For example, if the process metrics being considered are PMI values for each subprocess, the aggregate process metric could be the sum of these PMI values, representing the overall PMI for the integrated process.
  • the aggregate process metric may be a weighted combination of the individual process metrics, where different weights are assigned to different metrics based on their relative importance or impact on the overall process.
  • the aggregate process metric could be a weighted sum of PMI, SI, WI, and GWP, reflecting the combined environmental impact of the integrated process.
  • the aggregate process metric may be a more complex function that incorporates additional factors or constraints.
  • the aggregate process metric could be a multi-objective optimization function that considers not only the process metrics but also factors such as cost, yield, or specific environmental targets.
  • the process metrics estimator component 152 may also provide the capability to compare the estimated process metrics and the aggregate process metric with corresponding metrics from other processes or industry benchmarks. This comparison can aid in evaluating the relative performance and sustainability of the integrated process under consideration.
  • the process metrics estimator component 152 may interact with other components of the system 100, such as the search component 158 and the prediction engine 164, to explore alternative experimental IDs or prophetic experiments. When an alternative experimental ID or a prophetic experiment is selected, the process metrics estimator component 152 can update the aggregate process metric accordingly, reflecting the impact of the proposed change on the overall process metrics.
  • the process metrics estimator component 152 may provide visualizations or graphical representations of the estimated process metrics and the aggregate process metric. These visualizations can aid in interpreting the data and identifying areas for potential improvement or optimization.
  • the process metrics estimator component 152 can be designed with a modular architecture, allowing for the integration of various modules or sub-components responsible for estimating specific process metrics. This modular approach enables flexibility, scalability, and customization, as different modules can be added, removed, or updated independently based on the specific requirements or advancements in the field.
  • one module within the process metrics estimator component 152 could be dedicated to estimating process mass intensity (PMI) and solvent intensity (SI).
  • This module may leverage advanced algorithms and machine learning techniques to analyze the input and output materials, reaction conditions, and solvent usage data associated with each experimental ID. It may also incorporate industry-specific heuristics or empirical models to improve the accuracy of PMI and SI estimations.
  • Another module could focus on estimating water intensity (WI) and global warming potential (GWP).
  • This module may integrate with external databases or life cycle assessment (LCA) tools to obtain relevant data on the environmental impact of various materials and processes involved in the chemical synthesis. It could also employ computational fluid dynamics (CFD) simulations or thermodynamic calculations to estimate the energy consumption and associated GWP contributions.
  • CFD computational fluid dynamics
  • Yet another module within the process metrics estimator component 152 could be responsible for estimating cost-related metrics, such as raw material costs, energy costs, and overall process costs.
  • This module may interface with enterprise resource planning (ERP) systems, supply chain management systems, or market data feeds to obtain up-to-date information on material prices, energy costs, and other relevant cost factors.
  • ERP enterprise resource planning
  • the process metrics estimator component 152 may incorporate advanced uncertainty quantification techniques to provide confidence intervals, credible regions, or probability distributions for the estimated process metrics. These uncertainty estimates can be particularly valuable when dealing with incomplete or uncertain input data, or when accounting for inherent variability in the chemical synthesis processes.
  • the process metrics estimator component 152 may also be integrated with optimization algorithms or decision support systems.
  • the estimated process metrics could serve as objective functions or constraints in the optimization process, enabling the identification of optimal process parameters, reaction conditions, or material selections that minimize undesirable metrics (e.g., PMI, GWP) while maximizing desirable ones (e.g., yield, cost-effectiveness).
  • undesirable metrics e.g., PMI, GWP
  • desirable ones e.g., yield, cost-effectiveness
  • process metrics estimator component 152 could be designed to leverage distributed computing resources or cloud-based platforms. This would allow for parallel processing of multiple experimental IDs or the execution of computationally intensive simulations or calculations required for estimating certain process metrics. Cloud-based deployment could also facilitate collaborative efforts, where multiple users or research teams can contribute experimental data and access the estimated process metrics from various locations.
  • the process metrics estimator component 152 may incorporate machine learning capabilities to continuously improve its estimation accuracy. As more experimental data becomes available, the component could retrain its models or update its algorithms, leveraging techniques such as transfer learning or active learning to enhance the estimation performance over time.
  • the process metrics estimator component 152 could also be integrated with advanced visualization tools or dashboards, providing intuitive and interactive representations of the estimated process metrics. These visualizations could include interactive charts, graphs, Sankey diagrams, or even virtual reality (VR) or augmented reality (AR) environments, allowing users to explore the data from different perspectives and gain deeper insights into the chemical synthesis process.
  • advanced visualization tools or dashboards providing intuitive and interactive representations of the estimated process metrics.
  • These visualizations could include interactive charts, graphs, Sankey diagrams, or even virtual reality (VR) or augmented reality (AR) environments, allowing users to explore the data from different perspectives and gain deeper insights into the chemical synthesis process.
  • the process metrics estimator component 152 may support the integration of user-defined metrics or custom calculations. This would enable users to incorporate domain-specific knowledge, proprietary algorithms, or specialized requirements into the estimation process, tailoring the component to their specific needs or industry standards.
  • Fig. 4 may be values for a particular batch run or for a per kilogram of target material or may use any other unit of measurement.
  • Fig. 5 is the GUI showing a summary of the outputs of the linked experimental IDs of Fig. 4.
  • the synthesis tree component 156 provides a synthesis (illustrated in Fig. 6) as a synthesis tree 500 (as shown in the GUI in Figs. 7) corresponding to the one or more experimental IDs 142.
  • the search component 158 is configured for searching for structurally similar reaction subprocesses for each of the experimental IDs 142 to determine a plurality of similar experiments.
  • the search component 158 may be configured to search for similar reaction subprocesses for each of the experimental IDs 142 stored in the database 132. This search can help determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the selected one or more experimental IDs 142.
  • the search component 158 may identify similar experiments by looking for experiments with the same primary input molecule and target molecule as one or more of the selected experimental IDs 142.
  • the search component 158 may ignore or consider helper molecules, such as buffers or catalysts, when identifying similar experiments.
  • the search component 158 can facilitate the selection of an alternative experimental ID 142 for at least one of the initially selected one or more experimental IDs 142. This selection of an alternative experimental ID 142 can be made through the GUI component 160, which may display the search results from the search component 158.
  • the process metrics estimator component 152 can then update the aggregate process metric in accordance with the selected alternative experimental ID 142.
  • the search component 158 may identify recycled mass streams that are recycled to the beginning of a subprocess or the integrated process. This information can be used to estimate the effect of solvent recycling on the mass flow and process metrics, as illustrated in Fig. 11 of the patent application. [0039]
  • the search component 158 may utilize various techniques to identify similar experiments, such as chemical structure matching, reaction type matching, or machine learning models trained on historical experimental data. The search component 158 can provide a tool for process optimization by allowing users to explore alternative experimental IDs 142 and their corresponding process metrics, enabling informed decision-making during chemical process development.
  • the search component 158 may work in conjunction with other components of the mass-flow estimator component 112, such as the synthesis tree component 156, which provides a visual representation of the linked experimental IDs 142, and the prediction engine 164, which can suggest prophetic experiments based on predictive models or retrosynthesis tools.
  • the mass-flow estimator component 112 such as the synthesis tree component 156, which provides a visual representation of the linked experimental IDs 142, and the prediction engine 164, which can suggest prophetic experiments based on predictive models or retrosynthesis tools.
  • the search component 158 may employ chemical structure matching algorithms to identify experimental IDs 142 that involve the same input and output molecular structures. This approach can be particularly useful when the goal is to find alternative synthetic routes or optimize specific reaction steps within the integrated process.
  • the search component 158 may maintain a database or index of chemical structures associated with each experimental ID 142, enabling efficient querying and matching based on structural similarity metrics. These metrics can consider factors such as functional groups, bond connectivity, stereochemistry, and other relevant structural features.
  • the search component 158 can utilize well-established chemical structure representation formats, such as SMILES, InChi, or molecular fingerprints, to facilitate structure matching and comparison.
  • the search component 158 may employ reaction type matching, where it identifies similar experiments based on the types of chemical reactions involved. This approach can be beneficial when optimizing specific reaction classes or exploring alternative reagents or conditions for a particular transformation.
  • the search component 158 may leverage reaction classification algorithms or databases that categorize reactions based on their mechanistic patterns, functional group transformations, or other relevant criteria.
  • the search component 158 can incorporate machine learning techniques to identify similar experiments.
  • the search component 158 may train machine learning models on historical experimental data, including reaction conditions, reagents, yields, and other relevant features. These models can learn patterns and similarities between experiments, enabling the search component 158 to suggest relevant alternatives based on the provided experimental IDs 142.
  • the machine learning models employed by the search component 158 can range from simple similarity measures, such as k-nearest neighbors or cosine similarity, to more complex models like neural networks or decision trees. The choice of model can depend on factors such as the complexity of the chemical space, the availability of training data, and the desired trade-off between accuracy and computational efficiency.
  • the search component 158 may combine multiple search strategies, such as chemical structure matching and reaction type matching, to provide more comprehensive and relevant search results. This hybrid approach can leverage the strengths of different search techniques and provide users with a diverse set of alternative experimental IDs 142 to explore.
  • the search component 158 may incorporate user feedback and preferences to refine and enhance its search capabilities over time. For example, users may be able to provide ratings or rankings for the search results, which can be used to update the underlying search algorithms or models. This adaptive approach can ensure that the search component 158 remains relevant and aligned with the specific needs and objectives of the users.
  • the search component 158 may also integrate with external databases or resources to expand its search capabilities. For instance, it may access public or proprietary reaction databases, literature repositories, or patent databases to identify additional relevant experiments or synthetic routes. This integration can provide users with a broader perspective and facilitate the exploration of alternative experimental IDs 142 beyond the scope of the local database 132.
  • the search component 158 may provide additional information or analysis to support the decision-making process. This information can include statistical summaries of the identified similar experiments, such as distributions of yields, process metrics, or reaction conditions. The search component 158 may also highlight key differences or similarities between the alternative experimental IDs 142 and the initially selected ones, enabling users to make informed choices based on their specific requirements or constraints.
  • the searching may be activated by the button 502 on the webpage as is shown in Fig. 7. Also shown in the GUI on Fig. 7, is that the molecular role 504 is selectable as well.
  • the GUI of Fig. 7 When the GUI of Fig. 7 is presented to a user, it shows the synthesis steps with to thereby give the user the ability to modify the reactants. Although Fig. 7 only shows 2 steps, additional steps of the process may be shown by allowing the user to scroll down.
  • Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure where there is a drop-down menu to change experiments. [00106] Referring again to Fig. 1, each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs 142.
  • the GUI component 160 is configured for selecting an alternative experimental ID 142 for at least one of the one or more experimental IDs 142.
  • Figs 8A and 8B show Sankey Diagrams representing the mass flow from the selected experiment, e.g., the experiment of Fig. 6.
  • Fig. 9 shows an example result of the search component 158 in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 7.
  • Fig. 9 shows alternative experimental IDs and in some embodiments, it can be matched up to the ones that it replaces via another column, for example.
  • the resulting PMI vs. Yield values for each alternative experimental ID may be shown as in Fig. 10. Referring yet again to Fig.
  • the process metrics estimator component 152 updates the aggregate process metric in accordance with the alternative experimental ID 142 when a user selects the alternative experimental ID 142 to replace one of the initially entered in experimental ID 142.
  • Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow starting materials’, e.g., reactants, catalysts, etc. is estimated for the experiment. Thus, Fig. 11 shows another view off the GUI.
  • the synthesis tree component 156 of the system 100 may be used to identify a set of possible synthesis routes leading to a target material.
  • the search component 158 may be used to identify a set of similar experiments to each experimental ID 142 of the experimental IDs 142 and determine an optimal set of reactions to synthesize the target material.
  • the process metrics estimator component 152 can estimate the process metrics for each of the optimal set of reactions, and the aggregate process metric for the integrated processes.
  • the GUI component 160 can be used to visualize the synthesis route and adjust the experimental IDs 142, as well as adjust the process parameters to optimize the process metrics.
  • the experimental data corresponding to each experimental ID 142 includes information about one or more previously performed experiments which are stored in the subprocess metrics 144.
  • the aggregate process metric can be a process mass intensity, a solvent intensity, a water intensity, a global warming potential, cost, yield, an environmental metric, etc.
  • the GUI component 160 of the system 100 may include various user interface elements, such as dropdown menus, checkboxes, sliders, and text input fields, to enable users to select experimental IDs 142, adjust process parameters, and visualize the synthesis route.
  • the GUI component 160 can also be designed to be accessible via a web interface, allowing users to access the system 100 from anywhere with an internet connection.
  • the GUI component 160 is configured for selecting a role for a molecule such that the molecule is associates with an experimental ID 142 of the one or more experimental IDs 142.
  • the GUI component 160 can also be configured for adjusting a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
  • some of the experimental IDs 142 may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based upon the input parameters. Modifying the input parameters can update these output values.
  • the act of selecting the one or more experimental IDs 142 is performed by a user utilizing the GUI component 160.
  • the GUI component 160 can also be used to query a prediction engine 164 to determine a prophetic experiment. It can be used to replace at least one of the one or more experimental IDs 142 with the prophetic experiment and estimate a prophetic process metric corresponding to the prophetic experiment.
  • the system 100 can also update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs 142.
  • the prophetic process metric can include a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region.
  • the plurality of process metrics can be a ratio, such as PMI over Yield.
  • the one or more experimental IDs 142 can be selected out of order, and the act of linking the plurality of experiment IDs can include the act of ordering the one or more experimental IDs 142 to form the integrated process including all of the corresponding subprocesses, in order to synthesize the target material.
  • the system 100 is designed to efficiently estimate mass flow and enable users to adjust and optimize various process metrics to improve the efficiency and sustainability of chemical or industrial processes.
  • the mass-flow estimator component 112 may assign a confidence score. If the mass-flow estimator component 112 bases some or all of the estimates on data, models generated from data, or Monte Carlo simulation data, a confidence score can be assigned to the estimates of mass flow values to indicate the quality of the estimate. This could be included directly in the output or be derived from the standard deviation or variance in the sample data used to make the estimation, the min-max of mass flow data points, etc.
  • a frequentist confidence store may be derived using frequentist statistics.
  • a confidence score may use sample data of a distribution, hypothesis testing, p-values, significance testing, confidence intervals etc.
  • a confidence score is calculated for each (or a set of) sample values using posterior probabilities in a Bayesian estimate, which represent the updated belief about the mass flow values.
  • the confidence score for example, may be a credible interval of a posterior distribution or of a Bayesian estimator.
  • the mass-flow estimator component 112 also includes the communications component 162.
  • the communications component 162 may facilitate seamless communication and data exchange between multiple software applications, devices, and systems. That is, the communications component 162 may include protocol handling, message formatting, data serializing, encryption, and authentication to facilitate the communication with the computers 104 and/or the mobile device 106.
  • the communications component 162 may utilize a message formatting mechanism to format the messages into formats, such as XML, JSON, binary formats, and/or proprietary message formats.
  • the communications component 162 may utilize various encryption algorithms, such as RSA, AES, ECC, symmetric encryption, asymmetric encryption etc. to enable secure communications between the mass-flow estimator component 112 and the computers 104 and/or the mobile device 106.
  • the GUI component 160 can render a display for use by the computer 104 and/or the mobile device 106.
  • the GUI component 160 may be a webpage-based provider, such as flask, an HTML server, a web framework, etc.
  • the GUI component 160 may provide widgets, information, buttons, options, and menus to thereby facilitate a user’s interaction with the mass-flow estimator component 112.
  • the GUI component 160 can be used to log into user accounts 146 so that a user can create, save, or retrieve the experimental IDs 142 and/or the subprocess metrics 144, or otherwise interface with any account features. Additionally or alternatively, the GUI component 160 can save favorites, select default parameters, or adjust default values.
  • the GUI component 160 can direct other components to execute instructions based upon a workflow initiated by a user. That is, the GUI component 160 may receive events, such as a mouse click, button press, or GUI widget interaction to initiate a routine, series of steps, or series of acts. For example, the GUI component 160 may guide a user step-by-step on how to set up and work with the subprocess metrics 144 within the database 132.
  • the GUI component 160 may also be used to visualize the results of the process models and the waste estimates, in aggregate, in simulation, and/or may provide various visualization tools to analyze the data.
  • the data may be stored in the database 132 or the Electronic Notebook Component 150.
  • each user can log into a user account 146 to visual the results of their processes, the results of modifications to their processes on the entire synthesis chain, obtain a direct comparison and/or historical accuracy of their process mass flow values.
  • a resource dispatcher 110 may dispatch requests to perform an action to one or more virtual servers 122, each of which has a virtual processor 124, a virtual memory 126, and a virtual disk space 128.
  • the virtual servers 122 can be executed on one or more servers 121 on a server farm 119 as dispatched and activated by the resource dispatcher 110.
  • Fig. 2 show a block diagram illustration of a computing device 200 to calculate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure.
  • the computing device 200 of Fig. 2 may be the computer 104 or mobile device 106 of Fig. 1.
  • the computing device 200 includes an VO interface 210 to communicate therewithin.
  • the computing device 200 includes a data store 204, a processor 206, a network interface 208, a memory 225, and user I/O devices 226.
  • the data store 204 stores data and may be a hard drive, flash drive, thumb drive, volatile memory, non-volatile memory, semi-volatile memory etc.
  • the processor 206 can execute one or more processor-executable instructions 212, which may be stored in the data store 204 and/or the memory 225.
  • the processor 206 can execute processor-executable instructions 212 stored in memory 225 that was retrieved from the data store 204.
  • the memory 225 also includes program data 214 that may include information related to the processor-executable instructions 212.
  • the computing device 200 may include user I/O devices 226, such as a cursor device 230 (e.g., touchscreen or mouse), a keyboard 232 (virtual or physical), and/or a monitor 228 (which may be a touchscreen).
  • the computing device 200 communicates with the network 202 via a network interface 208.
  • the mass flow calculation functionality may reside wholly within the computing device 200 of Fig. 2.
  • the mass-flow estimator component 112 of Fig. 1 may reside within the processor-executable instructions 212 of Fig. 2 as mass flow calculation and mass-flow estimator component 242.
  • the electronic notebook component 150 may be the same or similar to the electronic notebook component 150, the process metrics estimator component 152, the linking component 154, the synthesis tree component 156, the search component 158, the GUI Component 160, the communications component 162, and the prediction engine 164 of Fig. 1, respectively.
  • the database 244 may be similar to the database 132 of Fig. 1.
  • the database 244 may, for example, be an Oracle, SQLite, postgresql, MariaDB, MySQL, or any other database embedded on the computing device 200.
  • the experimental IDs 246, the subprocess metrics 248, and the user accounts 250 of Figs. 2 may be similar or identical to the experimental IDs 142, the subprocess metrics 144, and the user accounts 146 of Fig. 1, respectively.
  • the mass-flow estimator component 242 may reside wholly on a local device (such as on the computers 104, the mobile device 106, etc.) may be partially within a cloud service provider 102, and/or may be organized in a hybrid local and cloud configuration.
  • the mass-flow estimator component 242 may be an application, may be executed on the computers 104, the mobile device 106, the cloud service provider 102, the computing device 200, etc. or some combination thereof.
  • Fig. 3 is a flowchart of an example process 300.
  • one or more process blocks of Fig. 3 may be performed by a device, such as a computing device.
  • the process involves selecting one or more experimental IDs from an electronic notebook.
  • Each experimental ID corresponds to a subprocess that has at least one input material and at least one output material. Those input and output materials might be the same or different.
  • the experimental ID may include experimental data corresponding to one or more previously performed experiments.
  • the experimental IDs may be selected by using a GUI, such as that shown in Fig. 4. Or, Fig. 4 may show the listed of experimental IDs after a list is inputted into a GUI and may be after linking.
  • a role is selected for a molecule associated with an experimental ID from the One or more experimental IDs.
  • Fig. 7 shows a drop-down box where a role can be selected.
  • Act 306 the one or more experimental IDs are linked to form an integrated process that includes all of the corresponding subprocesses, which together synthesize a target material. This is done automatically or with user interaction.
  • Act 308 involves estimating a plurality of process metrics, where each of the metrics corresponds to a respective experimental ID from the one or more experimental IDs.
  • an aggregate process metric is estimated, which corresponds to the integrated process, and is a function of the estimated plurality of process metrics.
  • the aggregate process metrics may be, for example, a process mass intensity, a solvent intensity, a water intensity, a global warming potential, a cost, an energy uptake, a waste score, a yield, and any other environmental metric.
  • the process then moves to Act 312, which involves providing a synthesis tree (e.g., as shown in Fig. 5) that corresponds to the one or more experimental IDs.
  • Act 314 searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs.
  • similar experiments are experiments with a primary input molecule and a target molecule that is the same as one or more experimental IDs. Helper molecules, such as buffers, may be different.
  • Act 316 selects an alternative experimental ID for at least one of the one or more experimental IDs. Then, in Act 318, the aggregate process metric is updated in accordance with the alternative experimental ID. In Act 320, a process parameter is updated. Act 322 involves updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. Act 324 queries a prediction engine to determine a prophetic experiment. Act 326 then replaces at least one of the one or more experimental IDs with the prophetic experiment.
  • a prophetic process metric is estimated, which corresponds to the prophetic experiment, and the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
  • the prophetic process metrics can have a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region, and the multiple process metrics can be expressed as a ratio, such as PMI over Yield.
  • the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
  • Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations so that different so that different experiments may be compared.
  • the process allows for the consideration and comparison of experimental data from one or more previously conducted experiments within each experimental ID.
  • the aggregate process metric may be based on process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield or environmental impact, either alone or in combination with other customizable variations.
  • a molecule associated with an experimental ID can be selected to play a specific role (e.g. product, starting material, reagent, solvent, catalyst), while process parameters can be adjusted to achieve optimal results.
  • process metrics can be updated to achieve target parameters.
  • the user can select experimental IDs using a graphical user interface. Additionally, the one or more experimental IDs can be selected out of order, but must be properly ordered to form an integrated process that includes all corresponding subprocesses that synthesize the target material.
  • the customizable variations can be implemented separately or in combination with one another, and the process may include more, fewer, or differently arranged blocks than those illustrated in Figure 3.
  • the blocks in the process can be performed simultaneously when necessary.
  • the process can provide a versatile and customizable process for optimizing chemical reactions through the consideration of various experimental and process parameters.
  • process 300 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 3. Additionally, or alternatively, two or more of the blocks of process 300 may be performed in parallel.
  • Figs. 14, 15 and 16 are described as follows to illustrate the operation of the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2
  • the predictive estimation of sustainability metrics (solvent intensity, water intensity, product carbon footprint) with an electronic notebook-data using the system 100 of Fig. 1 can utilize a prediction engine, e.g., retro-synthesis tools
  • the production of highly complex molecules such as APIs (active pharmaceutical ingredients) is a multi-step synthesis process that starts from already complex raw materials.
  • building blocks are for example the model compounds shown in Fig. 14.
  • the prediction engine 164 is a component of the mass-flow estimator component 112 in the system 100 for calculating mass flow and estimating sustainability metrics in chemical synthesis. This prediction engine 164 can serve various purposes and functionalities within the overall system.
  • the prediction engine 164 may be queried to determine a prophetic experiment.
  • the term "prophetic experiment” can refer to a hypothetical or simulated experiment that has not yet been physically performed.
  • the prediction engine 164 can utilize various techniques, such as machine learning models, computational chemistry methods, or rule-based algorithms, to generate predictions or suggestions for potential chemical reactions or processes that could be explored. These predicted or prophetic experiments can then be incorporated into the overall integrated process being analyzed by the system.
  • the prediction engine 164 may be used to replace at least one of the one or more experimental IDs 142 with a prophetic experiment suggested by the prediction engine 164.
  • the process metrics estimator component 152 can then estimate a prophetic process metric corresponding to this prophetic experiment.
  • the aggregate process metric which is a function of the estimated plurality of process metrics, can be updated to include the prophetic experiment in place of the replaced experimental ID(s) 142.
  • the prophetic process metric estimated by the process metrics estimator component 152 for the prophetic experiment may include a range of values.
  • This range of values can take various forms, such as a confidence interval, a confidence region, a credible interval, or a credible region. These ranges can provide an indication of the uncertainty or variability associated with the predicted or simulated experiment, allowing users to assess the reliability or robustness of the predictions made by the prediction engine 164.
  • the prediction engine 164 may employ retrosynthesis tools or algorithms to suggest potential synthesis routes or pathways for a target molecule. These retrosynthesis tools can analyze the target molecule and propose a series of chemical reactions or transformations that could be used to synthesize the target from simpler starting materials or building blocks. The prediction engine 164 can then use this suggested synthesis route, along with data from the electronic notebook component 150 and the subprocess metrics 144, to estimate process metrics and sustainability indicators for the proposed synthetic pathway.
  • the prediction engine 164 may be integrated with other components of the system, such as the search component 158, to identify similar reactions or processes from the database 132 that could inform or refine the predictions made by the prediction engine 164.
  • the search component 158 may identify similar experimental IDs 142 or subprocesses that have been previously performed, and the prediction engine 164 can use this information to improve the accuracy or reliability of its predictions.
  • the prediction engine 164 can also be used in conjunction with the synthesis tree component 156 and the GUI component 160.
  • the synthesis tree component 156 provides a visual representation of the integrated process, including the one or more experimental IDs 142 and their corresponding subprocesses.
  • the prediction engine 164 may suggest modifications or alternatives to this synthesis tree, which can be visualized and edited through the GUI component 160, allowing users to explore different scenarios and optimize the integrated process based on the predictions made by the prediction engine 164.
  • the prediction engine 164 may incorporate various types of data and models to make its predictions. This can include experimental data from the electronic notebook component 150, thermodynamic calculations, empirical correlations, or machine learning models trained on relevant chemical data.
  • the prediction engine 164 may also integrate with external databases or resources to obtain additional information or data relevant to the chemical processes being analyzed.
  • Model compound A is a halogenated aromatic with an additional nitrile group
  • compound B is a halogenated heteroaromatic molecule.
  • Such are so-called value- added complex intermediates in the chemical industry, which means that they are produced from base chemicals via one or more synthesis steps.
  • Such molecules usually carry more than one functional group (e.g. halogenic substituent, nitrile group, amine group, alcohol group, ester from boronic acid, etc.).
  • Starting materials should be base chemicals, available in bulk quantities, that are available in a life cycle inventory database (e.g. ecoinvent)
  • SynthiaTM is a retro-synthesis tool that suggests production routes for a certain molecule and the above-mentioned information.
  • An alternative to SynthiaTM is the tool ASKCOS from MIT (Massachusetts Institute of Technology), which is available free of charge.
  • SynthiaTM suggests a two-step synthesis from the base chemical maleic acid and hydrazine (Fig. 15).
  • Fig. 14 shows model compounds for raw materials to synthesize active pharmacal ingredients while Fig. 15 shows how it may be be presented in the GUI.
  • Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure.
  • Both chemicals are available in life cycle inventory database such as ecoinvent. Additionally, the type of reaction is shown and basic information on reaction conditions (e.g. solvent, catalyst, temperature during reaction) are mentioned.
  • the disclosed method does offer the automated generation of combined PMI / yield data of recorded chemical conversions and related processes, thereby preparing data sets, which can be used to allow a faster and more qualified plausibility testing of 'greener' retro-synthetic planning by systematic analysis of similar structures and chemical conversions.
  • This information can be used to search for similar reactions by the search component 158 to obtain the average amount of solvent, water, catalyst, etc. consumed in each step by looking at the average PMI for the specific reaction type that was suggested by the retro-synthesis tool.
  • the prediction engine 164 can also provide an average chemical yield for the respective synthesis steps as shown in Fig. 16. Based on the comparison of PMI and chemical yield for dozens of similar reactions documented in the ELN-system, an average amount of solvent, water, catalyst, etc. can be applied.
  • an uncertainty estimate may be made in the environmental footprint evaluation.
  • the deviation of the PMI data points to the average PMI as well as the deviation of data points for the chemical yield to the average yield can provide an indication on the accuracy on the environmental footprint evaluation.
  • Product carbon footprints can be estimated based on life cycle inventory database entries for raw materials, solvents, water, catalysts, etc.
  • Figs. 11 and 12 describe the function of the recycling functionality incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2. In this module the effect of solvent recycling on sustainability metrics is estimated.
  • Fig. 17 describes the function of the energy module incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2.
  • Fig. 17 may be a separate module used for energy calculation at each step using, e.g., input metrics.
  • energy consumption of the selected processes or subprocesses can be estimated based on basic physical phenomena and/or empirical relationships. Individual unit-operations are selected within the tool and important parameters are adjusted by the user. Examples for unit operations are, but are not limited to, heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction.
  • a summary of the energy contribution the unit-operations of the unit operations in the synthesis tree are visualized in the software, like shown in Fig. 18. The software allows adaption to various scales. The energy contribution might be included into the estimation of the final product carbon footprint.
  • Fig. 19 describes the possibility to include weight-based carbon footprints for each individual staring material. These can be assigned manually, can be retrieved from external date vendors (e.g. Ecoinvent), or can be assigned based on the individual role of each starting material. The output presented in the GUI can also be a weight-based carbon footprint estimation.
  • weight-based carbon footprints for each individual staring material. These can be assigned manually, can be retrieved from external date vendors (e.g. Ecoinvent), or can be assigned based on the individual role of each starting material.
  • the output presented in the GUI can also be a weight-based carbon footprint estimation.
  • a computer-implemented method for calculating mass flow and estimating sustainability metrics comprising: selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material; linking the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses wherein the integrated process synthesizes a target material; estimating a plurality of process metrics wherein each of the plurality of process metrics corresponds to a respective experimental IDs of the one or more experimental IDs; estimating an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics; providing a synthesis tree corresponding to the one or more experimental IDs; searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs; selecting an alternative experimental ID for at least one of the one or more experimental IDs; and updating the aggregate process
  • each experimental ID includes experimental data corresponding to one or more previously performed experiments.
  • the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
  • the subprocesses include at least one of chemical reactions or purifications.
  • each experimental ID is associated with specific equipment used in the subprocess, and the method further comprises adjusting equipment settings based on the subprocess requirements.
  • estimating the energy consumption metric comprises applying at least one of empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs.
  • estimating the energy consumption metric comprises applying at least one of empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs.
  • a data processing system comprising means for carrying out the method of any one of aspects 1 to 56.
  • a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56.
  • a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56.

Landscapes

  • Business, Economics & Management (AREA)
  • Engineering & Computer Science (AREA)
  • Human Resources & Organizations (AREA)
  • Strategic Management (AREA)
  • Economics (AREA)
  • Theoretical Computer Science (AREA)
  • Entrepreneurship & Innovation (AREA)
  • General Physics & Mathematics (AREA)
  • Marketing (AREA)
  • Tourism & Hospitality (AREA)
  • Physics & Mathematics (AREA)
  • General Business, Economics & Management (AREA)
  • Operations Research (AREA)
  • Quality & Reliability (AREA)
  • Chemical & Material Sciences (AREA)
  • Game Theory and Decision Science (AREA)
  • Development Economics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Primary Health Care (AREA)
  • Analytical Chemistry (AREA)
  • Manufacturing & Machinery (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Sustainable Development (AREA)
  • Data Mining & Analysis (AREA)
  • Educational Administration (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Computing Systems (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Stored Programmes (AREA)

Abstract

A method is disclosed that selects one or more experimental IDs from an electronic notebook. The method links the one or more experimental IDs to form an integrated process to synthesize a target material and estimate a plurality of process metrics wherein each of the plurality of process metrics corresponds to a respective experimental ID of the one or more experimental IDs. The method estimates an aggregate process metric corresponding to the integrated process. The aggregate process metric is a function of the estimated plurality of process metrics. The method provides a synthesis tree corresponding to the one or more experimental IDs and searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments. The method also selects an alternative experimental ID for at least one of the one or more experimental IDs and updates the aggregate process metric in accordance with the alternative experimental ID.

Description

METHOD AND SYSTEM FOR CALCULATING MASS FLOW AND ESTIMATING SUSTAINABILITY METRICS IN CHEMICAL SYNTHESIS
BACKGROUND
Relevant Field
[0001] The present disclosure relates to chemical synthesis. More particularly, the present disclosure relates to a method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis.
[0002] Description of Related Art
[0003] The production of chemicals through raw material processing can have negative impacts on the environment and human health, making sustainability a primary consideration of chemical synthesis. To account for these negative impacts, sustainable and environmentally friendly processes for chemical production can be developed that reduce pollution and protect population health. Thus, it is important to develop processes that manage emissions and waste to account for their environmental impact.
[0004] Making sustainable decisions during the early stages of research and development can prove to be more efficient than waiting until after production starts. This is because modifying established production processes after the fact can be complex due to interdependent processes. Planning and implementing modifications during the preproduction phase is comparatively simpler. Thus, making early development decisions helps to ensure that the production of chemicals and the processing of raw materials have minimal negative effects on the environment and health while considering processing efficiency. A data-driven approach that monitors and provides possible alternatives can support chemical process development while facilitating sustainable practices.
[0005] One indicator in chemical process development is the so-called Process Mass Intensity (“PMI”) (for general key figures, see also: Green Chem., 2015, 17, 3111; Processes 2022, 10, 1274). The PMI can be used as an indicator of both the cost-effectiveness and environmental compatibility of a process. The PMI provides information on how efficiently the process mass is used compared to the mass of the target material or product. The process mass refers to all chemicals, organic solvents, water, auxiliaries such as catalysts or pH buffers, rinsing media, etc. used in the production process. This mass-based resource consumption is measured in relation to the mass of a produced chemical compound (usually kg per kg). A low (small) PMI value indicates that the process is comparatively efficient and that the process mass is used optimally, while a high (large) PMI indicates that the process is inefficient and the process mass is not being used optimally, which possibly leads to a larger amount of waste and thus a high environmental impact. Yet another use of PMI for the evaluation of chemical processes was pioneered by Jimenez-Gonzales et al. (Org. Process Res. Dev. 2011, 15, 4, 912- 917), who showed a direct correlation between the PMI and the potential amount of eCO2 as a measure of global warming potential (GWP).
[0006] By analyzing the PMI, processes can be improved and optimized to reduce costs, increase yield, and effectively account for the environmental impact. Especially when optimizing a chemical synthesis process with respect to sustainability, it can be helpful to identify the production step that has the most significant impact on driving up the PMI.
[0007] One challenge in chemical process development is the acquisition of process- oriented and up-to-date PMI values; manual input and individual evaluation during the PMI calculation methods can make the determination of PMI slow, error-prone, and not reflecting the latest state of the process. Additionally, the heavy manual interaction needed to calculate PMI can be a significant administrative burden for which the process developers are very often not specialized.
[0008] In addition to the PMI, other process metrics can be used, such as Solvent Intensity (“SI”) and Water Intensity (“WI”). The SI and WI are calculated similarly as the PMI, but only consider solvents or aqueous components, respectively. Moreover, energy consumption is another process metric that is typically considered for environmental sustainability.
[0009] Generally, energy consumption plays a subordinate role to material efficiency in standard processes in fine and specialty chemicals. Nevertheless, to quantify it, the energy consumption of the individual process steps (unit operations) can be estimated using thermodynamic calculations and empirically gained correlations. (Parvatker et al., ACS Sustainable Chem. Eng. 2019, 7, 6580-6591, "GSK Studie: Org. Process Res. Dev. 2011, 15, 912-917"). Such calculations can involve significant manual effort.
[0010] In the field of sustainability indicator calculation, there are multiple methods and utilities available for use, including the three that follow.
[0011] DOZN ™ Tool: An online application that allows for a semi-quantitative evaluation of a single process step based on the 12 principles of green chemistry. It is not suitable for complex syntheses, like linear or branched multi-step syntheses, and requires detailed manual inputs. Mass-based resource consumption and Global Warming potential cannot be calculated. [0012] American Chemical Society (ACS) Green Chemistry Institute, Process Mass Intensity Calculator: An Excel -based tool used for single or multi-step reaction PMI calculation by manually entering reactants, reagents, solvents and aqueous systems, as well as the interdependencies of the above steps.
[0013] Chem Pager / Roche: A module connected to an ERP system that allows PMI calculation for established production processes. It is unsuitable for process development as the data source is limited to ERP system recipes.
[0014] All aforementioned methods share two significant drawbacks: they require manual data entry, and they operate on a case-by-case basis, i.e. they do not include any data of similar reactions of the existing portfolio. This makes the process of determining the PMI slow, error-prone, and potentially relying on outdated data, as described previously. The Roche model utilizing ERP data is an exception, but it is limited to analyzing already established production processes stored in the ERP system.
[0015] Some commercial providers offer a more detailed analysis known as Life Cycle Assessments (LCAs), which in a first step require the gathering of detailed customer data and in a second step link these data to other databases for analysis. However, such time consuming and costly analysis is typically only useful for products already in routine production and not for products in development phase. Examples of such providers include Sphera (GaBi) and Pre (SimaPro). Some software solutions for automated calculation of Product Carbon Footprints (PCFs) do exist, which pull information from ERP systems without requiring manual data collection. Examples of such providers include AllocNow, Acts and iPoint. Carbon Minds offers a generalized approach for the estimation of carbon footprints of base chemicals. It should be noted that mass-based resource consumption and process energies are often not recorded at product level in ERP systems and must be allocated accordingly. Therefore, also these tools are inadequate for evaluating the environmental impact of a product during the development phase. In summary, all currently existing solutions require significant manual intervention, are slow and/or are time consuming, and in some cases very expensive so that they are not suitable for daily use in product development.
[0016] SUMMARY
[0017] The present disclosure relates to a computer-implemented method for calculating mass flow and estimating sustainability metrics in chemical processes. The method may be initiated by selecting one or more experimental IDs from an electronic laboratory notebook (ELN), where each experimental ID corresponds to a subprocess with input and output materials. The selection process may be performed through a Graphical User Interface. Experimental data corresponding to one or more previously performed experiments may be associated with the selected experimental IDs.
[0018] Further steps of the method may include linking the selected experimental IDs to form an integrated process, which synthesizes a target material. The integrated process can include multiple corresponding subprocesses, and a synthesis tree corresponding to the one or more experimental IDs is provided by the software to support the process. The synthesis is adjustable by the user in some embodiments. Once the integration is established, the method estimates a plurality of process metrics, such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or environmental metric, wherein each of the plurality of process metrics corresponds to the respective experimental IDs of the one or more experimental IDs. The process metrics may also be a ratio, such as a PMI over Yield ratio. The process metrics may be provided for the whole process, parts of the process, or each of the subprocess steps. For easier comparison, the Process metrics may also be compared to one or more process metrics from other processes. The process metrics can also be intuitively visualized by graphs, like a Sankey-Diagram.
[0019] An aggregate process metric may be calculated, which is a function of the estimated plurality of process metrics. The aggregate process metrics may be, for example, a cost, a yield, an environmental metric, or the aggregation of any process metrics. The aggregate process metric may be a summation of the plurality of process metrics. The metrics may be on a unit basis, a mass basis, a volume basis, a batch basis, or in any unit known to one of ordinary skill in the relevant art. The method may include searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each of which corresponds to one or more respective experiments of the one or more experimental IDs. The search may look for reactions with the same input and output molecules and, in some embodiment, may ignore waste or support molecules such as catalysts, for example. The search may also look for recycled mass streams, that means, waste streams that are recycled to the beginning of a subprocess or a process. Based on this search, at least one alternative experimental ID may be selected for replacing one or more of the selected experimental IDs.
[0020] A prophetic experiment can also be determined by querying a prediction engine and replacing at least one of the one or more experimental IDs with the prophetic experiment. The method may then estimate a prophetic process metric corresponding to the prophetic experiment and update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs. The prophetic process metric may include a range of values. The range of values may be a confidence interval, a confidence region, a credible interval, and/or a credible region. Those cited ranges of values may also include multi-step processes.
[0021] In another embodiment, the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. The act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material. Thus, the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material. In a preferred embodiment, this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs. The plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
[0022] Some embodiments include a data processing system comprising means for the implementation of any of the acts described above. The disclosure also includes a computer program and a computer-readable medium that can execute and perform the steps described in any one of the acts described above.
[0023] In summary, some embodiments provide an intelligent and efficient computer- implemented method for calculating mass flow and estimating sustainability metrics in chemical processes. The method employs experimental data to select and link experimental IDs to form an integrated process, estimates process metrics, and calculates an aggregate process metric. Some embodiments allow for the selection of alternative experimental IDs, the adjustment of a process parameter, and querying of a prediction engine to identify a prophetic experiment.
[0024] In some embodiments, the method involves selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material. The method then links the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses, wherein the integrated process synthesizes a target material. The method further estimates a plurality of process metrics, wherein each of the process metrics corresponds to a respective experimental ID of the one or more experimental IDs. Additionally, the method estimates an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics. The method also provides a synthesis tree corresponding to the one or more experimental IDs. Furthermore, the method searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, with each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs. The method then selects an alternative experimental ID for at least one of the one or more experimental IDs and updates the aggregate process metric in accordance with the alternative experimental ID.
[0025] In some embodiments, the method may involve displaying the process metrics that are estimated. The plurality of process metrics estimated, such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other environmental metrics, can be visually displayed through the graphical user interface. This display of the estimated process metrics can provide users with a convenient overview and analysis of the efficiency, sustainability, and potential environmental impact of the integrated process. The visualization of process metrics can aid in interpreting the estimation results, identifying areas for potential improvement, and guiding decisions during chemical process development. In some implementations, advanced data visualization tools may be leveraged to provide intuitive charts, graphs, Sankey diagrams or other graphical representations of the estimated process metrics.
[0026] In some embodiments, each experimental ID corresponds to a subprocess and includes experimental data from one or more previously performed experiments or historic experiments. By linking experimental IDs to form an integrated process that synthesizes a target material, the method is able to incorporate experimental data from previously conducted experiments into the analysis of the overall integrated process. This previously conducted experimental data provides the necessary information to estimate process metrics for the subprocesses associated with each experimental ID in the integrated process, as well as an aggregate process metric for the overall integrated process. The experimental data may encompass details regarding factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other parameters relevant to calculating process metrics like process mass intensity, yield, cost, and environmental impact indicators. Overall, the consideration of previously conducted experimental data lends greater accuracy and reliability to the method's ability to analyze the efficiency and sustainability of chemical synthesis processes.
[0027] In some embodiments, the method further comprises comparing the aggregate process metric with the alternative experimental ID to the aggregate process metric before updating the aggregate process metric based on the alternative experimental ID. More specifically, the aggregate process metric calculated using the initially selected experimental ID(s) is compared to what the aggregate process metric would be if a particular alternative experimental ID was used instead. This comparison step allows for assessment of the impact of selecting a different experimental ID on the overall aggregate process metric prior to updating the metric. By enabling this comparison, users can make informed decisions about whether to replace an initial experimental ID with an alternative option, based on how that replacement would affect key aggregate metrics like process mass intensity, cost, yield, etc. Only after comparing the aggregate metrics is the aggregate process formally updated to incorporate the selected alternative experimental ID in place of the initial experimental ID. This embodies an optimization approach that empowers users to explore multiple experimental options and select the one that optimizes the aggregate process according to user-defined criteria.
[0028] In some embodiments, the aggregate process metric with the alternative experimental ID may be compared to the aggregate process metric prior to updating the aggregate process metric. This allows for the comparison of the process metrics before and after the replacement of an experimental ID. The results of this comparison may additionally be displayed through the graphical user interface or other visualization methods. By showing the comparison, users can evaluate the impact of selecting a particular alternative experimental ID on the overall aggregate process metric.
[0029] In some embodiments, the aggregate process metric estimated by the process metrics estimator component may be a process mass intensity (PMI). The PMI provides a quantitative indicator of the efficiency of a chemical process by measuring the total mass of materials used in the process per unit mass of product generated. A lower PMI value generally indicates a more efficient process with less waste. Thus, the aggregate process metric calculated as a function of the individual process metrics for each experimental ID may be this PMI value, representing the overall mass efficiency of the integrated chemical process under analysis. Using the PMI as the aggregate process metric enables effective assessment and comparison of the sustainability of alternative integrated processes in terms of their mass utilization. The system provides the flexibility to use PMI as the key aggregate metric for optimizing mass flow through the integrated synthesis process.
[0030] The aggregate process metric estimated by the method may represent one of several sustainability indicators, including solvent intensity, water intensity, or global warming potential. More specifically, in some embodiments, the aggregate process metric calculated as a function of the plurality of estimated process metrics for the individual experimental IDs may be the overall solvent intensity, water intensity, or global warming potential for the integrated process.
[0031] In some embodiments, the aggregate process metric estimated for the integrated process can include cost metrics, yield estimates, and various environmental metrics. Specifically, the aggregate process metric calculated by the process metrics estimator component may represent overall cost parameters such as total raw material costs, energy costs, and overall process economics. It may also incorporate chemical yield projections and optimizations based on the experimental data. Additionally, the aggregate process metric can encompass composite environmental impact measures that account for factors like greenhouse gas emissions, wastewater generation, solid waste production, and other sustainability indicators. By consolidating different parameters into one overarching metric, the system aims to provide users with a high-level quantification of the performance, economics, and environmental profile associated with the integrated process. The flexibility to compute aggregate metrics spanning cost, yield, and sustainability domains allows for multi-objective analysis and aids in the identification of optimal process configurations.
[0032] In some embodiments, the method may involve selecting a role for a molecule that is associated with an experimental ID of the one or more experimental IDs. For example, a user may utilize the GUI to select a specific role, such as product, starting material, reagent, solvent, catalyst, etc., for a given molecule linked to an experimental ID. Associating molecules with particular roles can help provide additional context and clarity regarding how various compounds are being utilized within the integrated process. Defining these molecule roles enables the system to make appropriate assumptions and calculations when estimating the process metrics. The ability to select molecule roles provides an added level of customization, allowing users to tailor the analysis to their specific needs.
[0033] In some embodiments of the method, a process parameter may be adjusted. This adjustment provides users with the capability to modify specific inputs or conditions associated with the integrated process in order to analyze the impacts on the overall process metrics. For example, users can tweak reaction temperatures, reagent quantities, solvent volumes, equipment settings, or other relevant process parameters. After adjusting the process parameter, at least one of the plurality of estimated process metrics can then be updated accordingly. This update allows users to immediately see how modifications to the process parameters influence metrics like process mass intensity, yield, cost, global warming potential, etc. By providing the functionality to adjust process parameters and view updated metrics, the method enables iterative optimization of the integrated chemical synthesis process. Users can repeatedly modify parameters and analyze the effects on mass flow, cost, sustainability indicators, and other output values of interest. Overall, the adjustment of process parameters combined with live updating of estimated metrics facilitates customizable scenario analysis and supports data- driven decision making during chemical process development.
[0034] In some embodiments, the method may involve adjusting a process parameter. Some of the experimental IDs may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based on the input parameters. Modifying the input parameters can update these output values. Thus, the method may further comprise updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. By enabling adjustment of parameters and updating process metrics accordingly, the method allows for customization and optimization of the integrated process to achieve target results.
[0035] In some embodiments, the selection of the one or more experimental IDs from the electronic notebook is performed by a user utilizing a Graphical User Interface (GUI). Specifically, the GUI provides an interface which enables the user to browse, search, or otherwise access the experimental IDs stored within the electronic notebook. The user can then manually select the desired experimental IDs through interactions with GUI elements such as checkboxes, dropdown menus, or search filters. This manual selection initiates the process of linking together the selected experimental IDs into an integrated process for the synthesis of a target material. The GUI therefore facilitates convenient user control over the initial experimental ID selection, while the subsequent linking and analysis steps can be automated by the system based on this user input. Overall, this GUI-enabled selection process allows for customizable construction of integrated processes according to the specific needs and priorities of each user.
[0036] In some embodiments, the method involves querying a prediction engine to determine a prophetic experiment. A prophetic experiment refers to a hypothetical or simulated experiment that has not yet been physically performed. The prediction engine can utilize various techniques to generate predictions or suggestions for potential chemical reactions or processes to be explored. The method then replaces at least one of the initially selected one or more experimental IDs with the prophetic experiment identified by the prediction engine. After replacing the experimental ID(s), the method estimates a prophetic process metric corresponding to the prophetic experiment. This prophetic process metric is then incorporated into the aggregate process metric calculation by updating the aggregate process metric to include the prophetic experiment in place of the replaced experimental ID(s). In essence, the prediction engine allows hypothetical experiments to be simulated within the overall integrated process, enabling users to explore different process variations and scenarios. The aggregate process metric reflecting the entire integrated process can then be re-estimated based on the incorporation of these prophetic experiments.
[0037] Optionally, the method may include a prophetic process metric corresponding to the prophetic experiment determined by a prediction engine. In some embodiments, the prophetic process metric includes a range of values, such as a confidence interval, a confidence region, a credible interval, or a credible region. This range provides an indication of the uncertainty or variability associated with the predicted experiment. Accounting for inherent uncertainty when incorporating prophetic experiments can help assess the reliability and robustness of the predictions.
[0038] In some embodiments, the prophetic process metric estimated for the prophetic experiment suggested by the prediction engine may comprise a range of values indicating uncertainty. This range could take various standard statistical forms for representing uncertainty, including a confidence interval, a confidence region, a credible interval, or a credible region. Using these types of value ranges allows users to assess the reliability and variability inherent in the predicted experiments from the prediction engine when evaluating the potential impact on the overall process metrics. Rather than providing only a single estimated number for the prophetic process metric, the prediction engine can quote a confidence interval, credible region, etc. to quantify the uncertainty in its estimations. This provides greater transparency into the accuracy of the predictive models used by the engine.
[0039] In some embodiments, the plurality of process metrics estimated may be represented as a ratio. For example, the ratio may be process mass intensity (PMI) over yield. Expressing the process metrics as a ratio can provide additional insights into the efficiency and sustainability of the integrated process by examining the relative values between different metrics. A lower ratio value may indicate improved performance, while a higher ratio value may highlight areas needing further optimization. Using customizable ratio metrics allows users to focus the analysis on parameters of greatest relevance to their specific needs and objectives. The ability to calculate and compare ratio-based process metrics adds an extra dimension of flexibility and customization to the overall method.
[0040] The plurality of process metrics may also be represented as a ratio. In some embodiments, this ratio is process mass intensity (PMI) over yield. Using the ratio of PMI over yield as a process metric enables evaluating both the mass efficiency and productivity of the integrated process in a simple, intuitive metric. A lower PMVyield ratio indicates greater efficiency and yield for the overall chemical synthesis pathway. Tracking how modifications to the process impact both PMI and yield through their ratio can further optimization efforts by balancing mass utilization and target output. Thus, the use of PMI over yield as a ratio metric provides a sustainability indicator for the integrated process.
[0041] In some embodiments, the one or more experimental IDs may be selected out of order by the user. However, the act of linking the plurality of experimental IDs includes the act of ordering the selected experimental IDs by the system to properly form the complete integrated process. This integrated process includes all of the corresponding subprocesses associated with the selected experimental IDs, arranged in the correct sequence. Automatically ordering the experimental IDs ensures that the full sequence of subprocesses is configured in the proper order to synthesize the target material. By rearranging the selected experimental IDs into the right process sequence, the system can link subprocesses that may have been originally selected out of order and still generate an integrated process that connects all associated subprocesses to produce the desired target output.
[0042] The method may involve subprocesses including chemical reactions, purifications, or a combination thereof. More specifically, in some embodiments each experimental ID selected from the electronic notebook corresponds to either a chemical reaction subprocess, a purification subprocess, or a subprocess involving both a chemical reaction and a purification. Therefore, when linking the experimental IDs to form the integrated process for synthesizing the target material, the resulting integrated process may feature multiple subprocesses including chemical reactions, purifications, or a combination of both chemical reactions and purifications. The ability to link together experimental IDs representing diverse types of chemical subprocesses provides flexibility in constructing integrated processes to synthesize desired target materials.
[0043] The method may involve generating a report that summarizes the estimated process metrics and the aggregate process metric. More specifically, in some embodiments, the method includes an additional step of producing a report that provides an overview of the plurality of process metrics calculated for each experimental ID as well as the aggregate process metric estimated for the integrated process. This report condenses the key information and metrics into a concise summary that allows users to easily review the overall efficiency, sustainability, cost, yield, and environmental impact indicators for the chemical process. By gathering the estimated metrics into a single report, users can assess the process performance holistically rather than examining individual metrics in isolation. The report provides a tool for convenient evaluation and comparison of different process configurations or alternatives explored using the system. In some implementations, the report may present the metrics in various graphical visualizations, such as bar charts, line plots, or Sankey diagrams, to enable intuitive interpretation of the data. The summary report equips users with a quick yet comprehensive perspective of the integrated process and its characterized performance based on the computed process metrics.
[0044] The method may involve verifying the compatibility of input and output materials when linking the one or more experimental IDs. This verification step acts to ensure continuity in the resulting integrated process that synthesizes the target material. Specifically, the input materials required for each experimental ID or subprocess are checked against the output materials produced by any preceding experimental IDs or subprocesses in the integrated process flow. Any mismatches in output versus required input materials can indicate possible breaks in continuity of the integrated process, which can then be addressed. By including this verification, an embodiment aims to confirm that the linkage of experimental IDs successfully creates an end-to-end integrated process with no gaps that would prevent the target synthesis. [0045] In some embodiments, the plurality of process metrics estimated by the process metrics estimator component may include energy consumption metrics. These energy consumption metrics can be calculated based on the input and output materials and the subprocesses utilized in each experimental ID of the integrated process. Specifically, factors such as the quantities and types of input materials, chemical transformations involved in the subprocesses, output materials produced, reaction conditions like temperature and pressure, and other relevant parameters can be used to estimate the energy consumption associated with each subprocess and experimental ID. Advanced techniques like computational fluid dynamics simulations, thermodynamic calculations, or empirical correlations based on historical data may be leveraged to quantify the energy consumption. By aggregating the energy consumption across all experimental IDs, an overall energy usage estimate can be obtained for the integrated process. This facilitates analysis of the environmental impact and cost-effectiveness of the overall synthesis route. The energy consumption metrics, along with the other estimated process metrics, allow users to make informed decisions when optimizing the integrated process.
[0046] The method may include a step of validating the estimated process metrics against predetermined criteria before updating the aggregate process metric. In some embodiments, there is a validation step where the estimated process metrics, such as process mass intensity, solvent intensity, water intensity, and other metrics, are checked against defined criteria or thresholds. The validation ensures accuracy and reliability of the metrics prior to using them to update the overall aggregate process metric. If the estimated process metrics for the individual experimental IDs meet the predetermined validation criteria, then they can be applied to update the combined aggregate metric for the full integrated process. However, if one or more process metrics fail to satisfy the validation checks, the method may require troubleshooting, metric re-estimation, or other corrective measures before the aggregate metric is updated. This validation step may act as a quality check on the estimated metrics, enhancing the reliability and robustness of the overall mass flow calculations and sustainability analysis enabled by the method.
[0047] In some embodiments, machine learning models are utilized for the searching for similar reaction subprocesses. The method involves searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, as previously described. The search can make use of machine learning algorithms and models that are trained on historical experimental data to identify patterns and similarities between experiments. These models enable the search to suggest relevant alternative experiments based on the provided experimental IDs. The utilization of machine learning models can expand the capabilities of the search and provide more comprehensive results, optimizing the process metrics and sustainability indicators.
[0048] In some embodiments, the method may include the step of alerting a user when the aggregate process metric exceeds a predetermined environmental impact threshold. More specifically, the aggregate process metric estimated as a function of the plurality of process metrics can be compared against a defined threshold representing the maximum allowable environmental impact. If the aggregate process metric is calculated to exceed this predetermined threshold, the system may generate and display an alert to notify the user. This alert feature enables users to recognize when certain sustainability or environmental impact constraints are violated due to excessive resource consumption or emissions associated with the integrated chemical synthesis process under analysis. The ability to define environmental impact limits and automatically trigger notifications when these constraints are surpassed can aid in adhering to sustainability guidelines and minimizing ecological footprints.
[0049] In some embodiments, each experimental ID may be associated with specific equipment used in the corresponding subprocess. For example, a reaction subprocess may utilize a particular reactor or purification equipment. The method may further comprise adjusting the settings or parameters of this equipment based on the requirements of that subprocess. For instance, if a reaction subprocess operates at a certain temperature and pressure, the reactor settings can be automatically adjusted to match those conditions. This automatic adjustment of equipment settings streamlines the experimental workflow and ensures alignment between the subprocess details contained within the experimental ID and the actual equipment configuration used to carry out that subprocess. By linking the experimental IDs to the relevant equipment in this manner, the method enables improved standardization, repeatability, and optimization of the subprocesses.
[0050] In some embodiments, the method may further involve automatically ordering chemicals and materials needed for the subprocesses. This automatic ordering is based on the input materials listed in the experimental IDs that were selected from the electronic notebook. By leveraging the input material information captured within the experimental ID data, the system can determine the required reagents, solvents, catalysts, and other chemicals needed to carry out each subprocess. It can then automatically generate purchase orders or materials requests to obtain the necessary supplies, ensuring analysts and researchers have the required ingredients on hand before commencing the chemical reactions or purifications encompassed within that subprocess. This just-in-time materials ordering facilitated by the automated system helps minimize inventory and procurement overhead for organizations frequently synthesizing new target compounds or materials.
[0051] In some embodiments, the synthesis tree provided to the user may include multiple alternative synthesis pathways that could be used to synthesize the target material. The method may provide the capability for the user to select one of these alternative synthesis pathways based on user-defined criteria or preferences. For example, the criteria could be related to optimizing particular process metrics like cost, environmental impact, or yield. The user interface allows the user to input these criteria and priorities. The method then facilitates the selection of the optimal synthesis pathway that best meets the specified criteria out of the alternatives included in the synthesis tree visualization. Enabling users to explore alternative synthesis routes and choose based on customizable metrics provides flexibility and aids in process optimization.
[0052] In some embodiments, the method may further involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs. This energy consumption metric can provide an indication of the overall energy requirements associated with the synthesis of the target material via the linked subprocesses. To estimate this metric, empirical correlations, thermodynamic calculations, or other techniques may be applied using parameters from the experimental data corresponding to each experimental ID. Such parameters can include temperature, pressure, flow rates, batch sizes, and other relevant factors that influence energy usage. By assessing the estimated energy consumption across the integrated process, opportunities can be identified to reduce energy demand through process or equipment optimization. Thus, evaluating the energy consumption metric allows for a more comprehensive analysis of the sustainability and cost-effectiveness of the chemical synthesis process under consideration.
[0053] In some embodiments, the method may involve estimating an energy consumption metric for the integrated process based on the one or more experimental IDs. Estimating the energy consumption metric can comprise applying empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs. For example, thermodynamic calculations and empirically gained correlations may be used to estimate the energy uptake of individual process steps, such as heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction. This approach leverages established techniques to assess the energy requirements and associated environmental impacts of the subprocesses linked to form the integrated process. By applying these energy estimation methods to the specific process parameters captured within the experimental IDs, the system can derive reliable energy consumption metrics for the overall integrated process.
[0054] In some embodiments, the method may involve estimating a carbon footprint metric for the integrated process. The carbon footprint metric may be estimated based on the one or more experimental IDs that were selected from the electronic notebook and linked to form the integrated process. Specifically, the carbon footprint estimation may utilize life cycle inventory data for the various input materials involved in the subprocesses and integrated process. By leveraging detailed life cycle impact information for the raw materials and chemical ingredients, a more accurate determination of the carbon footprint resulting from the overall synthesis process can be achieved. Thus, through consideration of the experimental data as well as life cycle inventory emissions data, the system can provide users with a carbon footprint metric that accounts for the environmental impact of material and energy flows throughout the multi-step process.
[0055] In some embodiments, the method further comprises providing a visualization of mass flow through the integrated process based on the one or more experimental IDs that were selected from the electronic notebook. This visualization may utilize graphical tools and diagrams that allow users to see how mass flows through each subprocess and the overall integrated process used to synthesize the target material. For example, Sankey diagrams or other types of flow charts could be generated to intuitively display quantities and mass balances across different stages of the chemical synthesis process defined by linking the experimental IDs. Enabling intuitive visualization of mass flow helps users identify areas of inefficient material usage and opportunities for reducing waste or environmental impact. The visualization component allows users to visually explore the mass flow impacts of selecting alternative experimental IDs or adjusting process parameters within the integrated process.
[0056] In some embodiments, the method includes a step of visualizing the mass flow through the integrated process based on the one or more experimental IDs. This visualization of the mass flow may employ a Sankey diagram, which allows an intuitive graphical representation of flows and their quantitative values within a system. Specifically, the Sankey diagram can illustrate the sequential flow of materials, energy transfers, or waste through each subprocess and throughout the overall integrated process. The thickness of the arrows in the Sankey diagram represents the magnitude or amount of mass flow. This Sankey diagram visualization provides users with an easily interpretable overview of how mass flows through and is transformed within the integrated chemical synthesis process under analysis. The visualization supports identification of mass intensive steps, recycling opportunities, yield losses, and other relevant mass flow characteristics, aiding in process assessment, optimization, and improvement.
[0057] The system may allow for the synthesis tree, which corresponds to the one or more experimental IDs, to be edited by the user through a graphical user interface. More specifically, in some embodiments the system includes functionality enabling user modification of the visualization of the overall synthesis route. This can facilitate optimization, process adjustments, or exploring alternative synthesis pathways.
[0058] In some embodiments, the act of searching for similar reaction subprocesses includes identifying subprocesses that have the same input and output molecular structures as a given experimental ID from among the one or more experimental IDs. Specifically, when searching for similar subprocesses, those processes that match the molecular structure of both the input material and the output material for a particular experimental ID can be recognized. By finding reaction subprocesses with identical input and output chemicals, alternative synthetic routes or optimizations for specific reaction steps within the integrated process may be determined. This approach focuses the search on processes that are highly comparable in terms of the chemical transformations being carried out.
[0059] In some embodiments, the method involves estimating a range of values for at least one of the plurality of process metrics based on the plurality of similar experiments identified in the search. More specifically, after searching for and determining similar reaction subprocesses and experiments for each experimental ID, a range of values may be calculated for process metrics like process mass intensity, solvent intensity, yield, cost, or other metrics. This range represents the variability in the metric across the multiple similar experiments found that correspond to the experimental ID(s) selected by the user. Providing such a range gives an indication of the uncertainty and potential fluctuation in the process metric, allowing for a more comprehensive analysis when evaluating and optimizing the integrated chemical synthesis process. The range could take the form of a confidence interval, confidence region, credible interval, credible region, or other representation of variability. This enhanced uncertainty quantification through metric value ranges enables more informed decision making during chemical process development.
[0060] Some embodiments of the method may involve suggesting an alternative solvent or an alternative reagent for at least one of the experimental IDs based on the plurality of similar experiments found by the search component. The search component identifies experiments with similar reactions and can determine if there may be better solvents or reagents that could be used for an experimental ID by analyzing the reagents and solvents used in those similar experiments. By leveraging the data on solvents and reagents from similar experiments, the system can recommend testing alternate solvents or reagents for an experimental ID that may improve yield, lower cost, reduce waste, or provide other benefits over the original solvent or reagent selected. This allows researchers to easily get suggestions on alternate reaction conditions to try that could optimize their process.
[0061] The method allows for the selection of one or more experimental IDs from an electronic notebook, wherein the selection is based on a target molecule specified by a user. In some embodiments, a user may first specify or input a desired target molecule they wish to synthesize. Based on this target molecule, relevant experimental IDs can then be retrieved and selected from the electronic notebook that correspond to subprocesses involved in the production or synthesis of the specified target material. Thus, the target molecule provides a means for the user to indicate what final product they intend to make, while the system identifies and selects the necessary reaction steps and corresponding experimental IDs that can lead to the target molecule. This target molecule-based selection of experimental IDs facilitates the process of linking subprocesses into an overall integrated process for synthesizing the user- specified target material. Overall, enabling the selection of experimental IDs based on a user- defined target molecule allows the system to automatically identify the building blocks needed for a user's molecule of interest based on available experimental data.
[0062] Some embodiments provide the capability of tracking modifications made to the integrated process over time, as well as the resulting impacts on the aggregate process metric. Specifically, as changes or optimizations are made to the linked experimental IDs that comprise the integrated process, the system can monitor and log these modifications. Concurrently, the impacts of the changes on process metrics like the overall process mass intensity, solvent intensity, yield, cost, or other aggregate sustainability indicators can be reestimated and recorded. By tracking edits to the process flow alongside corresponding shifts in the process metrics, users can evaluate how impactful or beneficial particular process tweaks have been historically. This logging functionality enables data-driven analysis of how the integrated process has evolved regarding sustainability and can guide future optimization efforts. Overall, the ability to trace modifications and quantify associated effects on the aggregate process metric can facilitate systematic improvements over multiple iterations.
[0063] The method may further involve estimating the impacts of solvent recycling or waste stream recycling on the plurality of process metrics. In some embodiments, the process metrics estimator component estimates how recycling solvents or waste streams back into the integrated process influences metrics such as process mass intensity, solvent intensity, water intensity, and global warming potential. For example, solvent recycling can reduce the amount of fresh solvent utilized in the process, thereby lowering the solvent intensity. Similarly, recycling waste streams may decrease the material inputs needed, which can positively impact sustainability indicators like process mass intensity. Quantifying these recycling impacts provides additional insights into optimization opportunities for improving the environmental performance of the integrated process.
[0064] In some embodiments, the method may involve integrating the estimated process metrics and aggregate process metric with an enterprise resource planning (ERP) system. The ERP system can provide comprehensive data management capabilities to track materials, processes, inventory, orders, accounting, and other operational data across the enterprise. Integrating the sustainability metrics estimated by the system with the ERP allows for holistic tracking of environmental factors alongside traditional business metrics. This integration enables seamless monitoring and optimization of both process efficiency and environmental impact through a centralized platform. Overall, incorporating the calculated process metrics and aggregate sustainability indicators into an organization's broader ERP infrastructure can facilitate comprehensive data analysis to inform sustainable decision-making across chemical research and manufacturing activities.
[0065] In some embodiments, the method involves storing the one or more experimental IDs, the plurality of process metrics, and the aggregate process metric in a database. This database storage allows the experimental data, process metrics, and aggregate metrics to be saved for later retrieval and analysis. By storing this information, users can access a historical record of prior experiments, process metrics calculations, and overall sustainability indicators. The database also enables data sharing, collaboration between multiple users, integration with other systems, and long-term tracking of process improvements over time. In certain implementations, the stored data may be retrieved at a later point from the database for additional analysis, comparison to alternative processes, or to demonstrate progress in process optimization efforts. Overall, the database storage provides persistence and easy access to valuable experimental records and sustainability metrics.
[0066] Retrieval of stored data. In some embodiments, the method may further involve retrieving the one or more experimental IDs that were stored in the database, along with the associated plurality of process metrics and the aggregate process metric. This allows a user to revisit and analyze previous experimental data that had been saved, enabling the comparison of different process options or tracking changes in process metrics over time. The storage and retrieval capabilities facilitate data management and reuse, avoiding duplication of effort and taking full advantage of experimental data from past synthesis routes or process configurations. Through integration with the existing database architecture, users can conveniently restore experimental IDs and process metrics that pertain to earlier experiments or process alternatives. [0067] In some embodiments, the method allows for the adjustment of a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. The act of linking the plurality of experiment IDs is conducted to form an integrated process including all of the corresponding subprocesses in an ordered configuration to synthesize the target material. Thus, the one or more experimental IDs may be selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material. In a preferred embodiment, this linking of a one or more experimental IDs is done in an automatic manner, based on the input and output chemical structures in the experimental IDs. The plurality of process metrics may be represented as a ratio. For example, the ratio may be PMI over Yield.
[0068] The method may further involve updating the plurality of process metrics and the aggregate process metric based on edited input data associated with at least one of the selected experimental IDs. Specifically, the system enables users to manually edit the input data corresponding to the experimental IDs through the graphical user interface. For example, a user could modify the quantities of reagents used in a reaction or change the reaction conditions like temperature and pressure. After such edits are made, the process metrics estimator component automatically recalculates and updates the plurality of process metrics to reflect the changes in input data. This would update metrics like process mass intensity, solvent intensity, yield percentage, etc. in accordance with the manual edits. Additionally, since the aggregate process metric is a function of the individual process metrics, it is also updated accordingly when changes are made to the input data. This dynamic update allows users to immediately see the impact of any input data modifications on the overall process metrics and sustainability indicators. It facilitates rapid scenario testing and process optimization during chemical process development.
[0069] In some embodiments, the method includes displaying changes between the edited input data that the user manually adjusted and the original input data that was initially associated with at least one experimental ID. Visually highlighting the changes between the original and edited input data or providing a side-by-side comparison can make it easier for users to see where modifications have been made and understand the impacts on the estimated process metrics and aggregate process metric. This capability enhances transparency and facilitates iterative optimization of the integrated process through assessment of different input data scenarios.
[0070] The method includes enabling a user to manually edit input data associated with at least one of the one or more experimental IDs. The plurality of process metrics and aggregate process metrics are updated based on the edited input data. Changes between the edited input data and the original input data associated with the at least one experimental ID may be displayed. Displaying the changes may involve visually highlighting the differences between the edited and original input data within the user interface. Alternatively, displaying the changes may involve providing a side-by-side comparison of the original and edited input data, allowing the user to clearly see how the data has been modified.
[0071] In some embodiments, the plurality of estimated process metrics may be provided for the integrated process as a whole, for parts of the integrated process, or for each individual subprocess step corresponding to the respective experimental IDs. That is, the process metric estimates generated by the process metrics estimator component can be presented at different levels of granularity depending on the specific requirements. For example, a cumulative or aggregate PMI value may be calculated and displayed for the complete multi-step integrated process to synthesize the target material. Alternatively, separate PMI values may be estimated and shown for each distinct subprocess, giving insights into the PMI contributions of individual steps. As another option, PMI metrics may be calculated and visualized for logical subsections of the integrated process, such as a sequence of reactions or a particular purification train. This flexibility in process metric estimation and visualization at multiple levels enables detailed analysis and comparison of different process options.
[0072] Optionally, the process metrics estimated for the integrated process may be compared to one or more process metrics from other processes. In some embodiments, this comparison to external process metrics can be performed in order to facilitate evaluation of the relative performance of the integrated process under consideration. By benchmarking against standardized metrics or metrics from alternative production methods, the comparison can provide additional context and enable easier assessment of the efficiency, cost-effectiveness, sustainability, or other relevant parameters associated with the integrated process synthesized from the selected experimental IDs. The capability to draw such comparisons against suitable benchmarks expands the utility of the estimated process metrics and aggregate process metrics calculated by the system.
[0073] The method may further involve providing the plurality of process metrics for each individual subprocess within the integrated process formed by linking the experimental IDs. In some embodiments, the process metrics estimator component can estimate process metrics such as process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield, or other relevant metrics for each subprocess corresponding to the respective experimental IDs selected by the user. By making these process metrics available at the subprocess level, users can gain deeper insight into the factors driving the aggregate process metrics for the overall integrated process. This subprocess-level reporting of process metrics can also aid in identifying particular steps that have an outsized contribution to undesirable metrics, allowing users to focus process optimization and improvement efforts on these key subprocesses within the integrated flow. Thus, providing process metrics for each subprocess enables more targeted, granular analysis that supports enhancements in efficiency, sustainability, and environmental impact across the chemical synthesis route.
[0074] In some embodiments, the method includes comparing at least one process metric of the plurality of process metrics to a corresponding process metric from a different integrated process. For example, a process mass intensity or solvent intensity value estimated for one of the linked experimental IDs may be compared to process mass intensity or solvent intensity values obtained from a separate, distinct integrated process. This enables assessing the relative efficiency or sustainability of the process under consideration compared to alternative processes for producing the same or similar target compounds. The ability to benchmark against historical data or industry standards supports evaluating opportunities for improving the integrated process through suitable modifications to reaction conditions, workup procedures, or selection of starting materials and reagents.
[0075] In some embodiments, at least one process metric of the plurality of process metrics may be visualized using a Sankey diagram. The Sankey diagram provides an intuitive way to represent the flow of mass, energy, cost, or other metrics through the integrated process. This visualization can help highlight inefficiencies, mass imbalances, and opportunities for process improvement in a graphical manner. By mapping the process data onto the nodes and edges of the Sankey diagram, users can gain insight into how changes in one subprocess may propagate through the integrated process to impact other metrics of interest. Some implementations may allow users to interact with the Sankey diagram representation of the process metrics, for example by selecting specific pathways or modifying process parameters to observe the effect on the overall mass flow in real time. Thus, the Sankey diagram enables an interactive, graphical analysis that complements the detailed numeric process metrics estimated by the system.
[0076] In some embodiments, the method includes identifying recycled mass streams that are recycled to the beginning of a subprocess or the integrated process. Specifically, the act of searching for similar reaction subprocesses by the search component may involve recognizing mass streams from waste or byproducts that are recycled and fed back into an earlier subprocess or the start of the overall integrated process. By detecting these recycled mass streams, the system can analyze the impact of recycling on the estimated mass flow and process metrics. This recycling functionality provides another tool to explore optimization opportunities and improve the efficiency of the chemical synthesis process under analysis. The search component is configured to identify these recycled streams within the historical data and suggest integration opportunities accordingly.
[0077] In some embodiments, the range of values estimated for the prophetic process metric may correspond to a multi-step process rather than just a single reaction step. For example, if the prophetic experiment suggested by the prediction engine replaces multiple linked experimental IDs with a proposed multi-step synthesis route, the predicted process metrics for this overall pathway may include confidence intervals or credible regions. These uncertainty ranges can account for the inherent variability when predicting outcomes across several chemical transformations or over an extended process with several distinct steps. By propagating uncertainties across the various stages, reliable estimates for the overall reliability of the predicted process metrics and sustainability indicators can be provided even for complex, integrated processes. This allows users to realistically assess the accuracy of recommendations from in silico tools when dealing with intricate, multi-stage synthetic routes.
[0078] In some embodiments, the method involves automatically linking the one or more experimental IDs to form the integrated process synthesizing the target material. This automatic linkage is performed based on the input and output chemical structures specified in each of the experimental IDs. That is, the software analyzes the chemical structures entering and exiting each subprocess to determine compatibility and continuity of materials flow. It then automatically connects the experimental IDs in the proper order to construct the overall integrated process including all necessary subprocesses, without requiring extensive manual input or oversight from the user. By leveraging the input and output chemical structures stored within the experimental ID data, the linking process can be streamlined and expedited. This automation enables more rapid set up and estimation of mass flow and sustainability metrics for the chemical synthesis route.
[0079] In some embodiments, the plurality of process metrics are represented as a ratio. Specifically, the ratio can be process mass intensity over yield. The process mass intensity reflects the mass-based resource consumption per unit mass of product, while the yield represents the amount of product obtained from the process. Taking the ratio of these two metrics provides a useful indicator of the efficiency and environmental impact of the integrated process, accounting for both the resource usage and productivity. A lower ratio indicates a more efficient and sustainable process. By visualizing or tracking changes in this ratio, users can assess the impacts of modifications or alternate pathways on the overall process performance. Thus, representing the plurality of process metrics as a process mass intensity over yield ratio enables straightforward yet comprehensive evaluation of chemical synthesis processes.
[0080] Some embodiments provide a data processing system comprising means for implementing the method according to any of the preceding claims. In some embodiments, the system may include both hardware and software components configured to execute the computer-implemented method for calculating mass flow and estimating sustainability metrics. For example, the system may include one or more processors, memory, storage, network interfaces, databases, and other computing resources capable of carrying out the various steps involved in selecting experimental IDs, linking them to form an integrated process, estimating process metrics, providing a synthesis tree, searching for similar reactions, selecting alternative experimental IDs, and updating aggregate metrics. The system is designed to implement the full functionality of the method efficiently through specialized algorithms, predictive models, data structures, or other technological means. The modular, customizable architecture of the system also allows for easy extension or updating of capabilities in accordance with advancements in the field. Overall, embodiments provide an intelligent data processing system leveraging automation and advanced analytics to revolutionize mass flow calculations, process optimization, and sustainability assessment during chemical process development.
[0081] In some embodiments, the disclosure includes a computer program containing instructions that, when executed by a computer, cause the computer to perform the method described herein. This computer program allows the computer to select one or more experimental IDs, link the IDs to form an integrated process, estimate process metrics for each ID, calculate an aggregate process metric, provide a synthesis tree, search for similar reactions, select alternative IDs, and update metrics accordingly. By executing these instructions, the computer can fully implement the method for calculating mass flow and estimating sustainability metrics in chemical processes. The computer program enables automated calculation of metrics, exploration of alternatives, and optimization of chemical synthesis procedures in a computerized environment.
[0082] In some embodiments, the method may be implemented as a computer program with instructions stored on a computer-readable medium. When these instructions are executed by a computer, they cause the computer to carry out the method comprising: selecting one or more experimental IDs from an electronic notebook, linking the IDs to form an integrated process synthesizing a target material, estimating process metrics for each ID, calculating an aggregate metric, providing a synthesis tree, searching for similar reactions, selecting alternative IDs, and updating the aggregate metric accordingly. The computer-readable medium allows the computational implementation of the method, enabling the computer to perform the selection, linking, estimation, calculation, provision, searching, alternative selection, and updating involved in analyzing mass flow and sustainability metrics for chemical processes. Thus, through stored instructions, a computer can carry out the full method by accessing and executing the appropriate program from the computer-readable medium on which it resides.
[0083] BRIEF DESCRIPTION OF THE DRAWINGS
[0084] These and other aspects will become more apparent from the following detailed description of the various embodiments of the present disclosure with reference to the drawings wherein: [0085] Fig. 1 shows a block diagram illustration of a cloud-based system to calculate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure;
[0086] Fig. 2 show a block diagram illustration of a computing device to estimate mass flow and estimate sustainability metrics in chemical synthesis in accordance with an embodiment of the present disclosure;
[0087] Fig. 3 is a flow-chart diagram of a method of estimating mass flow and estimate sustainability metrics in accordance with an embodiment of the present disclosure;
[0088] Fig. 4 shows a GUI of linked experimental IDs in accordance with an embodiment of the present disclosure;
[0089] Fig. 5 shows a summary of the output of the of linked experimental IDs of Fig. 4 in accordance with an embodiment of the present disclosure;
[0090] Fig. 6 shows a synthesis tree in accordance with an embodiment of the present disclosure;
[0091] Fig. 7 shows an example result of the search component, in which the input data can be modified in accordance with an embodiment of the present disclosure;
[0092] Figs. 8A and 8B show an example result of the search component in which the mass flow is visualized in accordance with an embodiment of the present disclosure;
[0093] Fig. 9 shows an example result of the search component in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 6 in accordance with an embodiment of the present disclosure;
[0094] Fig. 10 shows PMI vs. Yield values for each alternative experimental ID that the search component identifies in accordance with an embodiment of the present disclosure;
[0095] Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow is estimated in accordance with an embodiment of the present disclosure;
[0096] Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure;
[0097] Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations in accordance with an embodiment of the present disclosure;
[0098] Fig. 14 shows model compounds for raw materials to synthesize APIs in accordance with an embodiment of the present disclosure;
[0099] Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure; [00100] Fig. 16 shows an average PMI and chemical yield for similar reactions of one synthesis step that was suggested by a retro-synthesis tool using data is obtained using the similarity search function in accordance with an embodiment of the present disclosure;
[00101] Fig.17 shows the Input metric for the energy calculation module in accordance with an embodiment of the present disclosure;
[00102] Fig.18 shows the output metric for the energy calculation module in accordance with an embodiment of the present disclosure; and
[00103] Fig. 19 shows the GUI displaying weight based carbon footprints received from an external database, which are assigned based on data of individual raw materials, on a general role, or manually edited data in accordance with an embodiment of the present disclosure.
[00104] DETAILED DESCRIPTION
[0001] Fig. 1 shows a block diagram illustration of a cloud-based system 100 to estimate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure. The system 100 automatically calculates metrics without significant manual interaction or input. The system 100 includes a cloud service provider 102, one or more personal computers 104, and a mobile device 106. The system 100 also includes a mass Flow-estimator component 112. The mass flow-estimator component 112 can estimate mass flow in chemical synthesis and enable a user to modify the chemical synthesis process to change or reduce the mass flow as described herein. The mass flow estimator component 112 can be used in chemical process development to optimize certain values, such as costs, yield or environmental compatibility. In order to record the effects of parameter adjustments, these may be calculated early in the development and updated regularly. These data may be stored in a data pool, e.g., in subprocess metrics 144 of a database 132. The database 132 may be accessed via standard interfaces, such as Oracle DB interfaces.
[0002] The system 100 can be used for estimating mass flow in a chemical synthesis or an industrial process. For example, the system 100 may be used in the pharmaceutical industry to estimate the mass flow in drug manufacturing processes, may be used in the chemical industry to estimate the mass flow in the production of various chemical compounds or formulations, and/or may be used in the biotechnology industry to estimate the mass flow in the production of biologies.
[0003] The cloud service provider 102 may be configured to provide remote capabilities to estimate mass flow in a chemical synthesis or industrial process by providing remote access to the mass-flow estimator component 112. In some embodiments of the present disclosure, the cloud service provider 102 may be a hosted service such as a company that offers cloud computing services to businesses and individuals such that the cloud service provider 102 provides the infrastructure, software, and platforms required to host, manage, and deliver cloud-based services. For example, in some embodiments, the cloud service provider 102 may provide infrastructure as a service, platform as a service, software as a service, and/or may be an interface into a blockchain infrastructure that may or may not be hosted by the cloud service provider 102. The cloud service provider 102 may be configured to scale up or down its computing resources based upon demand from users at a given moment.
[0004] In yet other embodiments, the cloud service provider 102 may be implemented on a block chain. The cloud service provider 102 may utilize a distributed ledger to store and verify mass flow estimates generated by the mass-flow estimator component 112. Users may be authenticated and/or authorized by a secure digital certificate, encryption key, single-sign on, or other secure mechanism. The data may be stored and calculated in a secure and tamper proof manner to provide transparency and accountability to all users. The smart contracts may include executable code that defines a manufacturing process in terms of one or more or the components within the mass-flow estimator component 112 in a manner consistent with transparency and security settings.
[0005] Referring generally to the system 100, the personal computers 104 and the mobile device 106 communicate with each other via a network 108. The network 108 may be Wi-Fi, ethernet, Bluetooth, etc. and may utilize the internet and associated protocols, such as TCP/IP. The network 108 may be a local area network, a wide-area network, a physical bus (such as a Universal Serial Bus), the internet, or some combination thereof.
[0006] The personal computer 104 and mobile device 106 may interface with the cloud service provider 102 to the calculate mass flow in a chemical or industrial process as determined by the subprocess metrics 144. In some embodiments, a specialized application for interfacing with the mass-flow estimator component 112 may be used, such as a mobile application on the mobile device 106 or a desktop application on the personal computer 104. The communications may include transmitting data in HTML, XML, JSON, YAML, or any data format. The mass-flow estimator component 112 may provide user-level accounts to individuals through a typical login mechanism. The mass-flow estimator component 112 may be a web application, a webserver, a web service, etc. and may utilize one or more protocols to communicate data.
[0007] The cloud service provider 102 may provide the mass-flow estimator component 112 as a webpage, a webapp, a program for download and execution on the computer 104 or the mobile device 106. The mass-flow estimator component 112 includes various sub-components, such as an electronic notebook component 150, a component for linking experimental IDs 142, a process metrics estimator component 152, a synthesis tree component 156, a search component 158, a GUI component 160, and a prediction engine 164. [0008] The mass-flow estimator component 112 of the system 100 may be implemented as a cloud-based software application accessible via a web interface. The system 100 may also be implemented as a standalone software application installed on a local computer or server. In some embodiments, the system 100 may include specialized hardware such as a high-performance computer or server, or other types of computing and networking equipment.
[0009] The mass flow calculator component 112 of the system 100 may be implemented on a variety of computing platforms, including desktop, laptop, server, or cloud-based configurations. The software may be written in various programming languages, such as Python, Julia, Java, C++, or other languages. The system 100 may also incorporate various data visualization tools, such as d3.js, Plotly.js, or Matplotlib, to allow users to visualize the experimental data and the process metrics in an intuitive manner.
[0010] The electronic notebook component 150 enables the selection of a one or more experimental IDs 142 that are stored within a database 132. The electronic notebook component 150 may use a REST API and may be accessed via the Mass flow calculator component. Each experimental ID 142 corresponds to a subprocess having an input material and an output material. Information about each experimental ID 142 may be found in the subprocess metrics 144 also included in the database 132. The subprocess metrics 144 may include historical data such that each experimental ID 142 can be associated within one or more historical experiments as found within the subprocess metrics 144 to estimate or predict a metric.
[0011] The linking component 154 links the one or more experimental IDs 142 to form an integrated process including all of the corresponding subprocesses, thereby synthesizing a target material. The process metrics estimator component 152 estimates a plurality of process metrics (e.g., using the subprocess metrics 144), each of which corresponds to a respective experimental ID of the one or more experimental IDs 142. Fig. 4 and Fig 5. shows the linked experiments IDs 402 with the molecular mode 404, the PMI values without cleaning 406, the PMI values 408, the estimated carbon footprint (cradle-to-grave) 410, the solvent intensity 412, the water intensity 414, and the estimated carbon footprint (gate-to-gate) 416 as a result for each step or summarized over the reaction sequence. The process metrics estimator component 152 also estimates an aggregate process metric that is a function of the estimated plurality of process metrics (which can be column summations).
[0012] The process metrics estimator component 152 can estimate various process metrics related to the chemical synthesis process. This component can be used for evaluating the efficiency, sustainability, and environmental impact of the integrated process formed by linking the selected experimental IDs.
[0013] The process metrics estimator component 152 estimates a plurality of process metrics, where each of these metrics corresponds to a respective experimental ID from the one or more experimental IDs selected by the user. These process metrics can encompass a wide range of parameters, including but not limited to process mass intensity (PMI), solvent intensity (SI), water intensity (WI), global warming potential (GWP), cost, yield, and various environmental metrics.
[0014] The process metrics estimator component 152 may calculate these metrics based on the experimental data associated with each experimental ID, which may include information from one or more previously performed experiments. This data can be retrieved from the subprocess metrics 144 stored in the database 132 or obtained from other relevant sources.
[0015] In some embodiments, the process metrics estimator component 152 may employ advanced algorithms, machine learning techniques, or empirical models to estimate the process metrics accurately. It may consider factors such as reaction conditions, reagent quantities, solvent usage, energy consumption, and other relevant parameters to derive these metrics.
[0016] The process metrics estimator component 152 may be used to estimate an aggregate process metric corresponding to the integrated process. This aggregate process metric is a function of the estimated plurality of process metrics for the individual experimental IDs. The function used to calculate the aggregate process metric can vary depending on the specific requirements and the nature of the process being analyzed.
[0017] In some embodiments, the aggregate process metric may be a simple summation of the individual process metrics. For example, if the process metrics being considered are PMI values for each subprocess, the aggregate process metric could be the sum of these PMI values, representing the overall PMI for the integrated process.
[0018] In some embodiments, the aggregate process metric may be a weighted combination of the individual process metrics, where different weights are assigned to different metrics based on their relative importance or impact on the overall process. For instance, the aggregate process metric could be a weighted sum of PMI, SI, WI, and GWP, reflecting the combined environmental impact of the integrated process. [0019] In other embodiments, the aggregate process metric may be a more complex function that incorporates additional factors or constraints. For example, the aggregate process metric could be a multi-objective optimization function that considers not only the process metrics but also factors such as cost, yield, or specific environmental targets.
[0020] The process metrics estimator component 152 may also provide the capability to compare the estimated process metrics and the aggregate process metric with corresponding metrics from other processes or industry benchmarks. This comparison can aid in evaluating the relative performance and sustainability of the integrated process under consideration.
[0021] Furthermore, the process metrics estimator component 152 may interact with other components of the system 100, such as the search component 158 and the prediction engine 164, to explore alternative experimental IDs or prophetic experiments. When an alternative experimental ID or a prophetic experiment is selected, the process metrics estimator component 152 can update the aggregate process metric accordingly, reflecting the impact of the proposed change on the overall process metrics.
[0022] In some implementations, the process metrics estimator component 152 may provide visualizations or graphical representations of the estimated process metrics and the aggregate process metric. These visualizations can aid in interpreting the data and identifying areas for potential improvement or optimization.
[0023] In some specific embodiments, the process metrics estimator component 152 can be designed with a modular architecture, allowing for the integration of various modules or sub-components responsible for estimating specific process metrics. This modular approach enables flexibility, scalability, and customization, as different modules can be added, removed, or updated independently based on the specific requirements or advancements in the field.
[0024] For instance, one module within the process metrics estimator component 152 could be dedicated to estimating process mass intensity (PMI) and solvent intensity (SI). This module may leverage advanced algorithms and machine learning techniques to analyze the input and output materials, reaction conditions, and solvent usage data associated with each experimental ID. It may also incorporate industry-specific heuristics or empirical models to improve the accuracy of PMI and SI estimations.
[0025] Another module could focus on estimating water intensity (WI) and global warming potential (GWP). This module may integrate with external databases or life cycle assessment (LCA) tools to obtain relevant data on the environmental impact of various materials and processes involved in the chemical synthesis. It could also employ computational fluid dynamics (CFD) simulations or thermodynamic calculations to estimate the energy consumption and associated GWP contributions.
[0026] Yet another module within the process metrics estimator component 152 could be responsible for estimating cost-related metrics, such as raw material costs, energy costs, and overall process costs. This module may interface with enterprise resource planning (ERP) systems, supply chain management systems, or market data feeds to obtain up-to-date information on material prices, energy costs, and other relevant cost factors.
[0027] In some embodiments, the process metrics estimator component 152 may incorporate advanced uncertainty quantification techniques to provide confidence intervals, credible regions, or probability distributions for the estimated process metrics. These uncertainty estimates can be particularly valuable when dealing with incomplete or uncertain input data, or when accounting for inherent variability in the chemical synthesis processes.
[0028] The process metrics estimator component 152 may also be integrated with optimization algorithms or decision support systems. In such cases, the estimated process metrics could serve as objective functions or constraints in the optimization process, enabling the identification of optimal process parameters, reaction conditions, or material selections that minimize undesirable metrics (e.g., PMI, GWP) while maximizing desirable ones (e.g., yield, cost-effectiveness).
[0029] Furthermore, the process metrics estimator component 152 could be designed to leverage distributed computing resources or cloud-based platforms. This would allow for parallel processing of multiple experimental IDs or the execution of computationally intensive simulations or calculations required for estimating certain process metrics. Cloud-based deployment could also facilitate collaborative efforts, where multiple users or research teams can contribute experimental data and access the estimated process metrics from various locations.
[0030] In some implementations, the process metrics estimator component 152 may incorporate machine learning capabilities to continuously improve its estimation accuracy. As more experimental data becomes available, the component could retrain its models or update its algorithms, leveraging techniques such as transfer learning or active learning to enhance the estimation performance over time.
[0031] The process metrics estimator component 152 could also be integrated with advanced visualization tools or dashboards, providing intuitive and interactive representations of the estimated process metrics. These visualizations could include interactive charts, graphs, Sankey diagrams, or even virtual reality (VR) or augmented reality (AR) environments, allowing users to explore the data from different perspectives and gain deeper insights into the chemical synthesis process.
[0032] In some embodiments, the process metrics estimator component 152 may support the integration of user-defined metrics or custom calculations. This would enable users to incorporate domain-specific knowledge, proprietary algorithms, or specialized requirements into the estimation process, tailoring the component to their specific needs or industry standards.
[0033] The values shown in Fig. 4 may be values for a particular batch run or for a per kilogram of target material or may use any other unit of measurement. Fig. 5 is the GUI showing a summary of the outputs of the linked experimental IDs of Fig. 4.
[0034] The synthesis tree component 156 provides a synthesis (illustrated in Fig. 6) as a synthesis tree 500 (as shown in the GUI in Figs. 7) corresponding to the one or more experimental IDs 142. The search component 158 is configured for searching for structurally similar reaction subprocesses for each of the experimental IDs 142 to determine a plurality of similar experiments.
[0035] The search component 158 may be configured to search for similar reaction subprocesses for each of the experimental IDs 142 stored in the database 132. This search can help determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the selected one or more experimental IDs 142.
[0036] In some embodiments, the search component 158 may identify similar experiments by looking for experiments with the same primary input molecule and target molecule as one or more of the selected experimental IDs 142. The search component 158 may ignore or consider helper molecules, such as buffers or catalysts, when identifying similar experiments. [0037] The search component 158 can facilitate the selection of an alternative experimental ID 142 for at least one of the initially selected one or more experimental IDs 142. This selection of an alternative experimental ID 142 can be made through the GUI component 160, which may display the search results from the search component 158. The process metrics estimator component 152 can then update the aggregate process metric in accordance with the selected alternative experimental ID 142.
[0038] In some embodiments, the search component 158 may identify recycled mass streams that are recycled to the beginning of a subprocess or the integrated process. This information can be used to estimate the effect of solvent recycling on the mass flow and process metrics, as illustrated in Fig. 11 of the patent application. [0039] The search component 158 may utilize various techniques to identify similar experiments, such as chemical structure matching, reaction type matching, or machine learning models trained on historical experimental data. The search component 158 can provide a tool for process optimization by allowing users to explore alternative experimental IDs 142 and their corresponding process metrics, enabling informed decision-making during chemical process development.
[0040] The search component 158 may work in conjunction with other components of the mass-flow estimator component 112, such as the synthesis tree component 156, which provides a visual representation of the linked experimental IDs 142, and the prediction engine 164, which can suggest prophetic experiments based on predictive models or retrosynthesis tools.
[0041] There are various ways in which the search for similar reaction subprocesses can be carried out. In one embodiment, the search component 158 may employ chemical structure matching algorithms to identify experimental IDs 142 that involve the same input and output molecular structures. This approach can be particularly useful when the goal is to find alternative synthetic routes or optimize specific reaction steps within the integrated process.
[0042] The search component 158 may maintain a database or index of chemical structures associated with each experimental ID 142, enabling efficient querying and matching based on structural similarity metrics. These metrics can consider factors such as functional groups, bond connectivity, stereochemistry, and other relevant structural features. The search component 158 can utilize well-established chemical structure representation formats, such as SMILES, InChi, or molecular fingerprints, to facilitate structure matching and comparison.
[0043] In another embodiment, the search component 158 may employ reaction type matching, where it identifies similar experiments based on the types of chemical reactions involved. This approach can be beneficial when optimizing specific reaction classes or exploring alternative reagents or conditions for a particular transformation. The search component 158 may leverage reaction classification algorithms or databases that categorize reactions based on their mechanistic patterns, functional group transformations, or other relevant criteria.
[0044] Additionally, the search component 158 can incorporate machine learning techniques to identify similar experiments. In this embodiment, the search component 158 may train machine learning models on historical experimental data, including reaction conditions, reagents, yields, and other relevant features. These models can learn patterns and similarities between experiments, enabling the search component 158 to suggest relevant alternatives based on the provided experimental IDs 142. [0045] The machine learning models employed by the search component 158 can range from simple similarity measures, such as k-nearest neighbors or cosine similarity, to more complex models like neural networks or decision trees. The choice of model can depend on factors such as the complexity of the chemical space, the availability of training data, and the desired trade-off between accuracy and computational efficiency.
[0046] In some embodiments, the search component 158 may combine multiple search strategies, such as chemical structure matching and reaction type matching, to provide more comprehensive and relevant search results. This hybrid approach can leverage the strengths of different search techniques and provide users with a diverse set of alternative experimental IDs 142 to explore.
[0047] Furthermore, the search component 158 may incorporate user feedback and preferences to refine and enhance its search capabilities over time. For example, users may be able to provide ratings or rankings for the search results, which can be used to update the underlying search algorithms or models. This adaptive approach can ensure that the search component 158 remains relevant and aligned with the specific needs and objectives of the users. [0048] The search component 158 may also integrate with external databases or resources to expand its search capabilities. For instance, it may access public or proprietary reaction databases, literature repositories, or patent databases to identify additional relevant experiments or synthetic routes. This integration can provide users with a broader perspective and facilitate the exploration of alternative experimental IDs 142 beyond the scope of the local database 132. [0049] In addition to suggesting alternative experimental IDs 142, the search component 158 may provide additional information or analysis to support the decision-making process. This information can include statistical summaries of the identified similar experiments, such as distributions of yields, process metrics, or reaction conditions. The search component 158 may also highlight key differences or similarities between the alternative experimental IDs 142 and the initially selected ones, enabling users to make informed choices based on their specific requirements or constraints.
[00105] The searching may be activated by the button 502 on the webpage as is shown in Fig. 7. Also shown in the GUI on Fig. 7, is that the molecular role 504 is selectable as well. When the GUI of Fig. 7 is presented to a user, it shows the synthesis steps with to thereby give the user the ability to modify the reactants. Although Fig. 7 only shows 2 steps, additional steps of the process may be shown by allowing the user to scroll down. Fig. 12 shows an example result for the synthesis tree editing function in accordance with an embodiment of the present disclosure where there is a drop-down menu to change experiments. [00106] Referring again to Fig. 1, each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs 142. The GUI component 160 is configured for selecting an alternative experimental ID 142 for at least one of the one or more experimental IDs 142. Figs 8A and 8B show Sankey Diagrams representing the mass flow from the selected experiment, e.g., the experiment of Fig. 6. Fig. 9 shows an example result of the search component 158 in which an alternative experimental ID can be selected to replace one or more of the processes shown in Fig. 7. Thus, Fig. 9 shows alternative experimental IDs and in some embodiments, it can be matched up to the ones that it replaces via another column, for example. The resulting PMI vs. Yield values for each alternative experimental ID may be shown as in Fig. 10. Referring yet again to Fig. 1, the process metrics estimator component 152 updates the aggregate process metric in accordance with the alternative experimental ID 142 when a user selects the alternative experimental ID 142 to replace one of the initially entered in experimental ID 142. Fig. 11 shows an example result of the search component in which the effect of solvent recycling on mass flow starting materials’, e.g., reactants, catalysts, etc. is estimated for the experiment. Thus, Fig. 11 shows another view off the GUI.
[0050] In another embodiment, the synthesis tree component 156 of the system 100 may be used to identify a set of possible synthesis routes leading to a target material. The search component 158 may be used to identify a set of similar experiments to each experimental ID 142 of the experimental IDs 142 and determine an optimal set of reactions to synthesize the target material Then, the process metrics estimator component 152 can estimate the process metrics for each of the optimal set of reactions, and the aggregate process metric for the integrated processes. The GUI component 160 can be used to visualize the synthesis route and adjust the experimental IDs 142, as well as adjust the process parameters to optimize the process metrics.
[0051] The experimental data corresponding to each experimental ID 142 includes information about one or more previously performed experiments which are stored in the subprocess metrics 144. The aggregate process metric can be a process mass intensity, a solvent intensity, a water intensity, a global warming potential, cost, yield, an environmental metric, etc.
[0052] The GUI component 160 of the system 100 may include various user interface elements, such as dropdown menus, checkboxes, sliders, and text input fields, to enable users to select experimental IDs 142, adjust process parameters, and visualize the synthesis route. The GUI component 160 can also be designed to be accessible via a web interface, allowing users to access the system 100 from anywhere with an internet connection. [0053] The GUI component 160 is configured for selecting a role for a molecule such that the molecule is associates with an experimental ID 142 of the one or more experimental IDs 142. The GUI component 160 can also be configured for adjusting a process parameter and updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. For example, some of the experimental IDs 142 may have input parameters such that the output values (e.g., PMI, yield, etc.) vary based upon the input parameters. Modifying the input parameters can update these output values. The act of selecting the one or more experimental IDs 142 is performed by a user utilizing the GUI component 160.
[0054] The GUI component 160 can also be used to query a prediction engine 164 to determine a prophetic experiment. It can be used to replace at least one of the one or more experimental IDs 142 with the prophetic experiment and estimate a prophetic process metric corresponding to the prophetic experiment. The system 100 can also update the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs 142. The prophetic process metric can include a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region.
[0055] The plurality of process metrics can be a ratio, such as PMI over Yield. The one or more experimental IDs 142 can be selected out of order, and the act of linking the plurality of experiment IDs can include the act of ordering the one or more experimental IDs 142 to form the integrated process including all of the corresponding subprocesses, in order to synthesize the target material. Overall, the system 100 is designed to efficiently estimate mass flow and enable users to adjust and optimize various process metrics to improve the efficiency and sustainability of chemical or industrial processes.
[0056] In some embodiments, the mass-flow estimator component 112 may assign a confidence score. If the mass-flow estimator component 112 bases some or all of the estimates on data, models generated from data, or Monte Carlo simulation data, a confidence score can be assigned to the estimates of mass flow values to indicate the quality of the estimate. This could be included directly in the output or be derived from the standard deviation or variance in the sample data used to make the estimation, the min-max of mass flow data points, etc.
[0057] In one embodiment, a frequentist confidence store may be derived using frequentist statistics. For example, a confidence score may use sample data of a distribution, hypothesis testing, p-values, significance testing, confidence intervals etc. In additional embodiments, a confidence score is calculated for each (or a set of) sample values using posterior probabilities in a Bayesian estimate, which represent the updated belief about the mass flow values. Thus, the confidence score, for example, may be a credible interval of a posterior distribution or of a Bayesian estimator.
[0058] The mass-flow estimator component 112 also includes the communications component 162. The communications component 162 may facilitate seamless communication and data exchange between multiple software applications, devices, and systems. That is, the communications component 162 may include protocol handling, message formatting, data serializing, encryption, and authentication to facilitate the communication with the computers 104 and/or the mobile device 106. The communications component 162 may utilize a message formatting mechanism to format the messages into formats, such as XML, JSON, binary formats, and/or proprietary message formats. The communications component 162 may utilize various encryption algorithms, such as RSA, AES, ECC, symmetric encryption, asymmetric encryption etc. to enable secure communications between the mass-flow estimator component 112 and the computers 104 and/or the mobile device 106.
[0059] The GUI component 160 can render a display for use by the computer 104 and/or the mobile device 106. The GUI component 160 may be a webpage-based provider, such as flask, an HTML server, a web framework, etc. The GUI component 160 may provide widgets, information, buttons, options, and menus to thereby facilitate a user’s interaction with the mass-flow estimator component 112.
[0060] The GUI component 160 can be used to log into user accounts 146 so that a user can create, save, or retrieve the experimental IDs 142 and/or the subprocess metrics 144, or otherwise interface with any account features. Additionally or alternatively, the GUI component 160 can save favorites, select default parameters, or adjust default values. The GUI component 160 can direct other components to execute instructions based upon a workflow initiated by a user. That is, the GUI component 160 may receive events, such as a mouse click, button press, or GUI widget interaction to initiate a routine, series of steps, or series of acts. For example, the GUI component 160 may guide a user step-by-step on how to set up and work with the subprocess metrics 144 within the database 132.
[0061] The GUI component 160 may also be used to visualize the results of the process models and the waste estimates, in aggregate, in simulation, and/or may provide various visualization tools to analyze the data. The data may be stored in the database 132 or the Electronic Notebook Component 150. Thus, each user can log into a user account 146 to visual the results of their processes, the results of modifications to their processes on the entire synthesis chain, obtain a direct comparison and/or historical accuracy of their process mass flow values. [0062] A resource dispatcher 110 may dispatch requests to perform an action to one or more virtual servers 122, each of which has a virtual processor 124, a virtual memory 126, and a virtual disk space 128. The virtual servers 122 can be executed on one or more servers 121 on a server farm 119 as dispatched and activated by the resource dispatcher 110.
[0063] Fig. 2 show a block diagram illustration of a computing device 200 to calculate mass flow in chemical synthesis in accordance with an embodiment of the present disclosure. The computing device 200 of Fig. 2 may be the computer 104 or mobile device 106 of Fig. 1. The computing device 200 includes an VO interface 210 to communicate therewithin. The computing device 200 includes a data store 204, a processor 206, a network interface 208, a memory 225, and user I/O devices 226. The data store 204 stores data and may be a hard drive, flash drive, thumb drive, volatile memory, non-volatile memory, semi-volatile memory etc. The processor 206 can execute one or more processor-executable instructions 212, which may be stored in the data store 204 and/or the memory 225. For example, the processor 206 can execute processor-executable instructions 212 stored in memory 225 that was retrieved from the data store 204. The memory 225 also includes program data 214 that may include information related to the processor-executable instructions 212. The computing device 200 may include user I/O devices 226, such as a cursor device 230 (e.g., touchscreen or mouse), a keyboard 232 (virtual or physical), and/or a monitor 228 (which may be a touchscreen). The computing device 200 communicates with the network 202 via a network interface 208.
[0064] Although the computing device 200 of Fig. 2 may be used as part of the system 100 of Fig. 1, in some embodiments, the mass flow calculation functionality may reside wholly within the computing device 200 of Fig. 2. For example, the mass-flow estimator component 112 of Fig. 1 may reside within the processor-executable instructions 212 of Fig. 2 as mass flow calculation and mass-flow estimator component 242. The electronic notebook component 252, the process metrics estimator component 254, the linking component 256, the synthesis tree component 258, the search component 260, the GUI component 262, the communications component 265, and the prediction component 266 of Fig. 2 may be the same or similar to the electronic notebook component 150, the process metrics estimator component 152, the linking component 154, the synthesis tree component 156, the search component 158, the GUI Component 160, the communications component 162, and the prediction engine 164 of Fig. 1, respectively.
[0065] The database 244 may be similar to the database 132 of Fig. 1. The database 244 may, for example, be an Oracle, SQLite, postgresql, MariaDB, MySQL, or any other database embedded on the computing device 200. Thus, the experimental IDs 246, the subprocess metrics 248, and the user accounts 250 of Figs. 2 may be similar or identical to the experimental IDs 142, the subprocess metrics 144, and the user accounts 146 of Fig. 1, respectively.
[0066] Thus, in some embodiments the mass-flow estimator component 242 may reside wholly on a local device (such as on the computers 104, the mobile device 106, etc.) may be partially within a cloud service provider 102, and/or may be organized in a hybrid local and cloud configuration. In some embodiments, the mass-flow estimator component 242 may be an application, may be executed on the computers 104, the mobile device 106, the cloud service provider 102, the computing device 200, etc. or some combination thereof.
[0067] Fig. 3 is a flowchart of an example process 300. In some implementations, one or more process blocks of Fig. 3 may be performed by a device, such as a computing device. Starting with Act 302, the process involves selecting one or more experimental IDs from an electronic notebook. Each experimental ID corresponds to a subprocess that has at least one input material and at least one output material. Those input and output materials might be the same or different. The experimental ID may include experimental data corresponding to one or more previously performed experiments. The experimental IDs may be selected by using a GUI, such as that shown in Fig. 4. Or, Fig. 4 may show the listed of experimental IDs after a list is inputted into a GUI and may be after linking. Next, in Act 304, a role is selected for a molecule associated with an experimental ID from the One or more experimental IDs. For example, Fig. 7 shows a drop-down box where a role can be selected.
[0068] Then, in Act 306, the one or more experimental IDs are linked to form an integrated process that includes all of the corresponding subprocesses, which together synthesize a target material. This is done automatically or with user interaction. Act 308 involves estimating a plurality of process metrics, where each of the metrics corresponds to a respective experimental ID from the one or more experimental IDs. In Act 310, an aggregate process metric is estimated, which corresponds to the integrated process, and is a function of the estimated plurality of process metrics. The aggregate process metrics may be, for example, a process mass intensity, a solvent intensity, a water intensity, a global warming potential, a cost, an energy uptake, a waste score, a yield, and any other environmental metric. The process then moves to Act 312, which involves providing a synthesis tree (e.g., as shown in Fig. 5) that corresponds to the one or more experimental IDs.
[0069] Act 314 searches for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, where each similar experiment corresponds to one or more respective experiments of the one or more experimental IDs. In some embodiments, similar experiments are experiments with a primary input molecule and a target molecule that is the same as one or more experimental IDs. Helper molecules, such as buffers, may be different. Act 316 selects an alternative experimental ID for at least one of the one or more experimental IDs. Then, in Act 318, the aggregate process metric is updated in accordance with the alternative experimental ID. In Act 320, a process parameter is updated. Act 322 involves updating at least one of the plurality of process metrics in accordance with the adjusted process parameter. Act 324 queries a prediction engine to determine a prophetic experiment. Act 326 then replaces at least one of the one or more experimental IDs with the prophetic experiment.
[0070] In Act 328, a prophetic process metric is estimated, which corresponds to the prophetic experiment, and the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs. The prophetic process metrics can have a range of values, such as a confidence interval, a confident region, a credible interval, or a credible region, and the multiple process metrics can be expressed as a ratio, such as PMI over Yield.
[0071] Finally, in Act 330, the aggregate process metric is updated to include the prophetic experiment in place of the at least one of the one or more experimental IDs. Fig. 13 shows an example for displaying a comparison of at least two mass flow calculations so that different so that different experiments may be compared.
[0072] The process allows for the consideration and comparison of experimental data from one or more previously conducted experiments within each experimental ID. The aggregate process metric may be based on process mass intensity, solvent intensity, water intensity, global warming potential, cost, yield or environmental impact, either alone or in combination with other customizable variations. Furthermore, a molecule associated with an experimental ID can be selected to play a specific role (e.g. product, starting material, reagent, solvent, catalyst), while process parameters can be adjusted to achieve optimal results. In some cases, process metrics can be updated to achieve target parameters. The user can select experimental IDs using a graphical user interface. Additionally, the one or more experimental IDs can be selected out of order, but must be properly ordered to form an integrated process that includes all corresponding subprocesses that synthesize the target material. The customizable variations can be implemented separately or in combination with one another, and the process may include more, fewer, or differently arranged blocks than those illustrated in Figure 3. The blocks in the process can be performed simultaneously when necessary. Overall, the process can provide a versatile and customizable process for optimizing chemical reactions through the consideration of various experimental and process parameters.
[0073] Although Fig. 3 shows example blocks of process 300, in some implementations, process 300 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in Fig. 3. Additionally, or alternatively, two or more of the blocks of process 300 may be performed in parallel.
[0074] Figs. 14, 15 and 16 are described as follows to illustrate the operation of the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2 The predictive estimation of sustainability metrics (solvent intensity, water intensity, product carbon footprint) with an electronic notebook-data using the system 100 of Fig. 1 can utilize a prediction engine, e.g., retro-synthesis tools The production of highly complex molecules such as APIs (active pharmaceutical ingredients) is a multi-step synthesis process that starts from already complex raw materials. Such building blocks are for example the model compounds shown in Fig. 14.
[0075] The prediction engine 164 is a component of the mass-flow estimator component 112 in the system 100 for calculating mass flow and estimating sustainability metrics in chemical synthesis. This prediction engine 164 can serve various purposes and functionalities within the overall system.
[0076] In one embodiment, the prediction engine 164 may be queried to determine a prophetic experiment. The term "prophetic experiment" can refer to a hypothetical or simulated experiment that has not yet been physically performed. The prediction engine 164 can utilize various techniques, such as machine learning models, computational chemistry methods, or rule-based algorithms, to generate predictions or suggestions for potential chemical reactions or processes that could be explored. These predicted or prophetic experiments can then be incorporated into the overall integrated process being analyzed by the system.
[0077] Specifically, the prediction engine 164 may be used to replace at least one of the one or more experimental IDs 142 with a prophetic experiment suggested by the prediction engine 164. The process metrics estimator component 152 can then estimate a prophetic process metric corresponding to this prophetic experiment. The aggregate process metric, which is a function of the estimated plurality of process metrics, can be updated to include the prophetic experiment in place of the replaced experimental ID(s) 142.
[0078] The prophetic process metric estimated by the process metrics estimator component 152 for the prophetic experiment may include a range of values. This range of values can take various forms, such as a confidence interval, a confidence region, a credible interval, or a credible region. These ranges can provide an indication of the uncertainty or variability associated with the predicted or simulated experiment, allowing users to assess the reliability or robustness of the predictions made by the prediction engine 164.
[0079] In some embodiments, the prediction engine 164 may employ retrosynthesis tools or algorithms to suggest potential synthesis routes or pathways for a target molecule. These retrosynthesis tools can analyze the target molecule and propose a series of chemical reactions or transformations that could be used to synthesize the target from simpler starting materials or building blocks. The prediction engine 164 can then use this suggested synthesis route, along with data from the electronic notebook component 150 and the subprocess metrics 144, to estimate process metrics and sustainability indicators for the proposed synthetic pathway.
[0080] Additionally, the prediction engine 164 may be integrated with other components of the system, such as the search component 158, to identify similar reactions or processes from the database 132 that could inform or refine the predictions made by the prediction engine 164. For example, the search component 158 may identify similar experimental IDs 142 or subprocesses that have been previously performed, and the prediction engine 164 can use this information to improve the accuracy or reliability of its predictions.
[0081] The prediction engine 164 can also be used in conjunction with the synthesis tree component 156 and the GUI component 160. The synthesis tree component 156 provides a visual representation of the integrated process, including the one or more experimental IDs 142 and their corresponding subprocesses. The prediction engine 164 may suggest modifications or alternatives to this synthesis tree, which can be visualized and edited through the GUI component 160, allowing users to explore different scenarios and optimize the integrated process based on the predictions made by the prediction engine 164.
[0082] Furthermore, the prediction engine 164 may incorporate various types of data and models to make its predictions. This can include experimental data from the electronic notebook component 150, thermodynamic calculations, empirical correlations, or machine learning models trained on relevant chemical data. The prediction engine 164 may also integrate with external databases or resources to obtain additional information or data relevant to the chemical processes being analyzed.
[0083] The prediction engine 164 is not limited to the specific embodiments or functionalities described above. Depending on the specific implementation and requirements of the system, the prediction engine 164 may be adapted or extended to provide additional predictive capabilities or integrate with other components of the mass-flow estimator component 112 or the overall system 100. [0084] Model compound A is a halogenated aromatic with an additional nitrile group, whereas compound B is a halogenated heteroaromatic molecule. These are so-called value- added complex intermediates in the chemical industry, which means that they are produced from base chemicals via one or more synthesis steps. Such molecules usually carry more than one functional group (e.g. halogenic substituent, nitrile group, amine group, alcohol group, ester from boronic acid, etc.). In most cases, multiple synthesis pathways to produce these compounds are possible. The environmental footprint of such molecules can be hardly assessed because they are almost never found in environmental databases (e.g. ecoinvent). To estimate sustainability metrics such as the carbon footprint for model compound B, it is necessary to understand the synthesis pathway of the molecule and the respective amounts of chemicals, solvents, water, etc. consumed in each synthesis step.
[0085] Manual research of both, the synthesis pathway and resource consumption in each step, requires data collecting from many different literature sources and often takes several days. In comparison, using the system 100 of Fig. 1 in combination with the prediction engine 164 of Fig. 1 or 266 of Fig. 2 (e.g., retro-synthesis tools) can perform this assessment with reliable data from the electronic-lab notebook found in the electronic notebook component 150. To estimate the environmental footprint of model compound B from Fig. 14 (3,6- Dichloropyridazine) with ELN-data from the electronic notebook component 150, a retro- synthesis tool must provide the following information:
[0086] 1. Chemical synthesis route to produce 3,4-Dichloropyridazine
[0087] 2. Starting materials should be base chemicals, available in bulk quantities, that are available in a life cycle inventory database (e.g. ecoinvent)
[0088] 3. Reaction types should be mentioned to search for similar reactions in the system 100 to get the average amount of solvent and water used in the specific reaction type [0089] Synthia™ is a retro-synthesis tool that suggests production routes for a certain molecule and the above-mentioned information. An alternative to Synthia™ is the tool ASKCOS from MIT (Massachusetts Institute of Technology), which is available free of charge. [00107] For the model compound B, Synthia™ suggests a two-step synthesis from the base chemical maleic acid and hydrazine (Fig. 15). Fig. 14 shows model compounds for raw materials to synthesize active pharmacal ingredients while Fig. 15 shows how it may be be presented in the GUI. Fig. 15 shows a two-step synthesis route of model compound of Fig. 14 using a prediction engine in accordance with an embodiment of the present disclosure.
[0090] Both chemicals are available in life cycle inventory database such as ecoinvent. Additionally, the type of reaction is shown and basic information on reaction conditions (e.g. solvent, catalyst, temperature during reaction) are mentioned. The disclosed method does offer the automated generation of combined PMI / yield data of recorded chemical conversions and related processes, thereby preparing data sets, which can be used to allow a faster and more qualified plausibility testing of 'greener' retro-synthetic planning by systematic analysis of similar structures and chemical conversions.
[0091] This information can be used to search for similar reactions by the search component 158 to obtain the average amount of solvent, water, catalyst, etc. consumed in each step by looking at the average PMI for the specific reaction type that was suggested by the retro-synthesis tool. Furthermore, the prediction engine 164 can also provide an average chemical yield for the respective synthesis steps as shown in Fig. 16. Based on the comparison of PMI and chemical yield for dozens of similar reactions documented in the ELN-system, an average amount of solvent, water, catalyst, etc. can be applied.
[0092] In addition, an uncertainty estimate may be made in the environmental footprint evaluation. The deviation of the PMI data points to the average PMI as well as the deviation of data points for the chemical yield to the average yield can provide an indication on the accuracy on the environmental footprint evaluation.
[0093] Thus, the combination of a retro-synthesis tool such as Synthia™ or ASKCOS with the prediction engine 164 allows the estimation of sustainability metrics (e.g. solvent intensity, water intensity, carbon footprint) for the above mentioned model compounds in a three-step process:
[0094] 1. Starting materials and reaction types are provided by e.g. Synthia™.
[0095] 2. Average synthesis yield and average PMI (and derived from that: solvent intensity, water intensity, etc.) are obtained.
[0096] 3. Product carbon footprints can be estimated based on life cycle inventory database entries for raw materials, solvents, water, catalysts, etc.
[0097] Various alternatives and modifications can be devised by those skilled in the art without departing from the disclosure. Accordingly, the present disclosure is intended to embrace all such alternatives, modifications and variances. Additionally, while several embodiments of the present disclosure have been shown in the drawings and/or discussed herein, it is not intended that the disclosure be limited thereto, as it is intended that the disclosure be as broad in scope as the art will allow and that the specification be read likewise. Therefore, the above description should not be construed as limiting, but merely as exemplifications of particular embodiments. And, those skilled in the art will envision other modifications within the scope and spirit of the claims appended hereto. Other elements, steps, methods and techniques that are insubstantially different from those described above and/or in the appended claims are also intended to be within the scope of the disclosure.
[0098] Figs. 11 and 12 describe the function of the recycling functionality incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2. In this module the effect of solvent recycling on sustainability metrics is estimated. Fig. 17 describes the function of the energy module incorporated into the prediction engine 164 of Fig. 1 or the prediction engine 266 of Fig. 2. Fig. 17 may be a separate module used for energy calculation at each step using, e.g., input metrics.
[0099] In this module energy consumption of the selected processes or subprocesses can be estimated based on basic physical phenomena and/or empirical relationships. Individual unit-operations are selected within the tool and important parameters are adjusted by the user. Examples for unit operations are, but are not limited to, heating, refluxing, distillation, cooling, crystallization, drying, applying vacuum, filtration, chromatography, recovery, stirring, pumping, grinding, sublimation, inertisation, or extraction. A summary of the energy contribution the unit-operations of the unit operations in the synthesis tree are visualized in the software, like shown in Fig. 18. The software allows adaption to various scales. The energy contribution might be included into the estimation of the final product carbon footprint.
[0100] Fig. 19 describes the possibility to include weight-based carbon footprints for each individual staring material. These can be assigned manually, can be retrieved from external date vendors (e.g. Ecoinvent), or can be assigned based on the individual role of each starting material. The output presented in the GUI can also be a weight-based carbon footprint estimation.
[0101] The embodiments shown in the drawings are presented only to demonstrate certain examples of the disclosure. And, the drawings described are only illustrative and are non-limiting. In the drawings, for illustrative purposes, the size of some of the elements may be exaggerated and not drawn to a particular scale. Additionally, elements shown within the drawings that have the same numbers may be identical elements or may be similar elements, depending on the context.
[0102] Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun, e.g., "a," "an," or "the,” this includes a plural of that noun unless something otherwise is specifically stated. Hence, the term "comprising" should not be interpreted as being restricted to the items listed thereafter; it does not exclude other elements or steps, and so the scope of the expression "a device comprising items A and B" should not be limited to devices consisting only of components A and B. This expression signifies that, with respect to the present disclosure, the only relevant components of the device are A and B.
[0103] Furthermore, the terms "first," "second," "third," and the like, whether used in the description or in the claims, are provided for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances (unless clearly disclosed otherwise) and that the embodiments of the disclosure described herein are capable of operation in other sequences and/or arrangements than are described or illustrated herein.
[0104] Each of the characteristics and examples described herein, and combination thereof, may be said to be encompassed by the present disclosure. The present disclosure is thus drawn to, but not limited to, the following aspects:
[0105] (1) A computer-implemented method for calculating mass flow and estimating sustainability metrics, the method comprising: selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material; linking the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses wherein the integrated process synthesizes a target material; estimating a plurality of process metrics wherein each of the plurality of process metrics corresponds to a respective experimental IDs of the one or more experimental IDs; estimating an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics; providing a synthesis tree corresponding to the one or more experimental IDs; searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs; selecting an alternative experimental ID for at least one of the one or more experimental IDs; and updating the aggregate process metric in accordance with the alternative experimental ID.
[0106] (2) The method according to aspect 1, further comprising displaying the process metric.
[0107] (3) The method according to aspect 1, wherein each experimental ID includes experimental data corresponding to one or more previously performed experiments. [0108] (4) The method according to aspect 1, further comprising comparing the aggregate process metric with the alternative experimental ID to the aggregate process metric prior to updating the aggregate process metric in accordance with the alternative experimental ID. [0109] (5) The method according to aspect 4, further comprising displaying the comparison.
[0110] (6) The method according to aspect 1, wherein the aggregate process metric is a process mass intensity.
[0111] (7) The method according to aspect 1 , wherein the aggregate process metric is one of a solvent intensity, a water intensity, and a global warming potential.
[0112] (8) The method according to aspect 1, wherein the aggregate process metric is one of a cost, a yield, and an environmental metric.
[0113] (9) The method according to aspect 1, further comprising selecting a role for a molecule associate with an experimental ID of the one or more experimental IDs.
[0114] (10) The method according to aspect 1, further comprising adjusting a process parameter.
[0115] (11) The method according to aspect 10, further comprising updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
[0116] (12) The method according to aspect 1, wherein the act of selecting is performed by a user selecting the one or more experimental IDs utilizing a Graphical User Interface.
[0117] (13) The method according to aspect 1, further comprising: querying a prediction engine to determine a prophetic experiment; replacing at least one of the one or more experimental IDs with the prophetic experiment; estimating a prophetic process metric corresponding to the prophetic experiment; and updating the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
[0118] (14) The method according to aspect 13, wherein the prophetic process metric includes a range of values.
[0119] (15) The method according to aspect 14, wherein the range of values is one of a confidence interval, a confidence region, a credible interval, and a credible region.
[0120] (16) The method according to aspect 1, wherein the plurality of process metrics is a ratio.
[0121] (17) The method according to aspect 16, wherein the ratio is PMI over Yield.
[0122] (18) The method according to aspect 1, wherein the one or more experimental
IDs are selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material. [0123] (19) The method according to aspect 1, wherein the subprocesses include at least one of chemical reactions or purifications.
[0124] (20) The method according to aspect 1, further comprising generating a report summarizing the estimated process metrics and the aggregate process metric.
[0125] (21) The method according to aspect 1, wherein the linkage of the one or more experimental IDs includes verifying compatibility of input and output materials to thereby ensure continuity in the integrated process.
[0126] (22) The method according to aspect 1, wherein the process metrics include energy consumption metrics calculated based on the input and output materials and the subprocesses utilized.
[0127] (23) The method according to aspect 1, further comprising a step of validating the estimated process metrics against predetermined criteria before updating the aggregate process metric.
[0128] (24) The method according to aspect 1, wherein the searching for similar reaction subprocesses includes utilizing a machine learning model for the act of searching.
[0129] (25) The method according to aspect 1, further comprising the step of alerting a user when the aggregate process metric exceeds a predetermined environmental impact threshold.
[0130] (26) The method according to aspect 1, wherein each experimental ID is associated with specific equipment used in the subprocess, and the method further comprises adjusting equipment settings based on the subprocess requirements.
[0131] (27) The method according to aspect 1, further comprising a step of automatically ordering chemicals and materials needed for the subprocesses based on the input materials listed in the experimental IDs.
[0132] (28) The method according to aspect 1, wherein the synthesis tree provided includes alternative synthesis pathways, and the method further comprises selecting one of the alternative pathways based on user-defined criteria.
[0133] (29) The method of aspect 1, further comprising estimating an energy consumption metric for the integrated process based on the one or more experimental IDs.
[0134] (30) The method of aspect 29, wherein estimating the energy consumption metric comprises applying at least one of empirical correlations or thermodynamic calculations to process parameters associated with the one or more experimental IDs. [0135] (31) The method of aspect 1, further comprising estimating a carbon footprint metric for the integrated process based on the one or more experimental IDs and life cycle inventory data for input materials.
[0136] (32) The method of aspect 1, further comprising providing a visualization of mass flow through the integrated process based on the one or more experimental IDs.
[0137] (33) The method of aspect 32, wherein the visualization comprises a Sankey diagram.
[0138] (34) The method of aspect 1, further comprising enabling a user to edit the synthesis tree corresponding to the one or more experimental IDs via a graphical user interface. [0139] (35) The method of aspect 1, wherein searching for similar reaction subprocesses comprises identifying subprocesses having the same input and output molecular structures as a given experimental ID.
[0140] (36) The method of aspect 1, further comprising estimating a range of values for at least one of the plurality of process metrics based on the plurality of similar experiments. [0141] (37) The method of aspect 1, further comprising suggesting an alternative solvent or an alternative reagent for at least one of the experimental IDs based on the plurality of similar experiments.
[0142] (38) The method of aspect 1, wherein the one or more experimental IDs are selected from the electronic notebook based on a target molecule specified by a user.
[0143] (39) The method of aspect 1, further comprising tracking modifications made to the integrated process and impacts on the aggregate process metric over time.
[0144] (40) The method of aspect 1, further comprising estimating the impacts of one of solvent recycling and waste stream recycling on the plurality of process metrics.
[0145] (41) The method according to aspect 1, further comprising integrating the process metrics and aggregate process metric with an enterprise resource planning (ERP) system for comprehensive data management.
[0146] (42) The method of aspect 1, further comprising storing the one or more experimental IDs, the plurality of process metrics, and the aggregate process metric in a database.
[0147] (43) The method of aspect 42, further comprising retrieving the stored one or more experimental IDs, plurality of process metrics, and aggregate process metric from the database.
[0148] (44) The method of aspect 1, further comprising enabling a user to manually edit input data associated with at least one of the one or more experimental IDs. [0149] (45) The method of aspect 44, further comprising updating the plurality of process metrics and the aggregate process metric based on the edited input data.
[0150] (46) The method of aspect 45, further comprising displaying changes between the edited input data and original input data associated with the at least one experimental ID.
[0151] (47) The method of aspect 46, wherein displaying the changes comprises visually highlighting the changes or providing a side-by-side comparison of the original and edited input data.
[0152] (48) A method according to aspect 1, wherein the process metrics are provided for the whole process, parts of the process, or each of the subprocess steps.
[0153] (49) A method according to aspect 1 or 2, wherein the process metrics are compared to one or more process metrics from other processes for easier comparison.
[0154] (50) The method of aspect 1, further comprising providing the plurality of process metrics for each individual subprocess of the integrated process.
[0155] (51) The method of aspect 1, further comprising comparing at least one process metric of the plurality of process metrics to a corresponding process metric from a different integrated process.
[0156] (52) The method of aspect 1, further comprising visualizing at least one process metric of the plurality of process metrics using a Sankey diagram.
[0157] (53) The method of aspect 1, wherein the searching for similar reaction subprocesses includes identifying recycled mass streams that are recycled to the beginning of a subprocess or the integrated process.
[0158] (54) The method of aspect 11, wherein the range of values for the prophetic process metric corresponds to a multi-step process.
[0159] (55) The method of aspect 1, wherein the linking of the one or more experimental IDs to form the integrated process is performed automatically based on the input and output chemical structures of the experimental IDs.
[0160] (56) The method of aspect 55, wherein the ratio is process mass intensity over yield.
[0161] (57) A data processing system comprising means for carrying out the method of any one of aspects 1 to 56.
[0162] (58) A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56. [0163] (59) A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of aspects 1 to 56.

Claims

What is Claimed is:
1. A computer-implemented method for calculating mass flow and estimating sustainability metrics, the method comprising: selecting one or more experimental IDs from an electronic notebook, wherein each experimental ID corresponds to a subprocess having an input material and an output material; linking the one or more experimental IDs to form an integrated process including all of the corresponding subprocesses wherein the integrated process synthesizes a target material; estimating a plurality of process metrics wherein each of the plurality of process metrics corresponds to a respective experimental IDs of the one or more experimental IDs; estimating an aggregate process metric corresponding to the integrated process, wherein the aggregate process metric is a function of the estimated plurality of process metrics; providing a synthesis tree corresponding to the one or more experimental IDs; searching for similar reaction subprocesses for each of the experimental IDs to determine a plurality of similar experiments, each similar experiment corresponding to one or more respective experiments of the one or more experimental IDs; selecting an alternative experimental ID for at least one of the one or more experimental IDs; and updating the aggregate process metric in accordance with the alternative experimental ID.
2. The method according to claim 1, further comprising displaying the process metric.
3. The method according to claim 1 , wherein each experimental ID includes experimental data corresponding to one or more previously performed experiments.
4. The method according to claim 1, further comprising comparing the aggregate process metric with the alternative experimental ID to the aggregate process metric prior to updating the aggregate process metric in accordance with the alternative experimental ID.
5. The method according to claim 4, further comprising displaying the comparison.
52
RECTIFIED SHEET (RULE 91) ISA/EP
6. The method according to claim 1 , wherein the aggregate process metric is a process mass intensity.
7. The method according to claim 1, wherein the aggregate process metric is one of a solvent intensity, a water intensity, and a global warming potential.
8. The method according to claim 1, wherein the aggregate process metric is one of a cost, a yield, and an environmental metric.
9. The method according to claim 1 , further comprising selecting a role for a molecule associate with an experimental ID of the one or more experimental IDs.
10. The method according to claim 1, further comprising adjusting a process parameter.
11. The method according to claim 10, further comprising updating at least one of the plurality of process metrics in accordance with the adjusted process parameter.
12. The method according to claim 1, wherein the act of selecting is performed by a user selecting the one or more experimental IDs utilizing a Graphical User Interface.
13. The method according to claim 1, further comprising: querying a prediction engine to determine a prophetic experiment; replacing at least one of the one or more experimental IDs with the prophetic experiment; estimating a prophetic process metric corresponding to the prophetic experiment; and updating the aggregate process metric to include the prophetic experiment in place of the at least one of the one or more experimental IDs.
14. The method according to claim 13, wherein the prophetic process metric includes a range of values.
53
RECTIFIED SHEET (RULE 91) ISA/EP
15. The method according to claim 14, wherein the range of values is one of a confidence interval, a confidence region, a credible interval, and a credible region.
16. The method according to claim 1 , wherein the plurality of process metrics is a ratio.
17. The method according to claim 16, wherein the ratio is PMI over Yield.
18. The method according to claim 1 , wherein the one or more experimental IDs are selected out of order and the act of linking the plurality of experiment IDs includes the act of ordering the one or more experimental IDs to form the integrated process including all of the corresponding subprocesses in order configured to thereby synthesize the target material.
19. The method according to claim 1, wherein the subprocesses include at least one of chemical reactions or purifications.
20. The method according to claim 1 , further comprising generating a report summarizing the estimated process metrics and the aggregate process metric.
21. The method according to claim 1 , wherein the linkage of the one or more experimental IDs includes verifying compatibility of input and output materials to thereby ensure continuity in the integrated process.
22. The method according to claim 1, wherein the process metrics include energy consumption metrics calculated based on the input and output materials and the subprocesses utilized.
23. The method according to claim 1 , further comprising a step of validating the estimated process metrics against predetermined criteria before updating the aggregate process metric.
24. A data processing system comprising means for carrying out the method of claim 1.
54
RECTIFIED SHEET (RULE 91) ISA/EP
25. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of claim 1.
26. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of claim 1.
55
RECTIFIED SHEET (RULE 91) ISA/EP
PCT/EP2024/066006 2023-06-12 2024-06-11 Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis Ceased WO2024256360A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP24732622.6A EP4725021A1 (en) 2023-06-12 2024-06-11 Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23178759.9 2023-06-12
EP23178759 2023-06-12

Publications (1)

Publication Number Publication Date
WO2024256360A1 true WO2024256360A1 (en) 2024-12-19

Family

ID=86760182

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/EP2024/066006 Ceased WO2024256360A1 (en) 2023-06-12 2024-06-11 Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis

Country Status (3)

Country Link
EP (1) EP4725021A1 (en)
TW (1) TW202520282A (en)
WO (1) WO2024256360A1 (en)

Non-Patent Citations (8)

* Cited by examiner, † Cited by third party
Title
GREEN CHEM., vol. 17, no. 3111, 2015
GSK STUDIE: ORG. PROCESS RES. DEV., vol. 15, 2011, pages 912 - 917
JIMENEZ-GONZALES ET AL., ORG. PROCESS RES. DEV, vol. 15, no. 4, 2011, pages 912 - 917
LOUREIRO HUGO ET AL: "ChemPager: Now Expanded for Even Greener Chemistry", CHIMIA INTERNATIONAL JOURNAL FOR CHEMISTRY, vol. 73, no. 9, 18 September 2019 (2019-09-18), CH, pages 724, XP093208077, ISSN: 0009-4293, Retrieved from the Internet <URL:https://chimia.ch/chimia/article/download/2019_724/621> DOI: 10.2533/chimia.2019.724 *
PARVATKER ET AL., ACS SUSTAINABLE CHEM. ENG., vol. 7, 2019, pages 6580 - 6591
PROCESSES, vol. 10, 2022, pages 1274
SHARMA PANKAJ ET AL: "DOZNTM 2.0: A quantitative green chemistry evaluator for a sustainable future", JOURNAL OF ORGANOMETALLIC CHEMISTRY, ELSEVIER, AMSTERDAM, NL, vol. 970, 29 April 2022 (2022-04-29), XP087076775, ISSN: 0022-328X, [retrieved on 20220429], DOI: 10.1016/J.JORGANCHEM.2022.122367 *
SHERER EDWARD C. ET AL: "Driving Aspirational Process Mass Intensity Using Simple Structure-Based Prediction", ORGANIC PROCESS RESEARCH & DEVELOPMENT, vol. 26, no. 5, 18 April 2022 (2022-04-18), US, pages 1405 - 1410, XP093208123, ISSN: 1083-6160, Retrieved from the Internet <URL:https://pubs.acs.org/doi/pdf/10.1021/acs.oprd.1c00477> DOI: 10.1021/acs.oprd.1c00477 *

Also Published As

Publication number Publication date
EP4725021A1 (en) 2026-04-15
TW202520282A (en) 2025-05-16

Similar Documents

Publication Publication Date Title
US10636007B2 (en) Method and system for data-based optimization of performance indicators in process and manufacturing industries
US11120347B2 (en) Optimizing data-to-learning-to-action
US20250094841A1 (en) Hybrid Machine Learning
US20240420026A1 (en) Systems and methods for advanced prediction using machine-learning and statistical models
Pu et al. The analysis of strategic management decisions and corporate competitiveness based on artificial intelligence
Machireddy et al. Enhancing predictive analytics with AI-powered RPA in cloud data warehousing: A comparative study of traditional and modern approaches
Sakhrawi et al. Support vector regression for enhancement effort prediction of Scrum projects from COSMIC functional size
Gangadharan et al. Metaheuristic approaches in biopharmaceutical process development data analysis
Mostofi et al. Performance‐driven contractor recommendation system using a weighted activity–contractor network
Hou et al. A novel technology life cycle analysis method based on LSTM and CRF
JP2008171171A (en) Demand forecast method, demand forecast analysis server, and demand forecast program
WO2024256360A1 (en) Method and system for calculating mass flow and estimating sustainability metrics in chemical synthesis
Walton et al. Automated resonance fitting for nuclear data evaluation
US20220067628A1 (en) Directional stream value analysis system and server
CN120543180A (en) An automated management system for watch after-sales service
Costantini et al. On the use of mean square error and directional forecast accuracy for model selection: a simulation study
Mhaskey Unlocking business potential: The transformative impact of ERP analytics
Sokolovas Investigation of process automation with large language models
JP4738898B2 (en) Demand forecast method, demand forecast analysis server, and demand forecast program
Fekete et al. A comprehensive causal AI framework for analysing factors affecting energy consumption and costs in customised manufacturing
US20250013632A1 (en) Smart selection of data fields during data analysis
Islam et al. PREDICTIVE ANALYTICS IN SUPPLY CHAIN MANAGEMENT A REVIEW OF BUSINESS ANALYST-LED OPTIMIZATION TOOLS
Chain Review of Applied Science and Technology
JORDAN AI-POWERED PORTFOLIO MANAGEMENT IN PHARMACEUTICAL R&D5. AI-POWERED PORTFOLIO MANAGEMENT IN PHARMACEUTICAL R&D
Lingqa et al. User Interface Design of Safety Stock Prediction System Using Demand Response-ARMA Method Case Study: Jaya Lestari

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24732622

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2024732622

Country of ref document: EP

Effective date: 20260112

WWE Wipo information: entry into national phase

Ref document number: 2024732622

Country of ref document: EP

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2024732622

Country of ref document: EP

Effective date: 20260112

ENP Entry into the national phase

Ref document number: 2024732622

Country of ref document: EP

Effective date: 20260112

ENP Entry into the national phase

Ref document number: 2024732622

Country of ref document: EP

Effective date: 20260112

ENP Entry into the national phase

Ref document number: 2024732622

Country of ref document: EP

Effective date: 20260112

WWP Wipo information: published in national office

Ref document number: 2024732622

Country of ref document: EP