WO2024012017A1 - 反应物分子的预测、模型的训练方法、装置、设备及介质 - Google Patents

反应物分子的预测、模型的训练方法、装置、设备及介质 Download PDF

Info

Publication number
WO2024012017A1
WO2024012017A1 PCT/CN2023/092036 CN2023092036W WO2024012017A1 WO 2024012017 A1 WO2024012017 A1 WO 2024012017A1 CN 2023092036 W CN2023092036 W CN 2023092036W WO 2024012017 A1 WO2024012017 A1 WO 2024012017A1
Authority
WO
WIPO (PCT)
Prior art keywords
molecule
information
chemical bond
sample
completed
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/092036
Other languages
English (en)
French (fr)
Inventor
孟子乔
赵沛霖
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology Shenzhen Co Ltd
Original Assignee
Tencent Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology Shenzhen Co Ltd filed Critical Tencent Technology Shenzhen Co Ltd
Publication of WO2024012017A1 publication Critical patent/WO2024012017A1/zh
Priority to US18/597,636 priority Critical patent/US20240212796A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/50Molecular design, e.g. of drugs
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/20Identification of molecular entities, parts thereof or of chemical compositions
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/30Prediction of properties of chemical compounds, compositions or mixtures
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/70Machine learning, data mining or chemometrics

Definitions

  • the embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a prediction of reactant molecules, a model training method, device, equipment and medium.
  • the embodiments of the present application provide a method, device, equipment and medium for predicting reactant molecules and training models, which can be used to improve the reliability and accuracy of predicting reactant molecules.
  • the technical solutions are as follows:
  • embodiments of the present application provide a method for predicting reactant molecules, which method includes:
  • the computer equipment acquires product molecules, breaks bonds on the product molecules, and obtains molecules to be completed, where the product molecules refer to any compound molecule of the reactant molecules to be predicted;
  • the computer device calls a molecular completion model to complete the molecule to be completed, obtains a completion result, and determines the reactant molecules of the product molecule based on the completion result;
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking substructures in the sample compound molecules.
  • a method for training a molecular completion model includes:
  • the computer device acquires sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking substructures in the sample compound molecules;
  • the computer device determines a training loss based on the sample compound molecules, the sample molecules to be completed, and the molecule completion model;
  • the computer device updates the model parameters of the molecular completion model based on the training loss to obtain a trained molecular completion model.
  • a prediction device for reactant molecules is provided, which is provided in a computer device.
  • the device includes:
  • the first acquisition unit is used to acquire product molecules, break bonds on the product molecules, and obtain molecules to be completed, where the product molecules refer to any compound molecule of the reactant molecules to be predicted;
  • a completion unit is used to call a molecular completion model to complete the molecule to be completed, obtain a completion result, and obtain the reactant molecules of the product molecule based on the completion result;
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking substructures in the sample compound molecules.
  • a training device for a molecular completion model is also provided, which is provided in a computer device.
  • the device includes:
  • the second acquisition unit is used to acquire sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking substructures in the sample compound molecules;
  • the third acquisition unit is used to based on the sample compound molecules, the sample molecules to be completed and the molecule completion model, Get the training loss;
  • An update unit configured to update the model parameters of the molecular completion model based on the training loss to obtain a trained molecular completion model.
  • a computer device includes a processor and a memory, at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor, so that the The computer device implements any of the above-mentioned methods for predicting reactant molecules or training methods for molecular completion models.
  • a computer-readable storage medium is also provided. At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by the processor, so that the computer device implements any of the above. 1. The prediction method of reactant molecules or the training method of molecular completion model.
  • the computer program product includes a computer program or computer instructions.
  • the computer program or the computer instructions are loaded and executed by a processor, so that the computer device implements any of the above.
  • the above-mentioned prediction method of reactant molecules or the training method of molecular completion model is also provided.
  • the prediction process of reactant molecules relies on a molecular completion model.
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed. Since the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules, that is to say, the training process of the molecule completion model is based on data obtained based on the sample compound molecules themselves.
  • the training process is a self-supervised training process based on sample compound molecules. This self-supervised training process does not need to pay attention to whether the sample compound molecules are compounds in known synthesis reactions. Therefore, this self-supervised training process is not affected by known compounds. Due to the limitations of the synthetic reaction, the molecular completion model trained using this training process has strong generalization ability, which is conducive to expanding the adaptation scenarios, thereby helping to improve the prediction reliability and accuracy of the reactant molecules.
  • Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application.
  • Figure 2 is a flow chart of a method for predicting reactant molecules provided by the embodiment of the present application.
  • Figure 3 is a schematic representation of a product molecule provided by the embodiment of the present application.
  • Figure 4 is a schematic diagram of two stages of a reactant molecule prediction process provided by the embodiment of the present application.
  • Figure 5 is a flow chart of a training method for a molecular completion model provided by an embodiment of the present application
  • Figure 6 is a schematic diagram of three situations of masked substructures in a sample compound molecule provided by the embodiment of the present application.
  • Figure 7 is a schematic diagram of cleaving different chemical bonds in sample compound molecules provided by the embodiment of the present application.
  • Figure 8 is a schematic structural diagram of an initial chemical bond completion model provided by an embodiment of the present application.
  • Figure 9 is a schematic structural diagram of an initial atomic completion model provided by an embodiment of the present application.
  • Figure 10 is a schematic diagram of a predicting device for reactant molecules provided by an embodiment of the present application.
  • Figure 11 is a schematic diagram of a training device for a molecular completion model provided by an embodiment of the present application.
  • Figure 12 is a schematic structural diagram of a server provided by an embodiment of the present application.
  • Figure 13 is a schematic structural diagram of a terminal provided by an embodiment of the present application.
  • related technology predicts reactant molecules, it first obtains the molecule to be completed based on the product molecule, then determines the structure matching the molecule to be completed from multiple candidate structures, connects the matching structure to the molecule to be completed, and connects the The resulting molecules are used as predicted reactant molecules.
  • multiple candidate structures are obtained by comparing the differential structures between product molecules and reactant molecules in known synthesis reactions.
  • the prediction method of the above reactant molecules relies on multiple candidate structures extracted from known synthesis reactions.
  • the generalization ability of multiple candidate structures is limited by the known synthesis reactions.
  • the generalization ability is poor and the applicable scenarios are relatively limited. , thus easily reducing the prediction reliability and prediction accuracy of reactant molecules.
  • the prediction method of reactant molecules and the training method of molecule completion models provided by the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. .
  • Artificial Intelligence is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
  • artificial intelligence is a comprehensive technology of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence.
  • Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
  • Artificial intelligence technology is a comprehensive subject that covers a wide range of fields, including both hardware-level technology and software-level technology.
  • Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation/interaction systems, mechatronics and other technologies.
  • Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, machine learning/deep learning, autonomous driving, smart transportation and other major directions.
  • Machine Learning is a multi-field interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. and many other disciplines. It specializes in studying how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance.
  • Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence.
  • Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, teaching learning and other technologies.
  • artificial intelligence technology has been researched and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless driving, autonomous driving, and drones. , robots, smart medical care, smart customer service, Internet of Vehicles, autonomous driving, smart transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
  • Figure 1 shows a schematic diagram of an implementation environment provided by an embodiment of the present application.
  • the implementation environment may include: a terminal 11 and a server 12 .
  • the method for predicting reactant molecules provided by the embodiments of the present application can be executed by the terminal 11 , the server 12 , or both the terminal 11 and the server 12 . This is not limited by the embodiments of the present application.
  • the server 12 undertakes the main calculation work and the terminal 11 undertakes the secondary calculation work; alternatively, the server 12 undertakes the secondary calculation work and the terminal 11 undertakes the main computing work; alternatively, the server 12 and the terminal 11 adopt a distributed computing architecture for collaborative computing.
  • the training method of the molecular completion model provided by the embodiment of the present application can be executed by the terminal 11 , the server 12 , or both the terminal 11 and the server 12 . This is not limited by the embodiment of the present application.
  • the server 12 undertakes the main calculation work, and the terminal 11 undertakes the secondary calculation work; or the server 12 undertakes the secondary calculation work
  • the terminal 11 is responsible for the main computing work; alternatively, the server 12 and the terminal 11 adopt a distributed computing architecture for collaborative computing.
  • execution device of the prediction method of reactant molecules and the execution device of the training method of the molecule completion model may be the same or different, and this is not limited in the embodiments of the present application.
  • the terminal 11 can be any electronic product that can perform human-computer interaction with the user through one or more methods such as keyboard, touch pad, touch screen, remote control, voice interaction or handwriting device, such as PC (Personal Computer). , personal computer), mobile phone, smartphone, PDA (Personal Digital Assistant, personal digital assistant), wearable devices, PPC (Pocket PC, handheld computer), tablet computer, smart car, smart TV, smart speaker, smart voice interaction Equipment, smart home appliances, vehicle-mounted terminals, VR (Virtual Reality, virtual reality) equipment, AR (Augmented Reality, Augmented reality) equipment, etc.
  • the server 12 may be one server, a server cluster composed of multiple servers, or a cloud computing service center. The terminal 11 and the server 12 establish a communication connection through a wired or wireless network.
  • terminal 11 and server 12 are only examples. If other existing or possible terminals or servers that may appear in the future are applicable to this application, they should also be included in the protection scope of this application, and are hereby referred to as References are included here.
  • the prediction method of reactant molecules provided in the embodiments of the present application is used to predict the reactant molecules that generate the product molecule based on the given product molecule.
  • This prediction task can be called a retrosynthesis prediction task.
  • the retrosynthesis prediction task is very important in the field of chemistry. and the pharmaceutical field are of extremely important significance.
  • Traditional retrosynthesis prediction tasks are mostly implemented based on synthesis reaction templates. For example, a matching algorithm is first used to find a template that matches the product molecule from the synthesis reaction template, and then the reactant molecules are obtained based on the matched template.
  • This type of method has achieved certain results in retrosynthetic prediction tasks, but methods based on synthetic reaction templates have two obvious flaws: The first flaw is that methods based on synthetic reaction templates are difficult to generalize to new reaction types.
  • the synthesis reaction template needs to be updated frequently, and summarizing the template requires a lot of work from chemical experts, which is very costly.
  • the second flaw is that the synthesis reaction template only summarizes part of the reaction rules at the molecular level and cannot capture the correct information of the whole situation. , often leading to wrong predictions.
  • Deep learning technology can be used to directly predict the reactant molecules of product molecules without the need for synthetic reactions. template to match.
  • Deep learning technology can achieve powerful retrosynthetic prediction effects.
  • the powerful retrosynthetic prediction effect can help chemical experts discover possible synthesis pathways of product molecules, greatly improving the efficiency of research and development of new compounds.
  • the product molecules can be drug molecules. , it can greatly improve the research and development efficiency of new drugs in the pharmaceutical industry.
  • powerful retrosynthetic prediction effects can also reveal some hidden scientific laws, provide new scientific knowledge, and discover new synthetic pathways and even new synthetic reactions.
  • the prediction method of reactant molecules provided in the embodiments of this application is a method for realizing the retrosynthesis prediction task based on deep learning technology.
  • V P represents a set of atoms in the product molecule G P and the size of the set represents the number N of atoms in the product molecule G P (N is an integer not less than 1), that is,
  • N.
  • B P ⁇ R N ⁇ N ⁇ C represents the chemical bond connection information of the product molecule G P.
  • the chemical bond connection information represents the chemical bond connection between the atoms in the product molecule G P.
  • the chemical bond connection information is a three-dimensional matrix, C( C (an integer not less than 1) represents the number of types of chemical bonds that may exist between atoms.
  • the value of the element [i, j, c] in B P represents the relationship between atom i and atom j in the product molecule G P Whether it is connected through a chemical bond of type c.
  • the chemical bond connection information may also be called an adjacency matrix.
  • a P ⁇ R N ⁇ F represents the atomic characteristic information of the product molecule G P.
  • the atomic characteristic information includes the sub-characteristic information of each atom in the product molecule G P.
  • Each atom has a dimension of F (F is not less than 1
  • the sub-feature information of the integer) the method of obtaining the sub-feature information of the atom will be introduced below, and will not be described here.
  • the meanings of V R , B R , and A R can be found in the meanings of V P , B P , and A P , and will not be described in detail here.
  • the embodiments of the present application provide a method for predicting reactant molecules, which can be applied to the implementation environment shown in Figure 1 above.
  • the prediction method of reactant molecules is executed by a computer device, which may be a terminal 11 or a server. server 12, which is not limited in the embodiment of the present application.
  • the prediction method of reactant molecules provided by the embodiment of the present application may include the following steps 201 to 203.
  • step 201 the computer device acquires product molecules.
  • a product molecule refers to any compound molecule whose reactant molecules are to be predicted. By predicting the reactant molecules of a product molecule, the synthetic path of the product molecule can be deduced, thereby providing data support for the research and development of product molecules.
  • the embodiments of the present application do not limit the type of product molecules.
  • the product molecules may be drug molecules, clothing molecules, food molecules, etc.
  • the product molecule includes multiple atoms, and the multiple atoms are connected through chemical bonds.
  • the type and number of atoms included in the product molecule, and the chemical bond connection between the multiple atoms are related to the product molecule, and are not limited by the embodiments of this application.
  • product molecules may also be referred to as product molecules.
  • product molecules can be extracted from a compound molecule database
  • product molecules can also be selected from compound molecules disclosed in journals or articles, or technical personnel can upload compound molecules as product molecules, etc.
  • the embodiments of the present application do not limit the representation form of the product molecule, as long as it can indicate the status of the atoms in the product molecule and the chemical bond connection status between the atoms.
  • the representation form of the product molecule can be a name, a molecular formula, a string, etc.
  • the name of the product molecule can be obtained by naming the product molecule according to the compound naming rules.
  • the molecular formula of the product molecule is the most intuitive expression of the composition structure of the product molecule.
  • the product molecule can be determined intuitively according to the molecular formula of the product molecule.
  • the string of the product molecule is a string generated according to certain specifications, which can represent the product molecule more concisely.
  • the specification based on which the string of the product molecule is generated can refer to SMILES (Simplified Molecular Input Line entry Specification , simplified molecular linear input specification).
  • SMILES Simple Molecular Input Line entry Specification , simplified molecular linear input specification
  • the string of product molecules may also be called the SMILES expression of the product molecule.
  • SMILES is a specification that uses ASCII (American Standard Code for Information Interchange, American Standard Code for Information Interchange) strings to clearly describe the molecular structure. Each compound molecule has a unique SMILES expression.
  • the same product molecule can be expressed by the molecular formula shown in (1) in Figure 3 or by the SMILES expression shown in (2) in Figure 3 .
  • step 202 the computer equipment performs bond breaking on the product molecule to obtain the molecule to be completed.
  • the molecule to be completed is a molecule obtained by subjecting the product molecule to bond breaking.
  • the molecule to be completed may also be called a synthon.
  • the molecule to be completed can be regarded as the molecule obtained by removing some molecular structures from the reactant molecules in the process of synthesizing the product molecule. After identifying at least one molecule to be completed, it can be further predicted by completing the molecule to be completed.
  • reactant molecules For example, the molecular structure removed from the reactant molecule may be called a leaving group.
  • the number of molecules to be completed based on the product molecules may be one or multiple, which is related to the actual situation of the product molecules, and is not limited in the embodiments of the present application.
  • the computer device product molecule undergoes bond breaking to obtain the molecule to be completed, including the following steps 2021 to 2023.
  • Step 2021 The computer device obtains the graph structure information of the product molecule.
  • the graph structure information of the product molecule is information used to characterize the graph structure of the product molecule.
  • the graph structure of the product molecule may be the only certain graph structure obtained by converting the product molecule. For example, each atom in the product molecule is regarded as a node, and each chemical bond in the product molecule is regarded as an edge. From this perspective, the product molecule is converted into a graph structure.
  • the graph structure of the product molecule is a graph structure constructed with the atoms in the product molecule as nodes and the chemical bonds in the product molecule as edges.
  • the embodiments of the present application do not limit the type of graph structure information of the product molecule, as long as it can characterize the graph structure of the product molecule.
  • the graph structure information of the product molecule includes atomic characteristic information of the product molecule and chemical bond connection information of the product molecule.
  • the atomic characteristic information of the product molecule is used to characterize the characteristics of the atoms in the product molecule.
  • the chemical bond connection information of the product molecule is used to characterize the chemical bond connection between atoms in the product molecule.
  • the atomic feature information of the product molecule can be regarded as the information characterizing the nodes in the graph structure of the product molecule.
  • the chemical bonds of the product molecule The connection information can be regarded as information that characterizes the edges in the graph structure of the product molecule. Therefore, the atomic feature information of the product molecule and the chemical bond connection information of the product molecule can be used to characterize the graph structure of the product molecule.
  • the atomic characteristic information of the product molecule includes the sub-characteristic information of each atom in the product molecule.
  • the sub-characteristic information of the atom is used to characterize the characteristics of the atom.
  • the embodiment of the present application does not limit the expression method of the sub-characteristic information of the atom.
  • the representation form of the sub-feature information of an atom can be a matrix or a vector, etc.
  • the dimensions of the sub-feature information of different atoms are the same. The same dimensions can be set based on experience or can be flexibly adjusted according to the application scenario. This is not limited in the embodiments of the present application.
  • the atomic characteristic information of the product molecule can also be called the atomic characteristic matrix of the product molecule, and each row of elements in the atomic characteristic matrix represents the sub-characteristics of an atom.
  • Feature information when the representation form of the sub-characteristic information of an atom is a matrix, the atomic characteristic information of the product molecule can also be called the atomic characteristic matrix of the product molecule, and each row of elements in the atomic characteristic matrix represents the sub-characteristics of an atom.
  • the principle of obtaining the sub-feature information of each atom is the same.
  • the embodiment of this application takes the method of obtaining the sub-feature information of an atom as an example to illustrate.
  • the method of obtaining the sub-feature information of the atom includes: obtaining the attribute information of the atom; performing feature extraction on the attribute information of the atom to obtain the sub-feature information of the atom.
  • the attribute information of the atom is used to describe the attributes of the atom.
  • the attribute information of the atom is set based on experience or flexibly adjusted according to the application scenario.
  • the attribute information of the atom includes but is not limited to the element information, valence state information, degree information, At least one of the information on whether it belongs to a benzene ring.
  • the element information includes at least one of, but is not limited to, the ranking of the atom in the periodic table, the symbolic representation of the element, and the relative atomic mass. For example, carbon is ranked 6th in the periodic table of elements, its symbol is C, and its relative atomic mass is 12.01.
  • Valence information refers to the valence state of atoms in product molecules. Valence is also called chemical valence or atomic valence. Valence is the number of combinations of an atom or atomic group, base (root) of various elements with other atoms. Atoms may or may not have the same valence state in different compounds. For example, the valence state of carbon in CO (carbon monoxide) is +2, while the valence state of carbon in CO 2 (carbon dioxide) is +4.
  • Degree information includes the number of other atoms to which the atom is connected. For example, in CO 2 , a carbon atom is connected to two oxygen atoms, and both oxygen atoms are connected to a carbon atom. Then the degree information of the carbon atom can be 2.
  • the information about whether the atom belongs to the benzene ring indicates whether the atom is an atom constituting the benzene ring.
  • feature extraction is performed on the attribute information of the atom, and the extracted information is used as the sub-feature information of the atom.
  • the method of feature extraction from the attribute information of the atom can be set based on experience.
  • an atomic feature extraction model is called to extract features from the attribute information of the atom.
  • the atomic feature extraction model can be trained through supervised training based on the attribute information of the sample atoms and the feature labels of the sample atoms.
  • the chemical bond connection information of the product molecule is determined based on the chemical bond connection between atoms in the product molecule.
  • the chemical bond connection information of the product molecule may also be referred to as the adjacency matrix of the graph structure of the product molecule.
  • the chemical bond connection information of the product molecule is an N*N*C dimensional matrix, where N and C are both integers not less than 1, N represents the number of atoms in the product molecule, and C represents the type of candidate chemical bond. quantity.
  • candidate chemical bond types are set based on experience or flexibly adjusted according to application scenarios. This is not limited by the embodiments of the present application.
  • candidate chemical bond types cover common chemical bond types in synthesis reactions.
  • candidate chemical bond types include but are not limited to single Bonds, double bonds, triple bonds, aromatic bonds, ionic bonds, covalent bonds and metallic bonds, etc.
  • the value of the element [i, j, c] located at the i-th row, j-th column, a-th depth indicates whether atom i and atom j are connected by a chemical bond of type a.
  • the chemical bond refers to the a-th candidate chemical bond type among the C candidate chemical bond types. Among them, i and j are any value from 1 to N, and a is any value from 1 to C.
  • the chemical bond connection information of product molecules can be obtained by analyzing the chemical bond connection between atoms in the product molecule.
  • the chemical bond connection refers to whether the atoms are connected through chemical bonds, and if connected through chemical bonds, which chemical bonds are used to connect.
  • the graph structure information of the product molecule also includes the chemical bond connection information of the product molecule.
  • the chemical bond connection information includes sub-feature information of each chemical bond in the product molecule.
  • the method of obtaining the sub-feature information of the chemical bond includes: obtaining the attribute information of the chemical bond, performing feature extraction on the attribute information of the chemical bond, and obtaining the sub-feature information of the chemical bond.
  • the attribute information of chemical bonds is used to describe the attributes of chemical bonds.
  • the attribute information of chemical bonds can be set based on experience or flexibly adjusted according to application scenarios, which is not limited in the embodiments of the present application.
  • the attribute information of the chemical bond includes, but is not limited to, at least one of the bond type, conjugation characteristics, ring bond characteristics, bond energy, and bonding distance of the chemical bond.
  • the bond type indicates the type of chemical bond, such as single bond, double bond, triple bond, aromatic bond, ionic bond, covalent bond, metallic bond, etc.
  • the conjugation characteristic indicates whether the chemical bond is conjugated.
  • the ring bond characteristic indicates whether the chemical bond is part of a ring bond.
  • Bond energy is a physical quantity that measures the strength of chemical bonds from energy factors. Generally speaking, the greater the bond energy, the stronger the chemical bond and the less likely it is to break. Bonding distance refers to the shortest distance necessary to form a chemical bond between two or more atomic nuclei.
  • feature extraction is performed on the attribute information of the chemical bond, and the extracted information is used as sub-feature information of the chemical bond.
  • the method of feature extraction for the attribute information of the chemical bond can be set based on experience.
  • a chemical bond feature extraction model is called to extract features from the attribute information of the chemical bond.
  • the chemical bond feature extraction model can be trained through supervised training based on the attribute information of the sample chemical bond and the feature label of the sample chemical bond.
  • Step 2022 The computer device predicts the breaking probability of the chemical bond in the product molecule based on the graph structure information, and determines the chemical bond whose breaking probability meets the reference condition as the broken chemical bond in the product molecule.
  • the breaking probability of a chemical bond indicates the possibility that the chemical bond is a chemical bond formed in a synthesis reaction.
  • the breaking probability of a chemical bond indicates the possibility that the chemical bond is a chemical bond formed in a synthesis reaction.
  • the greater the breaking probability of a chemical bond the greater the probability of the chemical bond breaking.
  • the chemical bond is more likely to be a chemical bond formed during a synthesis reaction.
  • the greater the likelihood that the chemical bond is a chemical bond formed in a synthesis reaction the more reliable it is to break the bond of the product molecule based on the chemical bond.
  • the cleavage probability of chemical bonds in the product molecule refers to the cleavage probability of each chemical bond in the product molecule.
  • the breaking probability of chemical bonds can be predicted based on graph structure information.
  • the graph structure information of the product molecule can indicate the existence of chemical bonds in the product molecule, such as which chemical bonds exist and the atoms connected by each chemical bond. Based on the graph structure information, the probability of breaking each chemical bond in the product molecule can be predicted. .
  • the process of predicting the cleavage probability of chemical bonds in the product molecule can be implemented by running a pre-written program, or by calling a graph neural network model.
  • the embodiments of this application take an example of calling a graph neural network model to predict the breakage probability of chemical bonds in product molecules based on graph structure information.
  • the graph neural network model is a model that can process the graph structure information of compound molecules to predict the breaking probability of chemical bonds in compound molecules, that is, it can distinguish which chemical bonds in product molecules are more likely to break.
  • the embodiments of this application do not limit the model structure of the graph neural network model.
  • the graph neural network model can be any graph-based deep learning network model, and can be designed to be simple or complex.
  • the graph neural network model may refer to the Graph Convolutional Networks (GCN) model, the Graph Attention Networks (GAT) model, the Message Passing Neural Network (MPNN) model, etc.
  • the process of calling the graph neural network model based on the graph structure information to predict the breakage probability of chemical bonds in the product molecules is the internal processing process of the graph neural network model and is related to the model structure of the graph neural network model.
  • the process of calling the graph neural network model to predict the breaking probability of chemical bonds in the product molecule based on the graph structure information includes: calling the graph neural network model to extract the target characteristics of the chemical bonds in the product molecule based on the graph structure information;
  • the target feature of the chemical bond predicts the probability of breakage of the chemical bond in the product molecule.
  • the target feature of a chemical bond is a feature based on which the probability of breaking the chemical bond is predicted.
  • the cleavage probability of a chemical bond can be 1 or 0.
  • the process of predicting the cleavage probability of a chemical bond can be regarded as a process of binary prediction of chemical bonds.
  • a cleavage probability of 1 means that the chemical bond is most likely to be synthesized.
  • a break probability of 0 means that the chemical bond is extremely likely to be a chemical bond formed during a synthesis reaction.
  • the probability of breaking a chemical bond may also be any probability between 0 and 1. The embodiments of the present application do not add any probability to this. to limit.
  • the chemical bonds whose cleavage probability meets the reference conditions can be determined from the chemical bonds in the product molecule, and the chemical bonds whose cleavage probability meets the reference conditions are determined as the broken chemical bonds in the product molecule.
  • the broken chemical bond is the chemical bond that is used to obtain the molecule to be completed from the product molecule.
  • broken chemical bonds may also be referred to as reaction sites.
  • the chemical bonds whose rupture probability meets the reference conditions refer to the chemical bonds that are more likely to be formed during the synthesis reaction, that is, the chemical bonds that are more likely to be broken.
  • the reliability of the bond-breaking treatment of product molecules based on such chemical bonds is higher, thereby improving the acquisition
  • the reliability of the molecules to be completed is improved, thereby improving the reliability of the predicted reactant molecules.
  • the fracture probability that satisfies the reference condition is set based on experience, and the acquisition is flexibly adjusted according to the application scenario, which is not limited in the embodiments of the present application.
  • the fracture probability meeting the reference condition means that the fracture probability is not less than the probability threshold.
  • the probability threshold is set based on experience or flexibly adjusted according to the application scenario. For example, the probability threshold is 0.5, or the probability threshold is 0.8, etc.
  • the rupture probability satisfying the reference condition may also mean that the rupture probability is the largest L (L is an integer not less than 1) of the cleavage probabilities of each chemical bond, and the value of L is set based on experience, or Flexibly adjust according to the application scenario, for example, the value of L is 3, or the value of L is 2, etc.
  • the graph neural network model before calling the graph neural network model to predict the cleavage probability of chemical bonds in the product molecule based on the graph structure information, the graph neural network model needs to be trained first.
  • the process of training a graph neural network model includes: obtaining the graph structure information of the training compound molecules and the standard breaking probability of the chemical bonds in the training compound molecules; calling the graph neural network model based on the graph structure information of the training compound molecules, Predicting a trained break probability of a chemical bond in a training compound molecule; determining a reference loss based on a difference between a standard break probability and a trained break probability; updating model parameters of a graph neural network model based on the reference loss; responsive to the training process satisfying a first termination condition , determine the model currently trained as the graph neural network model that has been trained.
  • Meeting the first termination condition is set based on experience, or can be flexibly adjusted according to the application scenario.
  • meeting the first termination condition means that the reference loss converges, the reference loss is less than the first loss threshold, and the number of model parameter updates reaches the first count threshold, etc.
  • the first loss threshold and the first count threshold are set based on experience or flexibly adjusted according to application scenarios, which are not limited in the embodiments of the present application.
  • Training compound molecules refer to compound molecules that can obtain graph structure information and standard cleavage probabilities of chemical bonds.
  • the number of training compound molecules may be one or multiple, and the embodiments of the present application are not limited to this.
  • the principle of acquiring the graph structure information of the training compound molecules is the same as the principle of acquiring the graph structure information of the product molecules, which will not be described again here.
  • the graph structure information of the training compound molecules can be stored in a database corresponding to the training compound molecules, so that the graph structure information of the training compound molecules can be directly extracted from the database.
  • the standard breaking probability of the chemical bond in the training compound molecule is the real breaking probability of the chemical bond in the training compound molecule, and is used to provide supervision information for the training process of the graph neural network model.
  • the standard breaking probability of the chemical bond in the training compound molecule may also be called the ground-truth label of the breaking probability of the chemical bond in the training compound molecule.
  • the standard breaking probability of the chemical bond in the training compound molecule is stored in the database corresponding to the training compound molecule, so that the standard breaking probability of the chemical bond in the training compound molecule can be directly extracted from the database.
  • the training compound molecule is a molecule synthesized through a known synthesis reaction.
  • the standard breakage probability of the chemical bond in the training compound molecule can be calculated by comparing the training compound molecule and the synthesis of the training compound molecule. obtained by comparing the reactant molecules.
  • the graph neural network model is called to predict the chemical bonds in the training compound molecules based on the graph structure information of the training compound molecules.
  • the implementation principle of training the cleavage probability is the same as the implementation principle of calling the graph neural network model to predict the cleavage probability of the chemical bonds in the product molecule based on the graph structure information of the product molecule, and will not be described again here.
  • the reference loss is determined based on the difference between the standard break probabilities and the training break probabilities.
  • the embodiments of the present application do not limit the measurement method of the difference between the standard break probability and the training break probability.
  • the difference between the standard break probability and the training break probability refers to the difference between the standard break probability and the training break probability.
  • the cross-entropy difference, or the difference between the standard break probability and the training break probability refers to the mean square difference between the standard break probability and the training break probability, etc.
  • the reference loss can be calculated based on Formula 1:
  • L represents the reference loss
  • K (K is an integer not less than 1) represents the number of training compound molecules
  • express basis The determined chemical bond in the k-th training compound molecule is obtained by comprehensively considering each chemical bond in each training compound molecule to obtain the reference loss
  • y uv represents the standard breakage probability of the chemical bond b uv
  • Step 2023 The computer equipment breaks the product molecule based on the broken chemical bond to obtain the molecule to be completed.
  • breaking the bonds of product molecules based on broken chemical bonds means breaking the broken chemical bonds in the product molecules, and using each molecule obtained after breaking the broken chemical bonds in the product molecule as each molecule to be completed. It should be noted that the bond breaking of the product molecule may result in one molecule to be completed or multiple molecules to be completed, which is not limited in the embodiments of the present application.
  • the molecule to be completed is expressed as Among them, H (H is an integer not less than 1) represents the number of molecules to be completed, Represents the hth (h is any value from 1 to H) molecule to be completed.
  • the process of obtaining the molecule to be completed based on the product molecule can be regarded as modeling probability distribution the process of.
  • the process of determining molecules to be completed from product molecules can be called a reaction site prediction process, and the reaction site prediction process can be regarded as the first stage in the reactant molecule prediction process.
  • the first stage can be as shown in (1) in Figure 4.
  • the number of broken chemical bonds in the product molecule is one. After breaking the one broken chemical bond, two molecules to be completed can be obtained. It should be noted that the dotted circles on the two molecules to be completed in (1) in Figure 4 mark the two atoms connected by broken chemical bonds.
  • the process of bond breaking of product molecules based on the above steps 2021 to 2023 is only an exemplary implementation process, and the embodiments of the present application are not limited thereto.
  • the broken chemical bonds can also be selected from the chemical bonds of the product molecules based on experience, and then the product molecules are bond-broken based on the selected broken chemical bonds.
  • step 203 the computer device calls the molecular completion model to complete the molecules to be completed, obtains the completion results, and determines the reactant molecules of the product molecules based on the completion results, where the molecular completion model is based on the sample compound molecules and the sample to-be-completed molecules.
  • the complementary molecules are trained, and the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules.
  • masking the substructure in the sample compound molecule means hiding the substructure in the sample compound molecule.
  • the substructure in the sample compound molecule cannot be known. original state.
  • the method of masking the substructures in the sample compound molecules can be set based on experience, or can be flexibly adjusted according to the actual application scenario, as long as the substructures in the sample compound molecules can be hidden.
  • the method of masking the substructure in the sample compound molecule may be to replace the substructure in the sample compound molecule with a specific structure, or it may be to cover the substructure in the sample compound molecule.
  • a specific structure refers to a structure that is different from the structure of a real compound molecule so as to distinguish masked portions from unmasked portions.
  • the process of predicting reactant molecules based on the molecules to be completed is implemented by calling the molecular completion model.
  • the molecular completion model is trained based on sample compound molecules and sample to-be-completed molecules obtained by masking substructures in the sample compound molecules.
  • the sample molecules to be completed are obtained based on the sample compound molecules themselves.
  • the process of training the sample compound molecules and the sample to be completed molecules to obtain the molecular completion model can be regarded as masking the sample compound molecules and training the model to The process of reconstructing the masked parts. This kind of training process is the process of training the model using the self-supervised learning strategy.
  • the self-supervised learning task is in "Mask and Fill"
  • the basic idea is to mask part of the substructure of the compound molecule, and then train a molecular completion model to reconstruct the substructure.
  • the data based on model training is data obtained based on the sample compound molecules themselves. Regardless of whether the sample compound is a compound in a known synthesis reaction, it can be used as a model
  • the data on which the training is based that is to say, the training process of the model will not be limited by known synthetic reactions.
  • the molecular completion model trained using this training process has strong generalization ability, which is conducive to effective adaptation and The prediction scenarios of reactant molecules related to various synthetic reactions expand the adaptation scenarios, which is beneficial to improving the prediction reliability and accuracy of reactant molecules.
  • each molecule to be completed corresponds to a completion result. According to The completion result of each molecule to be completed can determine one reactant molecule of the product molecule.
  • the principle of calling the molecular completion model to complete each molecule to be completed is the same.
  • the embodiment of this application takes the process of calling the molecular completion model to complete the molecules to be completed as an example to illustrate. It should be noted that when there are multiple molecules to be completed, the molecular completion model can be called to complete multiple molecules to be completed in parallel to improve the efficiency of predicting reactant molecules.
  • the form of the completion result of the molecule to be completed is not limited, as long as the reactant molecule can be determined based on the completion result of the molecule to be completed.
  • the completion result indicates the completed molecule.
  • the completion result of the molecule to be completed can be the graph structure information of the completed molecule of the molecule to be completed.
  • the completed molecule of the molecule to be completed can be determined based on the graph structure information.
  • the completed molecule of the molecule to be completed is regarded as a reactant molecule.
  • the completion result of the molecule to be completed can also be indication information of the missing structure in the molecule to be completed.
  • the indication information includes information used to characterize the missing structure in the molecule to be completed and information used to indicate the missing structure in the molecule to be completed.
  • the missing structure can be determined according to the instruction information, and then the missing structure is connected to the molecule to be completed, and the molecule obtained after the connection is used as a reactant molecule .
  • the information used to characterize the missing structure in the molecule to be completed can be the graph structure information, molecular formula, molecular string, etc. of the missing structure in the molecule to be completed, which is not limited in the embodiments of the present application.
  • the completion result of the molecule to be completed can also be the classification result and connection position of the missing structure in the molecule to be completed.
  • the classification results include the matching probability between the missing structure in the molecule to be completed and the reference structure (such as a reference atom or a reference chemical bond), and the connection position prediction result indicates the position in the molecule to be completed that is to be connected to the missing structure.
  • the category of the missing structure in the molecule to be completed (that is, the category of a single atom or a single chemical bond) can be determined based on the classification results, and then the missing structure is connected to the missing structure at the position in the molecule to be completed.
  • the structure of the category is connected to the molecule to be completed, and the molecule obtained after the connection is used as a reactant molecule.
  • the process of calling the molecular completion model to complete the molecules to be completed and obtaining the completion results of the molecules to be completed is the internal processing process of the molecular completion model, which is related to the structure of the molecular completion model.
  • the embodiments of the present application do not be limited.
  • the molecular completion model is a flow-based generative model (Flow-based Generative Models), and the flow-based generative model is a reversible model. Next, the flow-based generative model is introduced.
  • flow-based generative model Flow-based Generative Models
  • Flow-based generative models directly maximize likelihood estimation (Maximum Likelihood Estimation).
  • the flow-based generation model can give a likelihood estimate of the generated results, that is, the flow-based generation model has stronger interpretability.
  • Equation 2 The expression for the log likelihood estimate logp ⁇ (x) can be It is derived from the reversible mapping combined with the flow-based generative model, as shown in Equation 2:
  • det( ⁇ ) represents the calculation of the determinant of the matrix
  • it is an n ⁇ n matrix, called the Jacobian matrix
  • n is the dimension of z
  • n is an integer not less than 1
  • p ⁇ (z) represents the probability of z under the parameter ⁇
  • p ⁇ (x) represents the probability of x under parameter ⁇ .
  • Flow-based generative models usually adopt a coupling layer (coupling layer) network layer design scheme to take into account both computational efficiency and model representation capabilities.
  • Formula 3 means to copy the first d (d is an integer not less than 1 and not greater than n) dimensions of the input x
  • Formula 4 means to copy the remaining dimensions of the input x (that is, from the d+1th dimension to the nth dimension ) to transform.
  • S ⁇ ( ⁇ ) and T ⁇ ( ⁇ ) in Formula 4 represent two transformation functions, both of which are used to output the same transformation information as the x d+1:n dimension.
  • S ⁇ ( ⁇ ) represents a scale function
  • T ⁇ ( ⁇ ) represents a transformation function.
  • represents the multiplication of the corresponding position elements of the matrix.
  • this coupling layer makes the Jacobian matrix become a diagonal matrix, and the determinant calculation of the diagonal matrix is the product of the diagonal elements, that is Among them, j represents any element in S ⁇ (z 1:d ).
  • the scale function and conversion function can be arbitrarily complex neural networks without increasing the calculation amount of the Jacobian matrix determinant, thereby improving calculation efficiency.
  • the molecule completion model is a flow-based generation model
  • the molecule completion model is called to complete the molecule to be completed
  • the process of obtaining the completion result of the molecule to be completed includes the following steps: 2031 to step 2034.
  • Step 2031 The computer device determines the target atomic feature hidden variable based on the atomic feature information of the molecule to be completed.
  • the target atomic feature latent variable is the atomic feature latent variable of the completed molecule, based on the chemical bond connection information of the molecule to be completed. Determine the target chemical bond connection hidden variable, which is the chemical bond connection hidden variable of the completed molecule.
  • the atomic feature information of the molecule to be completed is used to characterize the features of the atoms in the molecule to be completed, and the target atomic feature hidden variables are used to make assumptions about the features of the atoms in the completed molecule of the molecule to be completed.
  • the principle of obtaining the atomic characteristic information of the molecule to be completed is the same as the principle of obtaining the atomic characteristic information of the product molecule, and will not be described again here.
  • the process of obtaining the target atomic feature latent variable includes: sampling the atomic feature latent variable of the missing structure of the molecule to be completed from a known probability distribution; The atomic characteristic information of the whole molecule and the atomic characteristic hidden variables of the missing structure are used to determine the target atomic characteristic hidden variables.
  • the missing structure of the molecule to be completed refers to the structure of the molecule to be completed, and the atomic characteristic hidden variables of the missing structure are used to make assumptions about the characteristics of the atoms in the missing structure.
  • the atomic feature latent variables of the missing structure are sampled from known probability distributions, that is to say, the atomic feature latent variables of the missing structure are variables that obey the known probability distribution.
  • a known probability distribution is any distribution that can determine the probability of a variable that obeys the probability distribution.
  • the type of the known probability distribution can be set based on experience. This is not limited in the embodiments of the present application.
  • the known probability distribution can refer to Gaussian distribution can also refer to uniform distribution, etc.
  • the method of sampling the atomic feature latent variables of the missing structure may be random sampling.
  • the atomic feature latent variables of the missing structure and the atomic feature information of the molecule to be completed are both in the form of matrices, and the matrix of the atomic feature latent variables of the missing structure and the atomic feature information of the molecule to be completed are in the form , each row of elements corresponds to one atom.
  • the method of determining the target atomic feature latent variable can be: using the atomic feature information of the molecule to be completed and the missing structure
  • the atomic feature latent variables are spliced vertically, and based on the spliced matrix, the target atomic feature latent variables are obtained.
  • the method of obtaining the latent variable of the target atom feature can be: using the matrix obtained by splicing as the latent variable of the target atom feature.
  • the method of obtaining the target atom feature latent variable can also be: if the dimension of the spliced matrix is the first reference dimension, use the spliced matrix as the target atom feature latent variable; if the spliced matrix is obtained The dimension of the matrix is smaller than the first reference dimension, the spliced matrix is expanded to a matrix of the first reference dimension, and the expanded matrix is used as the target atom feature hidden variable.
  • This method can ensure that the dimension of the target atomic feature latent variable is the first reference dimension, thereby improving the standardization of the target atomic feature latent variable.
  • the first reference dimension is a preset parameter used to constrain the dimension of information on atomic characteristics.
  • the first reference dimension can be considered as the dimension of the atomic characteristic information of the largest reactant molecule, that is, the embodiment of the present application considers the first reference dimension Not less than the dimension of the matrix obtained by splicing.
  • the process of expanding the spliced matrix into a matrix of the first reference dimension may include: placing the spliced matrix at the upper left corner and adding 0 elements at other positions until a matrix of the first reference dimension is obtained.
  • the chemical bond connection information of the molecule to be completed is used to characterize the chemical bond connection between the atoms in the molecule to be completed, and the target chemical bond connection hidden variable is used to represent the chemical bond connection between the atoms in the completed molecule of the molecule to be completed. Situation is assumed.
  • the principle of obtaining the chemical bond connection information of the molecule to be completed is the same as the principle of obtaining the chemical bond connection information of the product molecule, and will not be described again here.
  • the process of determining the target chemical bond connection latent variable includes: sampling the chemical bond connection latent variable of the missing structure of the molecule to be completed from a known probability distribution; The chemical bond connection information of the whole molecule and the chemical bond connection hidden variables of the missing structure are used to obtain the target chemical bond connection hidden variables.
  • the chemical bond connection hidden variable of the missing structure is used to make assumptions about the chemical bond connection between atoms in the missing structure.
  • the chemical bond connection latent variables of the missing structure are sampled from known probability distributions, that is to say, the chemical bond connection latent variables of the missing structure are variables that obey the known probability distribution.
  • the method of sampling the chemical bonds connecting the hidden variables of the missing structure may be random sampling.
  • the chemical bond connection hidden variables of the missing structure can indicate the chemical bond connection between the atoms in the missing structure (hypothetical situation)
  • the chemical bond connection information of the molecule to be completed can indicate the chemical bond connection between the atoms in the molecule to be completed.
  • an indicator can be obtained to indicate the atoms in the missing structure and the atoms in the molecule to be completed (i.e.
  • the target information of the chemical bond connection situation (hypothetical situation) between the atoms in the completed molecule) is obtained, and the target chemical bond connection hidden variable is obtained based on the target information.
  • both the target information and the target chemical bond connection hidden variables are in the form of matrices.
  • the chemical bond connection between atoms in the completed molecule includes, in addition to the chemical bond connection between atoms in the missing structure and the chemical bond connection between atoms in the molecule to be completed, it also includes Chemical bonds between atoms in the missing result and atoms in the completed molecule.
  • the chemical bond connection between the atoms in the missing result and the atoms in the completed molecule can be set based on experience. For example, by default, there is no chemical bond connection between the atoms in the missing result and the atoms in the completed molecule.
  • the method of obtaining the target chemical bond connection latent variable based on the target information may be: using the target information as the target chemical bond connection latent variable.
  • the method of obtaining the target chemical bond connection hidden variable based on the target information can also be: if the dimension of the target information is the second reference dimension, use the target information as the target chemical bond connection hidden variable; if the dimension of the target information is smaller than the second reference dimension , expand the target information into a matrix of the second reference dimension, and use the expanded matrix as the target chemical bond connection hidden variable.
  • This method can ensure that the dimension of the hidden variable connected by the target chemical bond is the second reference dimension, thereby improving the standardization of the hidden variable connected by the target chemical bond.
  • the second reference dimension is a preset parameter used to constrain the dimension of information on chemical bond connection.
  • the second reference dimension can be considered as the dimension of chemical bond connection information of the largest reactant molecule. That is, the embodiment of the present application considers the second reference dimension Not smaller than the dimension of the target information.
  • the process of expanding the target information into a matrix of the second reference dimension may include: placing the target information at the upper left corner and adding 0 elements at other positions until a matrix of the second reference dimension is obtained.
  • Step 2032 The computer device calls the molecular completion model to transform the target chemical bond connection hidden variable to obtain the target chemical bond connection information.
  • the target chemical bond connection information is used to characterize the chemical bond connection between atoms in the completed molecule of the molecule to be completed.
  • the target chemical bond connection hidden variable is transformed.
  • the process of obtaining the target chemical bond connection information refers to the process of obtaining the target chemical bond connection information according to the molecule to be completed.
  • the process of predicting the characterization information of the chemical bond connection between the atoms in the completed molecule based on the hypothesis information of the chemical bond connection between the atoms in the completed molecule.
  • the implementation process of calling the molecular completion model to transform the target chemical bond connection latent variable and obtain the target chemical bond connection information includes: calling the molecular completion model to obtain the reference chemical bond connection information of the molecule to be completed based on the target chemical bond connection latent variable. and the reference chemical bond connection hidden variables of the missing structure; transform the reference chemical bond connection hidden variable of the missing structure based on the reference chemical bond connection information of the molecule to be completed to obtain the reference chemical bond connection information of the missing structure; based on the reference chemical bond connection of the molecule to be completed information and the reference chemical bond connection information of the missing structure to obtain the target chemical bond connection information.
  • the target chemical bond connection latent variable is a matrix of the second reference dimension.
  • the method of obtaining the reference chemical bond connection information of the molecule to be completed based on the target chemical bond connection latent variable is: using the target chemical bond connection latent variable for indication.
  • the information on the chemical bond connection between atoms in the molecule to be completed remains unchanged, and other information is set to 0 to obtain the reference chemical bond connection information of the molecule to be completed.
  • the reference chemical bond connection information of the molecule to be completed obtained in this way is also a matrix of the second reference dimension.
  • the method of obtaining the reference chemical bond connection latent variable of the missing structure based on the target chemical bond connection latent variable is: retaining information in the target chemical bond connection latent variable indicating the chemical bond connection status between atoms in the missing structure. Leave unchanged, set other information to 0, and obtain the reference chemical bond connection hidden variable of the missing structure.
  • the reference chemical bond connection hidden variable of the missing structure obtained in this way is also a matrix of the second reference dimension.
  • the reference chemical bond connection hidden variables of the missing structure are transformed based on the reference chemical bond connection information of the molecule to be completed, and the implementation method of obtaining the reference chemical bond connection information of the missing structure can be: based on the reference chemical bond connection information of the molecule to be completed. , obtain the first reference transformation information; use the first reference transformation information to transform the reference chemical bond connection hidden variables of the missing structure to obtain the reference chemical bond connection information of the missing structure.
  • the first reference transformation information may be obtained based on at least one transformation function (eg, the S ⁇ ( ⁇ ) transformation function and the T ⁇ ( ⁇ ) transformation function involved in Equation 6).
  • the first reference transformation information is used to transform the reference chemical bond connection hidden variables of the missing structure
  • the implementation process of obtaining the reference chemical bond connection information of the missing structure can be expressed by formula 6, where x d+1:n is used to represent missing
  • the reference chemical bond connection information of the structure use T ⁇ (z 1:d ) and S ⁇ (z 1:d ) to represent the first reference transformation information obtained based on z 1:d , use z 1:d to represent the molecule to be completed
  • the transformation will not change the dimension of the information, that is, the dimension of the reference chemical bond connection hidden variable of the missing structure is the same as the dimension of the reference chemical bond connection information of the missing structure.
  • the reference chemical bond connection information of the molecule to be completed and the reference chemical bond connection information of the missing structure are both matrices of the second reference dimension. Based on the reference chemical bond connection information of the molecule to be completed and the reference chemical bond connection information of the missing structure, the method is obtained
  • the target chemical bond connection information may be obtained by adding elements at corresponding positions in the reference chemical bond connection information of the molecule to be completed and the reference chemical bond connection information of the missing structure, and the matrix obtained after the addition is used as the target chemical bond connection information.
  • the Cartesian product between the reference chemical bond connection information of the molecule to be completed and the reference chemical bond connection information of the missing structure can also be used as the target chemical bond connection information.
  • the molecule completion model includes a chemical bond completion model.
  • This step 2032 can be implemented by calling the chemical bond completion model in the molecule completion model, that is, calling the chemical bond completion model in the molecule completion model to hide the target chemical bond connection.
  • the variables are transformed to obtain the target chemical bond connection information.
  • the chemical bond completion model can transform the chemical bond connection hidden variables of the molecule into the chemical bond connection information of the molecule.
  • the chemical bond connection hidden variable of the molecule is used to make assumptions about the chemical bond connection between the atoms in the molecule, and the chemical bond connection information of the molecule is used to characterize the chemical bond connection between the atoms in the molecule. That is to say, the target
  • the input of the chemical bond completion model is the hypothesis information of the chemical bond connection between the atoms in the molecule, and the output is the representation information of the chemical bond connection between the atoms in the molecule.
  • the chemical bond completion model is a reversible model, that is, there is an inverse model of the chemical bond completion model.
  • the inverse model of the chemical bond completion model can inversely transform the chemical bond connection information of the molecule into the chemical bond connection hidden variables of the molecule. That is to say, the input of the inverse model of the chemical bond completion model is the representation information of the chemical bond connection between the atoms in the molecule, and the output is the hypothesis information of the chemical bond connection between the atoms in the molecule.
  • the chemical bond completion model is obtained through training. During the training process, the model structure of the chemical bond completion model remains unchanged.
  • the model structure of the chemical bond completion model please refer to the model structure introduced in the embodiment shown in Figure 5, which will not be discussed here for the time being. Repeat.
  • Step 2033 The computer device calls the molecular completion model to transform the target atom feature hidden variables to obtain the target atom feature information.
  • the target atom characteristic information is used to characterize the characteristics of the atoms in the completed molecule of the molecule to be completed.
  • the target atom characteristic hidden variables are transformed.
  • the process of obtaining the target atom characteristic information refers to the process of obtaining the target atom characteristic information based on the completed molecule of the to-be-completed molecule.
  • the implementation process of calling the molecule completion model to transform the target atomic feature latent variables to obtain the target atomic feature information includes: calling the molecule completion model to obtain the reference atomic feature information of the molecule to be completed based on the target atomic feature latent variables. and the reference atomic feature latent variables of the missing structure; transform the reference atomic feature latent variables of the missing structure based on the reference atomic feature information of the molecule to be completed to obtain the reference atomic feature information of the missing structure; based on the reference atomic features of the molecule to be completed Information and reference atomic feature information of missing structures to obtain target atomic feature information.
  • the target atomic feature latent variable is a matrix of the first reference dimension.
  • the method of obtaining the reference atomic feature information of the molecule to be completed based on the target atomic feature latent variable is: using the target atomic feature latent variable to indicate The characteristic information of the atoms in the molecule to be completed remains unchanged, and other information is set to 0 to obtain the reference atomic characteristic information of the molecule to be completed.
  • the reference atomic feature information of the molecule to be completed obtained in this way is also a matrix of the first reference dimension.
  • the method of obtaining the reference atomic feature latent variable of the missing structure based on the target atomic feature latent variable is: keeping the information in the target atomic feature latent variable used to indicate the characteristics of the atoms in the missing structure unchanged, Other information is set to 0 to obtain the reference atomic feature hidden variables of the missing structure.
  • the hidden variable of the reference atomic feature of the missing structure obtained in this way is also a matrix of the first reference dimension.
  • the reference atomic feature hidden variables of the missing structure are transformed based on the reference atomic feature information of the molecule to be completed, and the implementation method of obtaining the reference atomic feature information of the missing structure can be: based on the reference atomic feature information of the molecule to be completed , obtain the second reference transformation information; use the second reference transformation information to transform the reference atomic feature latent variables of the missing structure to obtain the reference atomic feature information of the missing structure.
  • the second reference transformation information may be obtained based on at least one transformation function (eg, the S ⁇ ( ⁇ ) transformation function and the T ⁇ ( ⁇ ) transformation function involved in Equation 6).
  • the second reference transformation information is used to transform the reference atomic feature latent variables of the missing structure
  • the implementation process of obtaining the reference atomic feature information of the missing structure can be expressed by formula 6, where x d+1:n is used to represent missing
  • the reference atomic feature information of the structure use T ⁇ (z 1:d ) and S ⁇ (z 1:d ) to represent the second reference transformation information obtained based on z 1:d , use z 1:d to represent the molecule to be completed
  • z d+1:n is used to represent the reference atomic feature hidden variable of the missing structure. It should be noted that the transformation will not change the dimension of the information, that is, the dimension of the latent variable of the reference atomic feature of the missing structure is the same as the dimension of the reference atomic feature of the missing structure.
  • the reference atomic feature information of the molecule to be completed and the reference atomic feature information of the missing structure are both matrices of the first reference dimension.
  • the The target atom feature information can be obtained by adding elements at corresponding positions in the reference atom feature information of the molecule to be completed and the reference atom feature information of the missing structure, and the matrix obtained after the addition is used as the target atom feature information.
  • the Cartesian product between the reference atom feature information of the molecule to be completed and the reference atom feature information of the missing structure can also be used as the target atom feature information.
  • the process of transforming the target atom characteristic latent variable needs to consider the constraints of the target chemical bond connection information to ensure the reliability of the transformation process. In this case, it is necessary to obtain the constraint information of the target atom feature hidden variables based on the target chemical bond connection information, and then call the target atom completion model to transform the target atom feature hidden variables under the constraints of the constraint information.
  • the differential process of transforming target atomic feature latent variables under the constraints of constraint information is reflected in the basic The process of transforming the reference atomic feature latent variables of the missing structure using the reference atomic feature information of the molecule to be completed to obtain the reference atomic feature information of the missing structure, that is, the basis for transforming the reference atomic feature latent variables of the missing structure.
  • the reference atomic characteristic information of the molecule to be completed it also includes constraint information.
  • the embodiments of the present application do not limit the method of obtaining the constraint information of the target atomic characteristic latent variable. It can be set based on experience, and can also be flexibly adjusted according to the application scenario.
  • the method of obtaining the constraint information of the target atom feature latent variable may be to use the target chemical bond connection information as the constraint information of the target atom feature latent variable.
  • the method of obtaining the constraint information of the target atomic feature latent variable may also be to standardize the target chemical bond connection information to obtain the constraint information of the target atomic feature latent variable. Standardizing the target chemical bond connection information is used to improve the standardization of the target chemical bond connection information.
  • the method of standardization processing can be set based on experience or flexibly adjusted according to the application scenario. This is not limited in the embodiments of the present application.
  • normalization processing can be implemented by calling the Graph Normalization (Graphnorm) module.
  • the molecular completion model includes an atom completion model.
  • This step 2033 can be implemented by calling the atom completion model in the molecular completion model, that is, calling the atom completion model in the molecular completion model to hide the target atomic features.
  • the variables are transformed to obtain target atomic feature information.
  • the atom completion model can transform the atomic characteristic latent variables of a molecule into the atomic characteristic information of the molecule.
  • the atomic characteristic latent variables of the molecule are used to make assumptions about the characteristics of the atoms in the molecule, and the atomic characteristic information of the molecule is used to characterize the characteristics of the atoms in the molecule. That is to say, the input of the atom completion model is the Hypothesis information about the characteristics of the atoms, and the output is representation information about the characteristics of the atoms in the molecule.
  • the atomic completion model is a reversible model, that is, there is an inverse model of the atomic completion model.
  • the inverse model of the atomic completion model can inversely transform the atomic feature information of the molecule into the atomic feature latent variables of the molecule. That is to say, the input of the inverse model of the atom completion model is the representation information of the characteristics of the atoms in the molecule, and the output is the hypothesis information of the characteristics of the atoms in the molecule.
  • the atomic completion model is obtained through training. During the training process, the model structure of the atomic completion model remains unchanged.
  • the model structure of the atomic completion model please refer to the model structure of the atomic completion model in the embodiment shown in Figure 5. No further details will be given here.
  • Step 2034 The computer device determines the completion result of the molecule to be completed based on the target chemical bond connection information and the target atom feature information.
  • the completion result of the molecule to be completed is a result that can determine the completed molecule of the molecule to be completed, and the target chemical bond connection information can indicate the chemical bond connection between the atoms in the completed molecule of the molecule to be completed.
  • the target atom characteristic information can indicate the characteristics of the atoms in the completed molecule of the molecule to be completed, based on the chemical bond connection between the atoms in the completed molecule and the characteristics of the atoms in the completed molecule A completed molecule can be uniquely determined. Therefore, the completion result of the molecule to be completed can be obtained based on the target chemical bond connection information and target atom feature information.
  • the method of obtaining the completion result of the molecule to be completed can be: using the information including the target chemical bond connection information and the target atom feature information as the completion of the to-be-completed molecule result.
  • the method of obtaining the completion result of the molecule to be completed can also be: based on the target chemical bond connection information and the target atom feature information, determine the molecular formula or diagram of the to-be-completed molecule. Structure, use the molecular formula or graph structure as the completion result of the molecule to be completed.
  • a S and B S are the atomic characteristic information and chemical bond connection information of the molecule G s to be completed, and The atomic characteristic hidden variables and chemical bond connection hidden variables of the missing structure in the molecule to be completed; is the target atom characteristic hidden variable of the completed molecule; Connect hidden variables to the target chemical bonds of the completed molecule;
  • the pending information Includes target atom feature latent variables of the completed molecule Connect the hidden variable to the target chemical bond
  • V S , B S and A S respectively represent the atom set, chemical bond connection information and atomic feature information of the molecule to be completed.
  • VR , BR and A R respectively represent the atom set, chemical bond connection information and atomic feature information of the reactant molecule .
  • the flow-based generation model is a non-autoregressive generation model that can generate completion results at one time. Compared with the autoregressive generation model, the generation speed is faster and is used to improve the prediction of reactant molecules. efficiency.
  • the above steps 2031 to 2034 only take the molecular completion model as a flow-based generation model as an example to introduce the implementation process of calling the molecular completion model to complete the molecules to be completed.
  • the molecular completion model can also be other types of generative models, such as a variational autoencoder (VAE) generative model and a generative adversarial model (GAN), etc., which are not limited in the embodiments of the present application.
  • VAE variational autoencoder
  • GAN generative adversarial model
  • the molecular completion model can also be a convolutional neural network model, etc.
  • the molecular completion model is called to complete the molecule to be completed, and the process of obtaining the completion result of the molecule to be completed is also different.
  • the embodiments of the present application will not be introduced one by one here.
  • the molecule completion model is called to complete the molecule to be completed, and the process of obtaining the completion result of the molecule to be completed can be: calling the molecule completion model to extract the characteristics of the molecule to be completed; based on the molecule to be completed Features predict the characteristics of the completed molecule of the molecule to be completed; based on the characteristics of the completed molecule of the molecule to be completed, the completion result of the molecule to be completed is obtained.
  • the process of completing at least one molecule to be completed can be regarded as the second stage in the prediction process of reactant molecules.
  • the second stage on the basis of the molecules to be completed obtained in the first stage, atoms and chemical bonds are added to complete (or reduce) the molecules to be completed to the original reactant molecules.
  • the operations at this stage can be considered as modeling probability distributions It can be solved as a conditional generation problem.
  • G R represents the reactant molecule.
  • the embodiments of this application assume that one molecule to be completed corresponds to one reactant molecule, and the probability distribution that needs to be established can be expressed as in, express reactant molecules, Represents the hth (h is an integer not less than 1) molecule to be completed.
  • the molecular completion model is called to complete two molecules to be completed respectively, and two reactant molecules of the product molecule are obtained.
  • the prediction process of reactant molecules relies on a molecular completion model.
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed. Since the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules, that is to say, the training process of the molecule completion model is based on data obtained based on the sample compound molecules themselves.
  • the training process is a self-supervised training process based on sample compound molecules. This self-supervised training process does not need to pay attention to whether the sample compound molecules are compounds in known synthesis reactions. Therefore, this self-supervised training process is not affected by known compounds. Due to the limitations of the synthetic reaction, the molecular completion model trained using this training process has strong generalization ability, which is conducive to expanding the adaptation scenarios, thereby helping to improve the prediction reliability and accuracy of the reactant molecules.
  • the embodiment of the present application provides a method for training a molecular completion model, which method can be applied to the implementation environment shown in Figure 1 above.
  • the training method of the molecular completion model is executed by a computer device, which may be a terminal 11 or a server 12, which is not limited in the embodiments of the present application.
  • the training method of the molecular completion model provided by the embodiment of the present application includes the following steps 501 and 502.
  • step 501 the computer device acquires sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking substructures in the sample compound molecules.
  • the sample compound molecule is the compound molecule on which the molecular completion model is trained once.
  • the number of sample compound molecules is The amount may be one or multiple, which is not limited in the embodiments of this application.
  • the sample compound molecules can be extracted from any data set including compound molecules.
  • a data set including compound molecules may be a data set including known synthesis reactions, or may be a data set not including known synthesis reactions, which is not limited in the embodiments of the present application.
  • the acquisition of sample compound molecules is highly flexible and is not limited to data sets including known synthetic reactions, which is beneficial to improving the generalization ability of the trained model.
  • the data set including the compounds can be an unlabeled data set.
  • the sample to-be-completed molecule of the sample compound molecule can be obtained, and the sample to-be-completed molecule is obtained by masking the substructure in the sample compound molecule.
  • masking the substructure in the sample compound molecule means hiding the substructure in the sample compound molecule.
  • the substructure in the sample compound molecule cannot be known. original state.
  • the method of masking the substructures in the sample compound molecules can be set based on experience, or can be flexibly adjusted according to the actual application scenario, as long as the substructures in the sample compound molecules can be hidden.
  • the method of masking the substructure in the sample compound molecule may be to replace the substructure in the sample compound molecule with a specific structure, or it may be to cover the substructure in the sample compound molecule.
  • a specific structure refers to a structure that is different from the structure of a real compound molecule so as to distinguish masked portions from unmasked portions.
  • the substructure in the sample compound molecule refers to a part of the sample compound molecule.
  • the embodiments of the present application do not limit the complexity of the masked substructure in the sample compound molecule.
  • the masked substructure in the sample compound molecule The substructure of the code can be an atom in the sample compound molecule, a chemical bond in the sample compound molecule, or a structure composed of at least one atom and at least one chemical bond in the sample compound molecule, etc.
  • the case where the masked substructure in the sample compound molecule is an atom in the sample compound molecule can be shown in (1) in Figure 6 , where the masked substructure in the sample compound molecule is the sample compound molecule.
  • the situation of a chemical bond in a compound molecule can be shown in (2) in Figure 6.
  • the masked substructure in the sample compound molecule is a structure composed of at least one atom and at least one chemical bond in the sample compound molecule.
  • the situation can be shown as (3) in Figure 6.
  • the parts marked with question marks and occluded are the masked substructures.
  • the process of training the molecule completion model can be regarded as The process of training the molecular completion model based on the reconstruction task of a single atom or a single chemical bond.
  • the reconstruction task based on a single atom or a single chemical bond is a relatively simple task, which is beneficial to improving the convergence speed of model training.
  • training a molecule completion model based on the reconstruction task of a single atom or a single chemical bond can be regarded as a multi-classification problem, and the multi-classification problem is used to predict the category of the masked single atom or single chemical bond.
  • a structure composed of at least one atom and at least one chemical bond may be called a subgraph.
  • one of the sample compound molecules is composed of at least one atom and at least one
  • the level at which sample compounds are masked can be called the subgraph level. Masking at this level is beneficial to improving the completion ability of the model.
  • the masked substructures in the sample compound molecules may be empirically selected from the sample compound molecules. For example, you can select any central atom from the sample compound molecule, perform a g (g is an integer not less than 0) jump from the central atom, and use the substructure covered by the g jump as the masked substructure in the sample compound molecule. . This method of selecting the masked substructure in the sample compound molecule is relatively simple.
  • the masked substructures in the sample compound molecules can also be selected from the sample compound molecules by referring to the candidate structure set.
  • the masked substructures in the sample compound molecules are structures in the sample compound that belong to the candidate structure set.
  • a structure belonging to a candidate structure set refers to a structure that constitutes the candidate structure set.
  • the candidate structure set is a set of structures whose credibility meets the selection conditions. Structures whose credibility meets the selection conditions refer to structures with higher credibility.
  • the masked substructures are selected from the sample compound molecules, which is beneficial to improving the masked substructures. The rationality can avoid the collapse of the overall structure, thereby improving the completion performance of the model.
  • structures that meet the selection criteria with confidence may include products in known synthesis reactions.
  • Known synthesis reactions can be extracted from retrosynthesis data sets, and retrosynthesis data sets can be selected based on experience.
  • the retrosynthesis data set is the USPTO-50K data set, which contains 50,000 retrosynthesis reactions.
  • the synthesis reaction is a known synthesis reaction.
  • structures whose credibility satisfies the selection conditions may also include motifs (which can be called motifs, which are the basic structures that constitute any characteristic sequence), functional groups whose occurrence frequency is greater than the frequency threshold, and functional groups whose occurrence frequency is greater than the frequency threshold.
  • motifs which can be called motifs, which are the basic structures that constitute any characteristic sequence
  • functional groups whose occurrence frequency is greater than the frequency threshold and functional groups whose occurrence frequency is greater than the frequency threshold.
  • An algorithm for splitting molecules the structure obtained by splitting the reference molecule, etc.
  • the frequency of occurrence can be the frequency of occurrence in certain articles or journals, or the frequency of occurrence in retrosynthetic data sets, etc.
  • Reference molecules can be selected empirically.
  • the candidate structure set may also be called a substructure dictionary.
  • the construction method of the substructure dictionary can be selected flexibly, as long as the internal structure is ensured to be a trustworthy structure that meets the selection conditions. Under different construction methods, the average size and average frequency of occurrence of structures in the substructure dictionary may be different.
  • the implementation process of obtaining the sample molecule to be completed can be: select any one from the sample compound molecule Chemical bond, cut the chemical bond to obtain two structures, match the structure with a smaller number of atoms in the two structures with the candidate structure set, if the match is successful (that is, the structure with a smaller number of atoms belongs to the candidate structure set), then Determine that the cutting plan for the structure with a smaller number of atoms is reasonable. Use the structure with a smaller number of atoms as a masked substructure in the sample compound molecule. Mask the masked substructure to obtain the sample to be filled. Whole molecule.
  • the chemical bond is reselected for cutting.
  • This implementation process can be called the process of obtaining molecules to be completed from the sample based on molecular decomposition. This kind of molecular cutting can ensure that the masked substructure and the remaining parts retain meaningful structures, reducing the difficulty of completion (or reconstruction).
  • sample compound molecule under different masking schemes, different molecules to be completed can be obtained.
  • the sample compound molecule and each molecule to be completed can form a data pair (pair).
  • the model training process in the application embodiment is performed on the basis of data pairs. That is to say, the sample molecules to be completed in the embodiments of this application refer to any molecules to be completed in the sample compound molecules.
  • step 502 the computer device determines the training loss based on the sample compound molecules, the sample molecules to be completed, and the molecule completion model; updates the model parameters of the molecule completion model based on the training loss to obtain the trained molecule completion model.
  • the training loss is used to provide supervision information for the update of model parameters of the molecular completion model.
  • the implementation method of obtaining the training loss based on the sample compound molecules, sample molecules to be completed and the molecular completion model is related to the type of the molecular completion model. This application The examples are not limiting.
  • the computer device determines the training loss based on the sample compound molecules, the sample molecules to be completed, and the molecule completion model, including the following steps 5021 to 5025.
  • Step 5021 The computer device obtains the sample atomic characteristic information and the sample chemical bond connection information of the sample compound molecule.
  • the sample atomic characteristic information of the sample compound molecule is used to characterize the characteristics of the atoms in the sample compound molecule
  • the sample chemical bond connection information of the sample compound molecule is used to characterize the chemical bond connection between the atoms in the sample compound molecule.
  • the principle of obtaining the sample atomic characteristic information and the sample chemical bond connection information of the sample compound molecule is the same as the principle of obtaining the atomic characteristic information and chemical bond connection information of the product molecule in the embodiment shown in Figure 2, and will not be described again here.
  • the sample chemical bond connection information of the sample compound molecules is a matrix of the first reference dimension
  • the sample atomic characteristic information of the sample compound molecules is a matrix of the second reference dimension.
  • Step 5022 The computer device determines the atom mask information and the chemical bond mask information based on the difference between the sample compound molecule and the sample molecule to be completed.
  • the atom mask information indicates the masking status of the atoms in the sample compound molecule, such as which atoms are masked and which atoms are not masked.
  • the chemical bond mask information indicates the masking status of the chemical bonds between atoms in the sample compound molecule. For example, which chemical bonds are masked and which chemical bonds are not masked. Since the molecule to be completed in the sample is passed to the sample compound Therefore, by comparing the difference structures between the sample compound molecules and the sample molecules to be completed, the atom mask information and chemical bond mask information can be determined.
  • the dimensions of the atom mask information and the sample atom feature information are the same, for example, they are both matrices of the second reference dimension.
  • the value of the element at any position in the atomic mask information indicates the masking situation of the element at the same position in the sample atomic feature information. For example, if the value of the element at any position in the atomic mask information is 0, then indicates that the masking status of the elements located at the same position in the sample atomic feature information is masked; if the value of the element at any position in the atomic mask information is 1, it indicates that the elements located at the same position in the sample atomic feature information The element's mask status is not masked.
  • the dimensions of the chemical bond mask information and the sample chemical bond connection information are the same, for example, both are matrices of the first reference dimension.
  • the value of an element at any position in the chemical bond mask information indicates the masking situation of elements at the same position in the chemical bond connection information of the sample. For example, if the value of an element at any position in the chemical bond mask information is 0, It indicates that the masking status of the elements at the same position in the sample chemical bond connection information is masked; if the value of the element at any position in the chemical bond mask information is 1, it indicates that the elements at the same position in the sample chemical bond connection information are masked. The element's mask status is not masked.
  • Step 5023 The computer device calls the molecular completion model, performs inverse transformation on the sample chemical bond connection information based on the chemical bond mask information, and obtains the sample chemical bond connection hidden variables.
  • the sample chemical bond connection latent variable is used to make assumptions about the chemical bond connection between atoms in the sample compound molecule.
  • the molecular completion model is a reversible model that can not only transform the chemical bond connection latent variables of the molecule into the chemical bond connection information of the molecule, but also reversely transform the chemical bond connection information of the molecule into the chemical bond connection latent vector of the molecule. Since during the training process of the model, the sample chemical bond connection information of the sample compound molecules is known and relatively accurate information, the molecular completion model is called to implement the inverse transformation of the sample chemical bond connection information.
  • the computer device calls the molecular completion model to perform an inverse transformation on the sample chemical bond connection information based on the chemical bond mask information.
  • the process of obtaining the sample chemical bond connection hidden variables includes the following steps 50231 to 50233.
  • Step 50231 The computer device calls the molecular completion model to determine the first chemical bond connection information and the second chemical bond connection information based on the chemical bond mask information and the sample chemical bond connection information.
  • the first chemical bond connection information is the chemical bond connection information of the sample molecule to be completed
  • the second chemical bond connection information is the chemical bond connection information of the substructure.
  • the first chemical bond connection information represents the chemical bond connection between atoms in the sample molecule to be completed
  • the second chemical bond connection information represents the chemical bond connection between atoms in the masked substructure. Since the chemical bond mask information can indicate the masking of chemical bonds between atoms in the sample compound molecule, the masked chemical bonds are the chemical bonds between the atoms in the masked substructure, and the unmasked chemical bonds are It is the chemical bond between the atoms in the sample molecule to be completed, so based on the chemical bond mask information and the sample chemical bond connection information of the sample compound molecule, the first chemical bond connection information of the sample to be completed molecule and the second chemical bond of the substructure can be obtained Connection information.
  • the method of determining the first chemical bond connection information and the second chemical bond connection information may be: combining the unmasked portions of the sample chemical bond connection information with those indicated by the chemical bond mask information.
  • Information related to chemical bonds is retained, and other information is set to 0 to obtain the first chemical bond connection information;
  • information related to the masked chemical bonds indicated by the chemical bond mask information in the sample chemical bond connection information is retained, and other information is set to 0.
  • the dimensions of the first chemical bond connection information and the second chemical bond connection information are the same as the matrix of the sample chemical bond connection information, for example, both are matrices of the first reference dimension.
  • Step 50232 The computer device performs an inverse transformation on the second chemical bond connection information based on the first chemical bond connection information to obtain the chemical bond connection hidden variable of the substructure.
  • the chemical bond connection hidden variable of the substructure is used to make assumptions about the chemical bond connection between atoms in the substructure.
  • the chemical bond connection hidden variable of the substructure is a variable that obeys a known probability distribution, and the probability distribution is known Set based on experience, or flexibly adjust according to application scenarios.
  • the known probability distribution is Gaussian distribution, or uniform distribution, etc.
  • the implementation process of performing inverse transformation on the second chemical bond connection information based on the first chemical bond connection information to obtain the chemical bond connection hidden variables of the substructure includes: obtaining the first sample transformation information based on the first chemical bond connection information. ; Use the first sample transformation information to perform an inverse transformation on the second chemical bond connection information to obtain the chemical bond connection hidden variables of the substructure.
  • the first sample transformation information may be based on at least one transformation function (eg, S ⁇ ( ⁇ ) transformation function and T ⁇ ( ⁇ ) transformation function) to obtain. It should be noted that the inverse transformation will not change the dimension of the information, that is, the dimension of the hidden variable connected by the chemical bond of the substructure is the same as the dimension of the information connected by the second chemical bond.
  • Step 50233 The computer device determines the chemical bond connection latent variable of the sample based on the first chemical bond connection information and the chemical bond connection latent variable of the substructure.
  • the first chemical bond connection information and the chemical bond connection latent variables of the substructure are both matrices of the first reference dimension.
  • the method of obtaining the sample chemical bond connection latent variables can be: Add the first chemical bond connection information and the elements at corresponding positions in the chemical bond connection hidden variables of the substructure, and use the matrix obtained after the addition as the sample chemical bond connection hidden variable.
  • the Cartesian product between the first chemical bond connection information and the chemical bond connection latent variable of the substructure may also be used as the sample chemical bond connection latent variable.
  • the molecular completion model includes a chemical bond completion model.
  • the chemical bond completion model is a flow-based generation model, that is, the chemical bond completion model is used to complete chemical bond connection information in a flow-based generation manner.
  • the molecular completion model may also be called a Synthon Flow (Synthon Flow) model, and the chemical bond completion model may also be called a Synthon Bond Flow (Synthon Bond Flow, SB Flow for short) model.
  • the chemical bond completion model is a reversible model.
  • the relationship between the inverse model of the chemical bond completion model and the chemical bond completion model is: the input of the inverse model of the chemical bond completion model is the output of the chemical bond completion model, and the inverse model of the chemical bond completion model
  • the output of is the input of the chemical bond completion model.
  • the dimensions of the input and output information of the chemical bond completion model are the same.
  • the input of the chemical bond completion model is the hypothesis information of the chemical bond connection between the atoms in the molecule
  • the output is the characterization information of the chemical bond connection between the atoms in the molecule, that is, the inverse model of the chemical bond model.
  • the input is the representation information of the chemical bond connection between the atoms in the molecule
  • the output is the hypothesis information of the chemical bond connection between the atoms in the molecule.
  • step 5023 can be performed by calling the inverse of the chemical bond completion model in the molecule completion model.
  • Model implementation That is to say, the inverse model of the chemical bond completion model in the molecular completion model is called to perform an inverse transformation on the sample chemical bond connection information based on the chemical bond mask information to obtain the sample chemical bond connection hidden variables.
  • the chemical bond completion model includes at least one chemical bond completion module, and each chemical bond completion module has the same structure.
  • the embodiment of the present application takes the chemical bond completion model including one chemical bond completion module as an example to illustrate.
  • the atom completion model includes a squeeze (Squeeze) module, a normalization processing (Actnorm) module, an invertible convolution (Invertible Convolution) module, a split/mask ( Split/Mask) module, an affine coupling (Affine Coupling) module and at least one transformation information acquisition module.
  • the transformation information acquisition module includes a convolution sub-module, a normalization processing (Batchnorm) sub-module and an activation (Relu) sub-module.
  • the number of transformation information acquisition modules is l (l is an integer not less than 1)
  • the convolution kernel of the reversible convolution module is 1*1
  • the convolution of the convolution submodule in the transformation information acquisition module The core is 3*3.
  • the extrusion module is used to transform the dimensions of the input chemical bond connection information
  • the standardization processing module is used to standardize the information
  • the reversible convolution module is used to rearrange the category dimensions in the chemical bond connection information
  • split The /mask module is used to split the sample chemical bond connection information of the sample compound into two parts (the first chemical bond connection information and the second chemical bond connection information) based on the chemical bond mask information
  • the affine coupling module is used to implement the second chemical bond Inverse transformation of connection information
  • l transformation information acquisition modules are used to obtain the first sample transformation information.
  • the process of obtaining the sample chemical bond connection hidden variables can be: input the chemical bond mask information M B and the sample chemical bond connection information B R into the chemical bond completion model, in sequence After processing by the extrusion module, normalization processing module, reversible convolution module and splitting/masking module, the first chemical bond connection information is obtained and second chemical bond connection information Utilize l transformation information acquisition modules (each transformation information acquisition module includes a convolution sub-module, a normalization processing sub-module and an activation sub-module) to obtain the first chemical bond connection information Perform processing to obtain the first sample transformation information and The first sample transformation information is then used by the affine coupling module and To the second chemical bond connection information Perform inverse transformation to obtain the chemical bond connection hidden variables of the substructure. Based on the first chemical bond connection information and substructure chemical bonds connecting hidden variables Get sample chemical bond connection hidden variables
  • the structure of the chemical bond completion model shown in Figure 8 is only an illustrative example, and the embodiments of the present application are not limited thereto. In other words, the structure of the chemical bond completion model can also include more or fewer modules.
  • inverse transformation of the sample chemical bond connection information based on the chemical bond mask information may be directly inverse transformation of the sample chemical bond connection information based on the chemical bond mask information, or may refer to the process of inverse transformation based on the chemical bond mask information.
  • the chemical bond connection information is inversely transformed, and the processed chemical bond connection information can be obtained by calling the GLOW module (an information processing module) to process the sample chemical bond connection information.
  • Step 5024 The computer device calls the molecular completion model, performs inverse transformation on the sample atomic feature information based on the atomic mask information, and obtains the sample atomic feature latent variable.
  • the sample atom characteristic latent variable is used to make assumptions about the characteristics of the atoms in the sample compound molecule.
  • the molecular completion model is a reversible model that can not only transform the atomic characteristic latent variables of the molecule into the atomic characteristic information of the molecule, but also inversely transform the atomic characteristic information of the molecule into the atomic characteristic latent vector of the molecule. Since during the training process of the model, the sample atomic feature information of the sample compound molecules is known and relatively accurate information, the molecule completion model is called to implement the inverse transformation of the sample atomic feature information.
  • the computer device calls the molecular completion model to perform inverse transformation on the sample atomic feature latent variables based on the atomic mask information.
  • the process of obtaining the sample atomic feature latent variables includes the following steps 50241 to 50243.
  • Step 50241 The computer device calls the molecule completion model, and determines the first atomic feature information and the second atomic feature information based on the atomic mask information and the sample atomic feature information.
  • the first atomic feature information is the atomic feature of the sample molecule to be completed.
  • information and the second atomic characteristic information is the atomic characteristic information of the substructure.
  • the first atomic feature information is used to characterize the atoms in the molecule to be completed in the sample
  • the second chemical bond connection information is used to characterize the atoms in the masked substructure. Since the atom mask information can indicate the masking status of atoms in the sample compound molecule, the masked atoms are the atoms in the masked substructure, and the unmasked atoms are the atoms in the sample molecule to be completed. atoms, so based on the atom mask information and the sample atomic characteristics of the sample compound molecule, the first atomic characteristic information of the sample molecule to be completed and the second atomic characteristic information of the substructure can be obtained.
  • the first atomic feature information of the sample molecule to be completed and the second atomic feature information of the substructure can be obtained by: combining the sample atomic feature information with The information related to the unmasked atoms indicated by the atomic mask information is retained, and other information is set to 0 to obtain the first atomic feature information; the masked atoms indicated by the atomic mask information in the sample atomic feature information are Relevant information is retained, and other information is set to 0 to obtain the second atomic characteristic information.
  • the dimensions of the first atomic feature information and the second atomic feature information are the same as the matrix of the sample atomic feature information, for example, both are matrices of the second reference dimension.
  • Step 50242 The computer device performs an inverse transformation on the second atomic feature information based on the first atomic feature information to obtain the atomic feature latent variable of the substructure.
  • the atomic characteristic latent variable of the substructure is used to make assumptions about the characteristics of the atoms in the substructure.
  • the atomic characteristic latent variable of the substructure is a variable that obeys a known probability distribution.
  • the known probability distribution is set based on experience. Or it can be flexibly adjusted according to the application scenario.
  • the known probability distribution is Gaussian distribution, or uniform distribution, etc.
  • the implementation process of performing inverse transformation on the second atomic feature information based on the first atomic feature information to obtain the atomic feature latent variables of the substructure includes: obtaining the second sample transformation information based on the first atomic feature information;
  • the second sample transformation information is used to perform an inverse transformation on the second atomic feature information to obtain the atomic feature latent variable of the substructure.
  • the second sample transformation information may be based on at least one transformation function (eg, S ⁇ ( ⁇ ) transformation function and T ⁇ ( ⁇ ) transformation function) Obtain. It should be noted that the inverse transformation will not change the dimension of the information, that is, the dimension of the atomic feature latent variable of the substructure is the same as the dimension of the second atomic feature information.
  • the process of inverse transformation of the second atomic feature information based on the first atomic feature information needs to consider the constraints of the sample chemical bond connection information to ensure the reliability of the inverse transformation process.
  • the second sample transformation information is obtained by comprehensively considering the first atomic feature information and the sample constraint information.
  • the embodiments of this application do not limit the method of obtaining sample constraint information, which can be set based on experience or flexibly adjusted according to application scenarios.
  • the sample constraint information may be obtained by using the sample chemical bond connection information as the sample constraint information.
  • the method of obtaining the sample constraint information may also be to standardize the chemical bond connection information of the sample to obtain the sample constraint information. Standardizing the chemical bond connection information of the sample is used to improve the standardization of the chemical bond connection information of the sample.
  • the method of standardization processing can be set based on experience or flexibly adjusted according to the application scenario.
  • the embodiments of the present application are not limited to this.
  • normalization processing can be implemented by calling the graph normalization module.
  • Step 50243 The computer device obtains the sample atomic feature latent variable based on the first atomic feature information and the atomic feature latent variable of the substructure.
  • the first atomic feature information and the atomic feature latent variables of the substructure are both matrices of the second reference dimension.
  • the method of obtaining the sample atomic feature latent variables can be: Add the first atomic feature information and the elements at corresponding positions in the atomic feature latent variable of the substructure, and use the matrix obtained after the addition as the sample atomic feature latent variable.
  • the Cartesian product between the first atomic feature information and the atomic feature latent variable of the substructure can also be used as the sample atomic feature latent variable.
  • Formula 9 is used to combine the first atomic feature information of the sample to be completed molecule Remaining unchanged
  • Equation 10 is used for the second atomic feature information of the substructure to be masked Transformed to a hidden variable obeying Gaussian distribution
  • Formula 10 converts the first atomic feature information and sample constraint information obtained based on sample chemical bond connection information B R as an input condition.
  • S ⁇ and T ⁇ can adopt a neural network structure, such as a graph neural network structure.
  • the output dimensions of S ⁇ and T ⁇ are both related to the second atomic feature information. have the same dimensions, the processing logic of S ⁇ and T ⁇ can be expressed by Formula 11:
  • h A represents the output structure of S ⁇ or T ⁇ function (that is, or );
  • M A ⁇ 0,1 ⁇ represents the atomic mask information, which is used to mask the atomic feature information of non-sample molecules to be completed (that is, the masked substructure) to zero, and to mask the sample molecules to be completed.
  • the atomic characteristic information remains unchanged.
  • the dimensions of M A are the same as the dimensions of the sample atomic feature information.
  • v a ⁇ V S represents that atom a is an atom in the atom set V S composed of atoms in the sample molecule to be completed, and M[a,j] represents the element of the jth dimension in the sub-feature information of atom a. mask value.
  • Graphconv() represents the graph neural network structure; Wi and W 0 represent the parameters of the graph neural network structure.
  • the molecular completion model includes an atom completion model
  • the atom completion model is a flow-based generation model
  • the atomic completion model is used to complete atomic feature information based on flow generation.
  • the atomic completion model may also be called a Synthon Graph Flow (Synthetic Subgraph Structure Flow, SG Flow for short) model.
  • the atomic completion model is a reversible model.
  • the relationship between the inverse model of the atomic completion model and the atomic completion model is: the input of the inverse model of the atomic completion model is the output of the atomic completion model, and the inverse model of the atomic completion model
  • the output of is the input of the atomic completion model.
  • the dimensions of the input and output information of the atomic completion model are the same.
  • the input of the atom completion model is the hypothesis information of the characteristics of the atoms in the molecule, and the output is the characteristic representation information of the atoms in the molecule. That is to say, the input of the inverse model of the chemical bond model is the characteristics of the atoms in the molecule. representation information, and the output is hypothesis information of atoms in the molecule.
  • step 5024 can be implemented by calling the inverse model of the atomic completion model in the molecule completion model. That is to say, the inverse model of the atom completion model in the molecular completion model is called to inversely transform the sample atom feature information based on the atom mask information to obtain the sample atom feature latent variable.
  • the atomic completion model includes at least one atomic completion module, and each atomic completion module has the same structure.
  • This embodiment of the present application takes the atomic completion model including one atomic completion module as an example for description.
  • the atomic completion model includes a normalization processing (Actnorm) module, a split/mask (Split/Mask) module, an affine coupling (Affine Coupling) module, Figure Standardization (Graphnorm) module and transformation information acquisition module.
  • the transformation information acquisition module includes at least one reference processing module and a multi-layer perceptron (MLP) module.
  • Each reference processing module includes a graph convolution sub-module, a normalization processing (Batchnorm) sub-module and an activation (Relu) sub-module. module.
  • the number of reference processing modules is l (l is an integer not less than 1)
  • the splitting/masking module is used to split the sample atomic feature information of the sample compound into two parts based on the atomic mask information (first atomic feature information and second atomic feature information);
  • the standardization processing module is used to perform standardization operations on each row of the matrix in each batch;
  • the affine coupling module is used to implement the second atomic feature information Inverse transformation;
  • the transformation information acquisition module is used to obtain the second sample transformation information.
  • the process of obtaining the sample atomic feature latent variables can be: input the atomic mask information M A and the sample atomic feature information A R into the atomic completion model, in sequence After processing by the standardized processing module and splitting/masking module, the first atomic feature information is obtained and second atomic characteristic information
  • the sample chemical bond connection information B R is standardized through the graph normalization module to obtain the sample constraint information.
  • a transformation information acquisition module including l reference processing modules (each reference processing module includes a graph convolution sub-module, a normalization processing sub-module and an activation sub-module) and an MLP module to obtain the first atomic feature information and sample constraint information Perform processing to obtain the second sample transformation information and The second sample transformation information is then used by the affine coupling module and For the second atomic characteristic information Perform inverse transformation to obtain the atomic characteristic hidden variables of the substructure. Based on the first atomic characteristic information Atomic characteristic hidden variables of substructures Obtain sample atomic feature hidden variables
  • the structure of the atomic completion model shown in Figure 9 is only an illustrative example, and the embodiments of the present application are not limited thereto. In other words, the structure of the atomic completion model can also include more or fewer modules.
  • Step 5025 The computer device connects the latent variables and the sample atomic feature latent variables based on the sample chemical bonds to determine the training loss.
  • the training loss is used to measure the prediction quality of the sample chemical bond connection latent variables and the sample atomic feature latent variables.
  • the training loss is a numerical value. The greater the training loss, the worse the prediction quality of the sample chemical bond connection hidden variables and the sample atomic feature latent variables, which means the worse the performance of the molecular completion model; the smaller the training loss, the sample chemical bond connection. The better the prediction quality of hidden variables and sample atomic features, the better the performance of the molecular completion model.
  • the implementation process of determining the training loss based on the sample chemical bond connection latent variable and the sample atomic feature latent variable includes: determining the first likelihood function value based on the sample chemical bond connection latent variable and the first sample transformation information; Based on the sample atomic feature latent variable and the second sample transformation information, determine the second likelihood function value; based on the first likelihood function value and the second likelihood function value, determine the target likelihood function value; combine it with the target likelihood function value
  • the value with a negative correlation is determined as the training loss. For example. Use the opposite number of the target likelihood function value as the training loss, or use the opposite number of the target likelihood function value The opposite of the value is used as the training loss, etc.
  • the first likelihood function value is used to measure the probability of calling the chemical bond completion model to obtain the sample chemical bond connection information based on the given sample chemical bond connection hidden variables.
  • the first likelihood function value can be based on the following formula 12 Calculated:
  • B R represents the sample chemical bond connection information
  • the second likelihood function value is used to measure the probability of obtaining the sample atomic feature information by calling the chemical bond completion model based on the given sample atomic feature hidden variables.
  • the second likelihood function value can be based on the following formula 13 Calculated:
  • Equation 14 the process of obtaining the target likelihood function value is as shown in Equation 14:
  • G R represents the sample compound molecule
  • BR represents the sample chemical bond connection information
  • the process of obtaining the target likelihood function value can be as follows:
  • sample compound molecule information G R (A R , B R ) and its mask information M (including atom mask information and chemical bond mask information), the inverse model of the atom completion model and the inverse model of the chemical bond completion model Among them, A R represents the sample atomic characteristic information, and BR represents the sample chemical bond connection information;
  • the process of determining the training loss based on steps 5021 to 5025 is only an exemplary implementation process, and the embodiments of the present application are not limited thereto.
  • the implementation of obtaining the training loss based on the sample compound molecules, the sample molecules to be completed, and the molecule completion model can also be: the computer device calls the molecule completion model to complete the sample molecules to be completed, and obtains Based on the completion results, the predicted completed molecules are determined; based on the difference between the predicted completed molecules and the sample compound molecules, the training loss is determined.
  • the implementation process of calling the molecule completion model to complete the sample molecules to be completed is as shown in step 203 in the embodiment shown in Figure 2, which will not be described again here.
  • the predicted completion molecules can be obtained.
  • the predicted completion molecules are the completed molecules of the sample to be completed molecules predicted by the molecular completion model.
  • the sample compound molecule is the real completed molecule of the sample molecule to be completed. Based on the difference between the predicted completed molecule and the sample compound molecule, training for providing supervision information for model parameter update of the molecule completion model can be obtained. loss.
  • the embodiments of the present application do not limit the way to measure the difference between two molecules. For example, based on the difference in atoms in the two molecules (difference in the number of atoms, difference in atom type, difference in atomic characteristics, etc.) and The difference between chemical bonds (difference in number of chemical bonds, difference in type of chemical bond, difference in characteristics of chemical bonds, etc.) determines the difference between two molecules. For example, molecular features of two molecules are extracted, and the difference between the two molecular features is taken as the difference between the two molecules. For example, the molecular features of two molecules can be extracted by calling a molecular feature extraction model.
  • the two molecular features are vectors or matrices of the same dimension, and the difference between the two molecular features can be determined based on the difference between elements at corresponding positions in the two molecular features.
  • the similarity between two molecular features can also be calculated, and a value that is negatively correlated with the similarity (eg, the opposite number of the similarity) is used as the difference between the two molecular features.
  • the model parameters of the molecular completion model are updated based on the training loss.
  • the process of updating the model parameters of the molecule completion model based on the training loss can be implemented based on the gradient descent method, that is, the update gradient of the model parameters of the molecule completion model is obtained based on the training loss, and the molecule completion model is updated based on the update gradient.
  • model parameters For example, for the case where the training loss is determined based on steps 5021 to 5025, the process of updating the model parameters of the molecule completion model based on the training loss may also refer to the process of updating the inverse model of the molecule completion model based on the training loss.
  • the model training process is an iterative process. After updating the model parameters of the molecule completion model based on the training loss, a molecule completion model that has been trained once is obtained. It is determined whether the current training process meets the target termination condition. If the current training process meets the target If the termination condition is determined, the molecule completion model trained once can be used as the trained molecule completion model; if the current training process does not meet the target termination condition, the new training loss can be obtained by referring to steps 501 and 502, and then use The new training loss updates the model parameters of the currently obtained molecule completion model, and so on, until the current training process meets the target termination condition, and the molecule completion model obtained when the target termination condition is met is used as the post-training molecule completion Model. It should be noted that the sample compound molecules and sample to-be-completed molecules based on which the new training loss is obtained may be partially or completely changed, or may not be changed.
  • satisfying the target termination are set based on experience, or flexibly adjusted according to the application scenario.
  • This embodiment of the present application does not add to limit.
  • satisfying the target termination condition may mean that the training loss converges, the training loss is less than the second loss threshold, the number of update times of the model parameters reaches the second number threshold, etc.
  • the second loss threshold and the second times threshold can be set based on experience, or can be flexibly adjusted according to application scenarios, which are not limited in the embodiments of the present application.
  • the process of training the molecular completion model can be a curriculum learning process.
  • tasks from easy to difficult can be constructed to gradually train the molecular completion model.
  • the difficulty of the task can be measured based on the complexity of the masked substructures in the sample compound molecules or the masking ratio of the masked substructures in the sample compound molecules.
  • the model can receive more information and the molecule completion task will be easier; If the complexity of the masked substructures in the sample compound molecules is high or the proportion of masked substructures in the sample compound molecules is high, the model can receive less information and the molecule completion task will be more difficult.
  • the training effect of the molecular completion model can be gradually improved.
  • complex tasks such as tasks constructed under a subgraph-level masking scheme
  • satisfying the target termination condition may also refer to completing training of the molecule completion model based on the most difficult task.
  • training the molecule completion model based on tasks of any difficulty refers to training the molecule completion model based on sample data of any difficulty.
  • the sample data is sample compound molecules and their samples that match the any difficulty. numerator to be completed).
  • Completion of training the molecule completion model based on tasks of any difficulty may mean that during the process of training the molecule completion model based on sample data of any difficulty, the loss reaches convergence, or the loss is less than a certain loss threshold, or The number of training times reaches a certain number threshold, etc.
  • the process of training the molecular completion model can also be a process of pre-training and then fine-tuning, where the sample data based on pre-training is obtained based on compound molecules in any unlabeled compound molecule data set,
  • the sample data on which fine-tuning is based is based on known product molecules from synthesis reactions.
  • meeting the target termination condition may mean that the fine-tuning of the molecular completion model is completed. For example, during the training of the molecular completion model based on the sample data based on fine-tuning, the loss reaches convergence, or the loss is less than a certain loss. threshold, or the number of training times reaches a certain threshold, etc.
  • model training process can be set based on experience, or according to the computing power and application of the computer equipment. Scenes, etc. can be flexibly adjusted, which is not limited by the embodiments of the present application.
  • the training process of the molecular completion model provided by the embodiments of the present application can be pre-trained on an unlabeled large data set. This operation indirectly performs data augmentation and expands more available effective information, thus It can improve the generalization ability of the molecular completion model and adapt to a wider range of reactant molecule prediction scenarios, thereby improving the prediction reliability and accuracy of reactant molecules.
  • the pre-trained molecule completion task is very relevant to the completion task of molecules to be completed in the reactant molecule prediction process, after pre-training, it can be used in data sets that include known synthesis reactions (such as retrosynthesis data). Fine-tune the molecular completion model based on the set) so that the model will have stronger generalization ability.
  • the completion task of molecules to be completed in the reactant molecule prediction process can be considered as some special cases of the general molecule completion task.
  • the embodiments of this application use molecular reconstruction based on graph representation as a self-supervised learning task.
  • the solution can expand previous molecular self-supervised learning tasks and better apply the idea of "Mask and Fill" to graph-structured data. field.
  • Apply the self-supervised learning strategy to the field of prediction of reactant molecules, learn the model on self-supervised tasks (such as molecular structure completion) on large molecular data sets, and then fine-tune on retrosynthetic data sets. Finally, it can be given The product molecules directly predict the reactant molecules, and then the synthesis path is deduced.
  • the model trained in this way can improve the prediction ability of reactant molecules, break through data bottlenecks, and have stronger generalization capabilities.
  • embodiments of the present application can use a flow-based generation model to complete molecules.
  • the flow-based generation model is a non-autoregressive generation model and can be generated at one time. Compared with the autoregressive model, the generation efficiency is higher and the inference speed is faster. , can achieve similar or higher prediction accuracy, and the flow-based generation model can give the likelihood function value of the prediction result, which has better interpretability.
  • the training method of the molecular completion model trains the molecular completion model based on the sample compound molecules and the sample molecules to be completed. Since the sample molecule to be completed is masked by substructures in the sample compound molecule Obtained, that is to say, the data based on the training process of the molecular completion model is the data obtained on the basis of the sample compound molecules themselves.
  • This training process is a self-supervised training process based on the sample compound molecules. This kind of self-supervised training process is based on the sample compound molecules.
  • the supervised training process does not need to pay attention to whether the sample compound molecules are compounds in known synthesis reactions. Therefore, this self-supervised training process is not limited by known synthesis reactions.
  • the molecular completion model trained by this training process is It has strong generalization ability, which is conducive to expanding the adaptability to scenarios, thereby helping to improve the prediction reliability and accuracy of reactant molecules.
  • an embodiment of the present application provides a device for predicting reactant molecules, which is provided in a computer device.
  • the device includes:
  • the first acquisition unit 1001 is used to acquire product molecules, break bonds on the product molecules, and obtain molecules to be completed.
  • the product molecules refer to any compound molecule of the reactant molecules to be predicted;
  • the completion unit 1002 is used to call the molecular completion model to complete the molecules to be completed, obtain the completion results, and determine the reactant molecules of the product molecules based on the completion results;
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed, and the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules.
  • the completion unit 1002 is used to determine the target atomic feature latent variable based on the atomic feature information of the molecule to be completed, and the target atomic feature latent variable is the atomic feature latent variable of the completed molecule; based on The chemical bond connection information of the molecule to be completed is used to determine the target chemical bond connection hidden variable.
  • the target chemical bond connection hidden variable is the chemical bond connection hidden variable of the completed molecule; the molecule completion model is called to transform the target chemical bond connection hidden variable to obtain the target chemical bond Connect the information, transform the target atom feature hidden variables to obtain the target atom feature information; determine the completion result of the molecule to be completed based on the target chemical bond connection information and the target atom feature information.
  • the first acquisition unit 1001 is used to obtain the graph structure information of the product molecule; based on the graph structure information, predict the rupture probability of the chemical bond in the product molecule, and determine the chemical bond whose rupture probability meets the reference condition as the product molecule
  • the broken chemical bonds in the product are broken; the product molecules are broken based on the broken chemical bonds to obtain the molecules to be completed.
  • the prediction process of reactant molecules relies on a molecular completion model.
  • the molecular completion model is trained based on sample compound molecules and sample molecules to be completed. Since the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules, that is to say, the training process of the molecule completion model is based on data obtained based on the sample compound molecules themselves.
  • the training process is a self-supervised training process based on sample compound molecules. This self-supervised training process does not need to pay attention to whether the sample compound molecules are compounds in known synthesis reactions. Therefore, this self-supervised training process is not affected by known compounds. Due to the limitations of the synthetic reaction, the molecular completion model trained using this training process has strong generalization ability, which is conducive to expanding the adaptation scenarios, thereby helping to improve the prediction reliability and accuracy of the reactant molecules.
  • an embodiment of the present application provides a training device for a molecular completion model.
  • the device includes:
  • the second acquisition unit 1101 is used to acquire sample compound molecules and sample molecules to be completed.
  • the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules;
  • the third acquisition unit 1102 is used to determine the training loss based on the sample compound molecules, the sample molecules to be completed, and the molecule completion model;
  • the update unit 1103 is used to update the model parameters of the molecule completion model based on the training loss to obtain the trained molecule completion model.
  • the third acquisition unit 1102 is used to acquire the sample atomic feature information and sample chemical bond connection information of the sample compound molecules; determine the atomic mask based on the difference between the sample compound molecules and the sample molecules to be completed. information and chemical bond mask information; call the molecular completion model to inversely transform the sample chemical bond connection information based on the chemical bond mask information to obtain the sample chemical bond connection hidden variables, and inversely transform the sample atom feature information based on the atom mask information to obtain the sample atoms Feature latent variable; determine the training loss based on the sample chemical bond connection latent vector and the sample atomic feature latent vector.
  • the third acquisition unit 1102 is used to call the molecular completion model based on the chemical bond mask information. information and sample chemical bond connection information to determine the first chemical bond connection information and the second chemical bond connection information.
  • the first chemical bond connection information is the chemical bond connection information of the sample molecule to be completed, and the second chemical bond connection information is the chemical bond connection information of the substructure; based on The first chemical bond connection information performs an inverse transformation on the second chemical bond connection information to obtain the chemical bond connection hidden variables of the substructure; based on the first chemical bond connection information and the chemical bond connection hidden variables of the substructure, the sample chemical bond connection hidden variables are determined.
  • the third acquisition unit 1102 is used to determine the first atomic feature information and the second atomic feature information based on the atomic mask information and the sample atomic feature information, and the first atomic feature information is the sample to be completed.
  • the atomic feature information of the molecule, the second atomic feature information is the atomic feature information of the substructure; perform inverse transformation on the second atomic feature information based on the first atomic feature information to obtain the atomic feature hidden variable of the substructure; based on the first atomic feature information and the atomic characteristic latent variables of the substructure to determine the sample atomic characteristic latent variables.
  • the third acquisition unit 1102 is used to call the molecule completion model to complete the sample molecules to be completed, obtain the completion results, and determine the predicted completion molecules based on the completion results; based on the prediction completion The difference between the molecule and the sample compound molecule determines the training loss.
  • the substructure is a structure in the sample compound molecule that belongs to the candidate structure set
  • the candidate structure set is a set of structures whose credibility meets the selection conditions.
  • the training device for a molecular completion model trains a molecular completion model based on sample compound molecules and sample molecules to be completed. Since the sample molecules to be completed are obtained by masking the substructures in the sample compound molecules, that is to say, the training process of the molecule completion model is based on data obtained based on the sample compound molecules themselves.
  • the training process is a self-supervised training process based on sample compound molecules. This self-supervised training process does not need to pay attention to whether the sample compound molecules are compounds in known synthesis reactions. Therefore, this self-supervised training process is not affected by known compounds. Due to the limitations of the synthetic reaction, the target molecule completion model trained using this training process has strong generalization ability, which is conducive to expanding the adaptation scenarios, thereby helping to improve the prediction reliability and accuracy of the reactant molecules.
  • a computer device in an exemplary embodiment, includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors, so that the computer device implements any of the above-mentioned prediction methods of reactant molecules or training methods of molecule completion models.
  • the computer device can be a server or a terminal. Next, the structures of the server and terminal are introduced respectively.
  • FIG 12 is a schematic structural diagram of a server provided by an embodiment of the present application.
  • the server may vary greatly due to different configurations or performance, and may include one or more processors (Central Processing Units, CPUs) 1201 and one or A plurality of memories 1202, wherein at least one computer program is stored in the one or more memories 1202, and the at least one computer program is loaded and executed by the one or more processors 1201, so that the server implements the above method embodiments Provides prediction methods for reactant molecules or training methods for molecular completion models.
  • the server can also have components such as wired or wireless network interfaces, keyboards, and input and output interfaces to facilitate input and output.
  • the server can also include other components for implementing device functions, which will not be described again here.
  • Figure 13 is a schematic structural diagram of a terminal provided by an embodiment of the present application.
  • the terminal can be: PC, mobile phone, smartphone, PDA, wearable device, PPC, tablet computer, smart car machine, smart TV, smart speaker, smart voice interaction device, smart home appliances, car terminal, VR device, AR device.
  • the terminal may also be called user equipment, portable terminal, laptop terminal, desktop terminal, and other names.
  • the terminal includes: a processor 1301 and a memory 1302.
  • the processor 1301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc.
  • the processor 1301 can adopt DSP (Digital Signal Processing, digital signal processing), FPGA (Field-Programmable Gate Array, field programmable gate array), PLA (Programmable Logic Array, programmable logic array). One less form of hardware to implement.
  • Memory 1302 may include one or more computer-readable storage media, which may be non-transitory. Memory 1302 may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1301 to enable the terminal to implement the method embodiments of the present application. Provides prediction methods for reactant molecules or training methods for molecular completion models.
  • the terminal optionally further includes: a peripheral device interface 1303 and at least one peripheral device.
  • the processor 1301, the memory 1302 and the peripheral device interface 1303 may be connected through a bus or a signal line.
  • Each peripheral device can be connected to the peripheral device interface 1303 through a bus, a signal line, or a circuit board.
  • the peripheral device includes: at least one of a radio frequency circuit 1304 or a display screen 1305.
  • the peripheral device interface 1303 may be used to connect at least one I/O (Input/Output) related peripheral device to the processor 1301 and the memory 1302 .
  • the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one of the processor 1301, the memory 1302, and the peripheral device interface 1303 or Both of them can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
  • the radio frequency circuit 1304 is used to receive and transmit RF (Radio Frequency, radio frequency) signals, also called electromagnetic signals. Radio frequency circuit 1304 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals.
  • RF Radio Frequency, radio frequency
  • the display screen 1305 is used to display UI (User Interface, user interface).
  • the UI can include graphics, text, icons, videos, and any combination thereof.
  • display screen 1305 also has the ability to collect touch signals on or above the surface of display screen 1305 .
  • the touch signal can be input to the processor 1301 as a control signal for processing.
  • the display screen 1305 can also be used to provide virtual buttons and/or virtual keyboards, also called soft buttons and/or soft keyboards.
  • Figure 13 does not constitute a limitation of the terminal, and may include more or fewer components than shown, or combine certain components, or adopt different component arrangements.
  • a computer-readable storage medium is also provided. At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor of the computer device, so that the computer Implement any of the above methods for predicting reactant molecules or training methods for molecular completion models.
  • the above computer-readable storage medium can be read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), read-only compact disc (Compact Disc Read-Only Memory) , CD-ROM), tapes, floppy disks and optical data storage devices, etc.
  • ROM Read-Only Memory
  • RAM Random Access Memory
  • CD-ROM Compact Disc Read-Only Memory
  • tapes floppy disks and optical data storage devices, etc.
  • a computer program product is also provided.
  • the computer program product includes a computer program or computer instructions.
  • the computer program or computer instructions are loaded and executed by the processor to enable the computer to implement any of the above reactants. Methods for predicting molecules or methods for training molecular completion models.
  • the information including but not limited to user equipment information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • signals involved in this application All are authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
  • the product molecules involved in this application were obtained with full authorization.

Landscapes

  • Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Chemical & Material Sciences (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Computing Systems (AREA)
  • Theoretical Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Evolutionary Computation (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Artificial Intelligence (AREA)
  • Medicinal Chemistry (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种反应物分子的预测、模型的训练方法、装置、设备及介质,属于人工智能技术领域。该反应物分子的预测方法包括:计算机设备获取产物分子(201);计算机设备对产物分子进行断键,得到待补全分子(202);计算机设备调用分子补全模型对待补全分子进行补全,得到补全结果,基于补全结果确定产物分子的反应物分子(203),分子补全模型基于样本化合物分子以及样本待补全分子训练得到,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到。反应物分子的预测过程依赖泛化能力较强的分子补全模型实现,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。

Description

反应物分子的预测、模型的训练方法、装置、设备及介质
本申请要求于2022年07月14日提交、申请号为202210830979.X、发明名称为“反应物分子的预测、模型的训练方法、装置、设备及介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请实施例涉及人工智能技术领域,特别涉及一种反应物分子的预测、模型的训练方法、装置、设备及介质。
背景技术
随着人工智能技术的兴起和快速发展,给定一个产物分子,预测其反应物分子的应用场景越来越广泛,如,化学合成场景、药品制备场景等。
发明内容
本申请实施例提供了一种反应物分子的预测、模型的训练方法、装置、设备及介质,可用于提高反应物分子的预测可靠性和准确性。所述技术方案如下:
一方面,本申请实施例提供了一种反应物分子的预测方法,所述方法包括:
计算机设备获取产物分子,对所述产物分子进行断键,得到待补全分子,所述产物分子是指待预测反应物分子的任一化合物分子;
所述计算机设备调用分子补全模型对所述待补全分子进行补全,得到补全结果,基于所述补全结果确定所述产物分子的反应物分子;
其中,所述分子补全模型基于样本化合物分子以及样本待补全分子训练得到,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到。
还提供了一种分子补全模型的训练方法,所述方法包括:
计算机设备获取样本化合物分子以及样本待补全分子,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到;
所述计算机设备基于所述样本化合物分子、所述样本待补全分子和分子补全模型,确定训练损失;
所述计算机设备基于所述训练损失更新所述分子补全模型的模型参数,得到训练后的分子补全模型。
另一方面,提供了一种反应物分子的预测装置,设置于计算机设备中,所述装置包括:
第一获取单元,用于获取产物分子,对所述产物分子进行断键,得到待补全分子,所述产物分子是指待预测反应物分子的任一化合物分子;
补全单元,用于调用分子补全模型对所述待补全分子进行补全,得到补全结果,基于所述补全结果获取所述产物分子的反应物分子;
其中,所述分子补全模型基于样本化合物分子以及样本待补全分子训练得到,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到。
还提供了一种分子补全模型的训练装置,设置于计算机设备中,所述装置包括:
第二获取单元,用于获取样本化合物分子以及样本待补全分子,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到;
第三获取单元,用于基于所述样本化合物分子、所述样本待补全分子和分子补全模型, 获取训练损失;
更新单元,用于基于所述训练损失更新所述分子补全模型的模型参数,得到训练后的分子补全模型。
另一方面,提供了一种计算机设备,所述计算机设备包括处理器和存储器,所述存储器中存储有至少一条计算机程序,所述至少一条计算机程序由所述处理器加载并执行,以使所述计算机设备实现上述任一所述的反应物分子的预测方法或分子补全模型的训练方法。
另一方面,还提供了一种计算机可读存储介质,所述计算机可读存储介质中存储有至少一条计算机程序,所述至少一条计算机程序由处理器加载并执行,以使计算机设备实现上述任一所述的反应物分子的预测方法或分子补全模型的训练方法。
另一方面,还提供了一种计算机程序产品,所述计算机程序产品包括计算机程序或计算机指令,所述计算机程序或所述计算机指令由处理器加载并执行,以使计算机设备实现上述任一所述的反应物分子的预测方法或分子补全模型的训练方法。
本申请实施例提供的技术方案至少带来如下有益效果:
本申请实施例提供的技术方案,反应物分子的预测过程依赖分子补全模型实现,分子补全模型是基于样本化合物分子和样本待补全分子训练得到的。由于样本待补全分子通过对样本化合物分子中的子结构进行掩码得到,也就是说,分子补全模型的训练过程所依据的数据是在样本化合物分子本身的基础上得到的数据,此种训练过程为一种基于样本化合物分子的自监督训练过程,此种自监督训练过程无需关注样本化合物分子是否为已知的合成反应中的化合物,因而此种自监督训练过程并不会受已知的合成反应的限制,利用该训练过程训练得到的分子补全模型的泛化能力较强,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
附图说明
图1是本申请实施例提供的一种实施环境的示意图;
图2是本申请实施例提供的一种反应物分子的预测方法的流程图;
图3是本申请实施例提供的一种产物分子的表示形式的示意图;
图4是本申请实施例提供的一种反应物分子预测过程的两个阶段的示意图;
图5是本申请实施例提供的一种分子补全模型的训练方法的流程图;
图6是本申请实施例提供的一种样本化合物分子中的被掩码的子结构的三种情况的示意图;
图7是本申请实施例提供的一种对样本化合物分子中不同的化学键进行切割的示意图;
图8是本申请实施例提供的一种初始化学键补全模型的结构示意图;
图9是本申请实施例提供的一种初始原子补全模型的结构示意图;
图10是本申请实施例提供的一种反应物分子的预测装置的示意图;
图11是本申请实施例提供的一种分子补全模型的训练装置的示意图;
图12是本申请实施例提供的一种服务器的结构示意图;
图13是本申请实施例提供的一种终端的结构示意图。
具体实施方式
相关技术在预测反应物分子时,先根据产物分子获取待补全分子,然后从多个候选结构中确定与待补全分子匹配的结构,将该匹配的结构与待补全分子连接,将连接后得到的分子作为预测的反应物分子。其中,多个候选结构通过对已知的合成反应中的产物分子和反应物分子之间的差异结构进行比对得到。
上述反应物分子的预测方法依赖从已知的合成反应中提取的多个候选结构,多个候选结构的泛化能力受限于已知的合成反应,泛化能力较差,适应的场景较为局限,从而容易降低反应物分子的预测可靠性和预测准确性。
在示例性实施例中,本申请实施例提供的反应物分子的预测方法以及分子补全模型的训练方法可应用于各种场景,包括但不限于云技术、人工智能、智慧交通、辅助驾驶等。
人工智能(Artificial Intelligence,AI)是利用数字计算机或者数字计算机控制的机器模拟、延伸和扩展人的智能,感知环境、获取知识并使用知识获得最佳结果的理论、方法、技术及应用系统。换句话说,人工智能是计算机科学的一个综合技术,人工智能企图了解智能的实质,并生产出一种新的能以人类智能相似的方式做出反应的智能机器。人工智能也就是研究各种智能机器的设计原理与实现方法,使机器具有感知、推理与决策的功能。
人工智能技术是一门综合学科,涉及领域广泛,既有硬件层面的技术也有软件层面的技术。人工智能基础技术一般包括如传感器、专用人工智能芯片、云计算、分布式存储、大数据处理技术、操作/交互系统、机电一体化等技术。人工智能软件技术主要包括计算机视觉技术、语音处理技术、自然语言处理技术以及机器学习/深度学习、自动驾驶、智慧交通等几大方向。
本申请实施例提供的方案涉及人工智能技术中的机器学习技术,机器学习(Machine Learning,ML)是一门多领域交叉学科,涉及概率论、统计学、逼近论、凸分析、算法复杂度理论等多门学科。专门研究计算机怎样模拟或实现人类的学习行为,以获取新的知识或技能,重新组织已有的知识结构使之不断改善自身的性能。机器学习是人工智能的核心,是使计算机具有智能的根本途径,其应用遍及人工智能的各个领域。机器学习和深度学习通常包括人工神经网络、置信网络、强化学习、迁移学习、归纳学习、示教学习等技术。
随着人工智能技术研究和进步,人工智能技术在多个领域展开研究和应用,例如常见的智能家居、智能穿戴设备、虚拟助理、智能音箱、智能营销、无人驾驶、自动驾驶、无人机、机器人、智能医疗、智能客服、车联网、自动驾驶、智慧交通等,相信随着技术的发展,人工智能技术将在更多的领域得到应用,并发挥越来越重要的价值。
图1示出了本申请实施例提供的一种实施环境的示意图。该实施环境可以包括:终端11和服务器12。
本申请实施例提供的反应物分子的预测方法可以由终端11执行,也可以由服务器12执行,还可以由终端11和服务器12共同执行,本申请实施例对此不加以限定。对于本申请实施例提供的反应物分子的预测方法由终端11和服务器12共同执行的情况,服务器12承担主要计算工作,终端11承担次要计算工作;或者,服务器12承担次要计算工作,终端11承担主要计算工作;或者,服务器12和终端11二者之间采用分布式计算架构进行协同计算。
本申请实施例提供的分子补全模型的训练方法可以由终端11执行,也可以由服务器12执行,还可以由终端11和服务器12共同执行,本申请实施例对此不加以限定。对于本申请实施例提供的分子补全模型的训练方法由终端11和服务器12共同执行的情况,服务器12承担主要计算工作,终端11承担次要计算工作;或者,服务器12承担次要计算工作,终端11承担主要计算工作;或者,服务器12和终端11二者之间采用分布式计算架构进行协同计算。
需要说明的是,反应物分子的预测方法的执行设备与分子补全模型的训练方法的执行设备可以相同,也可以不同,本申请实施例对此不加以限定。
可选地,终端11可以是任何一种可与用户通过键盘、触摸板、触摸屏、遥控器、语音交互或手写设备等一种或多种方式进行人机交互的电子产品,例如PC(Personal Computer,个人计算机)、手机、智能手机、PDA(Personal Digital Assistant,个人数字助手)、可穿戴设备、PPC(Pocket PC,掌上电脑)、平板电脑、智能车机、智能电视、智能音箱、智能语音交互设备、智能家电、车载终端、VR(Virtual Reality,虚拟现实)设备、AR(Augmented Reality, 增强现实)设备等。服务器12可以是一台服务器,也可以是由多台服务器组成的服务器集群,或者是一个云计算服务中心。终端11与服务器12通过有线或无线网络建立通信连接。
本领域技术人员应能理解上述终端11和服务器12仅为举例,其他现有的或今后可能出现的终端或服务器如可适用于本申请,也应包含在本申请保护范围以内,并在此以引用方式包含于此。
本申请实施例提供的反应物分子的预测方法,用于根据给定的产物分子,预测生成该产物分子的反应物分子,该预测任务可称为逆合成预测任务,逆合成预测任务对于化学领域以及制药领域等均有着极为重要的意义。传统的逆合成预测任务大多基于合成反应模板实现,例如,首先通过匹配算法从合成反应模板中找到与产物分子匹配的模板,再根据匹配到的模板得到反应物分子。这类方法在逆合成预测任务上取得了一定的效果,但基于合成反应模板的方法有两个比较明显的缺陷:第一个缺陷是基于合成反应模板的方法很难泛化到新的反应类型上,导致合成反应模板需要被频繁更新,而总结模板需要化学专家大量的工作,成本很高;第二个缺陷是合成反应模板只总结了部分分子级别的反应规律,无法抓住全局的正确信息,常常会导致错误的预测。
随着深度学习技术的兴起,为了克服基于合成反应模板的方法的缺陷,深度学习模型被广泛应用在逆合成预测任务中,利用深度学习技术能够直接预测产物分子的反应物分子,无需与合成反应模板进行匹配。通过深度学习技术能够实现强有力的逆合成预测效果,强有力的逆合成预测效果能够帮助化学专家发现产物分子的可能的合成路径,大大提高新化合物的研发效率,如,产物分子可以为药物分子,则可以大大提高制药产业新药的研发效率。此外,强有力的逆合成预测效果还可以揭示一些隐藏的科学规律,提供新的科学知识,发现新的合成路径乃至新的合成反应。本申请实施例提供的反应物分子的预测方法即为一种基于深度学习技术实现逆合成预测任务的方法。
示例性地,逆合成预测任务可以表示为(GP→GR),GP=(VP,BP,AP)表示一系列产物分子,GR=(VR,BR,AR)表示一系列反应物分子。VP代表产物分子GP中的原子的集合且集合的大小表示产物分子GP中的原子的个数N(N为不小于1的整数),也即|VP|=N。
BP∈RN×N×C表示产物分子GP的化学键连接信息,该化学键连接信息表示产物分子GP中的原子之间的化学键连接情况,该化学键连接信息是一个三维的矩阵,C(C为不小于1的整数)表示原子之间可能存在的化学键的类型的数量,BP中的元素[i,j,c]的取值表示产物分子GP中的原子i和原子j之间是否通过类型为c的化学键连接,若[i,j,c]的取值为1,则表示产物分子GP中的原子i和原子j之间通过类型为c的化学键连接;若[i,j,c]的取值为0,则表示产物分子GP中的原子i和原子j之间未通过类型为c的化学键连接。示例性地,化学键连接信息还可以称为邻接矩阵。
AP∈RN×F表示产物分子GP的原子特征信息,该原子特征信息包括产物分子GP中的每个原子的子特征信息,每个原子都具有维度为F(F为不小于1的整数)的子特征信息,原子的子特征信息的获取方式将在下文中介绍,此处暂不赘述。VR,BR,AR的含义可以参见VP,BP,AP的含义,此处不再加以赘述。
当前的逆合成预测任务通常关注单产物分子以及单步逆合成(多步逆合成可由单步逆合成结果组合得到),即产物分子的数量为一个,也即|GP|=1,而反应物分子的数量为T(T为不小于1的整数)个,即|GR|=T,T≥1。合成反应遵循atom-mapping(原子匹配)原则,即产物分子中的原子与反应物分子中的原子一一对应,故产物分子和反应物分子共享同一个原子集合。由于逆合成预测任务在产物分子一端没有提供副产物分子,所以产物分子中的原子的数量通常小于反应物分子中的原子的数量,也即|VR|≥|VP|。
本申请实施例提供一种反应物分子的预测方法,该方法可应用于上述图1所示的实施环境。该反应物分子的预测方法由计算机设备执行,该计算机设备可以为终端11,也可以为服 务器12,本申请实施例对此不加以限定。如图2所示,本申请实施例提供的反应物分子的预测方法可以包括如下步骤201至步骤203。
在步骤201中,计算机设备获取产物分子。
产物分子是指待预测反应物分子的任一化合物分子,通过预测产物分子的反应物分子,能够推导该产物分子的合成路径,从而为产物分子的研发提供数据支持。本申请实施例对产物分子的类型不加以限定,例如,产物分子可以为药物分子、衣物分子、食品分子等。产物分子包括多个原子,多个原子之间通过化学键连接,产物分子包括的原子的类型和数量,以及多个原子之间的化学键连接情况与产物分子有关,本申请实施例对不加以限定。示例性地,产物分子还可以称为生成物分子。
本申请实施例对产物分子的获取方式不加以限定,示例性地,可以从化合物分子数据库中提取产物分子,也可以从期刊或文章中公开的化合物分子中选取产物分子,还可以将技术人员上传的化合物分子作为产物分子等。
本申请实施例对产物分子的表示形式不加以限定,只要能够指示出产物分子中的原子的情况以及原子之间的化学键连接情况即可。示例性地,产物分子的表示形式可以为名称,也可以为分子式,还可以为字符串等。示例性地,产物分子的名称可以根据化合物命名规则对产物分子进行命名得到。产物分子的分子式为产物分子的组成结构的最为直观的表示方式,根据产物分子的分子式能够直观确定产物分子。
产物分子的字符串为按照一定的规范生成的字符串,能够对产物分子进行较为精简的表示,示例性地,生成产物分子的字符串所依据的规范可以是指SMILES(Simplified Molecular Input Line entry Specification,简化分子线性输入规范)。示例性地,对于生成产物分子的字符串所依据的规范为SMILES的情况,产物分子的字符串还可以称为产物分子的SMILES表达式。SMILES是一种用ASCII(American Standard Code for Information Interchange,美国信息交换标准代码)字符串明确描述分子结构的规范,每个化合物分子具有唯一的SMILES表达式。
示例性地,对于同一产物分子,可以利用图3中的(1)所示的分子式表示,也可以利用图3中的(2)所示的SMILES表达式表示。
在步骤202中,计算机设备对产物分子进行断键,得到待补全分子。
待补全分子是通过对产物分子进行断键处理得到的分子,在一些实施例中,待补全分子还可以称为合成子(synthon)。待补全分子可视为反应物分子在合成产物分子的过程中去掉一些分子结构后得到的分子,在确定至少一个待补全分子后,即可进一步通过对待补全分子进行补全,来预测反应物分子。示例性地,反应物分子去掉的分子结构可以称为离去基团(leaving group)。
需要说明的是,基于产物分子获取的待补全分子的数量可能为一个,也可能为多个,这与产物分子的实际情况有关,本申请实施例对此不加以限定。
在一种可能实现方式中,计算机设备产物分子进行断键,得到待补全分子的实现过程包括以下步骤2021至步骤2023。
步骤2021:计算机设备获取产物分子的图结构信息。
产物分子的图结构信息为用于表征产物分子的图结构的信息,产物分子的图结构可以是将产物分子转换得到的唯一确定的图结构。示例性地,将产物分子中的每个原子视为一个节点,将产物分子中的每个化学键视为一条边,从该角度出发,将产物分子转换为图结构。也就是说,产物分子的图结构为以产物分子中的原子为节点,以产物分子中的化学键为边构建得到的图结构。
本申请实施例对产物分子的图结构信息的类型不加以限定,只要能够表征产物分子的图结构即可。示例性地,产物分子的图结构信息包括产物分子的原子特征信息和产物分子的化学键连接信息。示例性地,产物分子的原子特征信息用于对产物分子中的原子的特征进行表征。产物分子的化学键连接信息用于对产物分子中的原子之间的化学键连接情况进行表征。
由于产物分子的图结构以产物分子中的原子为节点,以产物分子中的化学键为边,所以产物分子的原子特征信息可视为表征产物分子的图结构中的节点的信息,产物分子的化学键连接信息可视为表征产物分子中的图结构中的边的信息,因此,能够利用产物分子的原子特征信息和产物分子的化学键连接信息来对产物分子的图结构进行表征。接下来,分别介绍产物分子的原子特征信息的获取方式以及产物分子的化学键连接信息的获取方式。
产物分子的原子特征信息包括产物分子中的每个原子的子特征信息,原子的子特征信息用于对该原子的特征进行表征,本申请实施例对原子的子特征信息的表示方式不加以限定,例如,原子的子特征信息的表示形式可以为矩阵,也可以为向量等。需要说明的是,不同原子的子特征信息的维度相同,该相同的维度可以根据经验设置,也可以根据应用场景灵活调整,本申请实施例对此不加以限定。示例性地,对于原子的子特征信息的表示形式为矩阵的情况,产物分子的原子特征信息还可以称为产物分子的原子特征矩阵,该原子特征矩阵中的每行元素均表示一个原子的子特征信息。
获取每个原子的子特征信息的原理相同,本申请实施例以获取原子的子特征信息的方式为例进行说明。示例性地,获取原子的子特征信息的方式包括:获取原子的属性信息;对该原子的属性信息进行特征提取,得到该原子的子特征信息。原子的属性信息用于描述原子的属性,原子的属性信息根据经验设置,或者根据应用场景灵活调整,示例性地,原子的属性信息包括但不限于原子的元素信息、价态信息、度信息、是否属于苯环的信息中至少一种。
元素信息包括原子在元素周期表的排行、元素的符号表示、相对原子质量中至少一种但不限于此。例如,碳元素在元素周期表中排第6,碳元素的符号表示为C,碳元素的相对原子质量为12.01。价态信息是指原子在产物分子中的价态,价态又称化合价或者原子价,价态是各种元素的一个原子或原子团、基(根)与其他原子相互化合的数目。原子在不同化合物中的价态可能相同,也可能不相同。例如,在CO(一氧化碳)中碳的价态为+2价,而在CO2(二氧化碳)中碳的价态为+4价。度信息包括连接了该原子的其他原子的数量。例如CO2,碳原子与两个氧原子相连接,两个氧原子均分别与碳原子相连接。那么碳原子的度信息可以为2。是否属于苯环的信息指示原子是否为构成苯环的原子。
在获取原子的属性信息后,对该原子的属性信息进行特征提取,将提取得到的信息作为该原子的子特征信息。示例性地,对该原子的属性信息进行特征提取的方式可以根据经验设置,例如,调用原子特征提取模型对该原子的属性信息进行特征提取。示例性地,原子特征提取模型可以基于样本原子的属性信息以及样本原子的特征标签,通过监督训练的方式训练得到。
产物分子的化学键连接信息基于产物分子中的原子之间的化学键连接情况确定。示例性地,产物分子的化学键连接信息还可以称为产物分子的图结构的邻接矩阵。示例性地,产物分子的化学键连接信息为一个N*N*C维的矩阵,其中,N和C均为不小于1的整数,N表示产物分子中的原子的数量,C表示候选化学键类型的数量。候选化学键类型根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定,示例性地,候选化学键类型涵盖合成反应中常见的化学键类型,例如,候选化学键类型包括但不限于单键、双键、三键、芳香键、离子键、共价键和金属键等。
产物分子的化学键连接信息中位于第i行第j列第a深度的元素[i,j,c]的取值表示原子i和原子j之间是否通过类型为a的化学键连接,类型为a的化学键是指C个候选化学键类型中的第a个候选化学键类型。其中,i和j均为1~N中的任一取值,a为1~C中的任一取值。若[i,j,a]的取值为1,则表示原子i和原子j之间通过类型为a的化学键连接;若[i,j,a]的取值为0,则表示原子i和原子j之间未通过类型为a的化学键连接。产物分子的化学键连接信息能够通过分析产物分子中的原子之间的化学键连接情况得到,化学键连接情况是指是否通过化学键连接,以及在通过化学键连接时,通过哪种化学键连接等。
在示例性实施例中,产物分子的图结构信息还包括产物分子的化学键连接信息,通过额外考虑化学键连接信息,能够为后续的化学键的断裂概率的预测过程提供更多的数据支持, 从而提高化学键的断裂概率的预测准确性。示例性地,化学键连接信息包括产物分子中每个化学键的子特征信息。
示例性地,化学键的子特征信息的获取方式包括:获取化学键的属性信息,对化学键的属性信息进行特征提取,得到化学键的子特征信息。示例性地,化学键的属性信息用于描述化学键的属性,化学键的属性信息可以根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定。示例性地,化学键的属性信息包括但不限于化学键的键类型、共轭特征、环键特征、键能、键合距离中的至少一种。
键类型表示化学键所属的类型,如单键、双键、三键、芳香键、离子键、共价键和金属键等。共轭特征表示化学键是否共轭。环键特征表示化学键是否为环键的一部分。键能是从能量因素衡量化学键强弱的物理量。一般来说,键能越大,化学键越牢固,化学键越不容易断裂。键合距离是指两个或以上的原子核之间形成化学键所必需的最短距离。
在获取化学键的属性信息后,对该化学键的属性信息进行特征提取,将提取到的信息作为该化学键的子特征信息。示例性地,对化学键的属性信息进行特征提取的方式可以根据经验设置,例如,调用化学键特征提取模型对化学键的属性信息进行特征提取。示例性地,化学键特征提取模型可以基于样本化学键的属性信息以及样本化学键的特征标签,通过监督训练的方式训练得到。
步骤2022:计算机设备基于图结构信息,预测产物分子中的化学键的断裂概率,将断裂概率满足参考条件的化学键确定为产物分子中的断裂化学键。
化学键的断裂概率指示该化学键为在合成反应中形成的化学键的可能性,化学键的断裂概率与该化学键为在合成反应中形成的化学键的可能性呈正相关关系,也即化学键的断裂概率越大,说明该化学键为在合成反应中形成的化学键的可能性越大。示例性地,化学键为在合成反应中形成的化学键的可能性越大,根据该化学键对产物分子进行断键的可靠程度越大。需要说明的是,产物分子中的化学键的断裂概率是指产物分子中的各个化学键的断裂概率。
化学键的断裂概率可以基于图结构信息预测得到。产物分子的图结构信息能够指示出产物分子中的化学键的存在情况,例如存在哪些化学键,以及每个化学键连接的原子的情况等,根据图结构信息能够预测出产物分子中的各个化学键的断裂概率。在示例性实施例中,基于图结构信息,预测产物分子中的化学键的断裂概率的过程可以通过运行预先编写的程序实现,也可以通过调用图神经网络模型实现。
本申请实施例以调用图神经网络模型基于图结构信息,预测产物分子中的化学键的断裂概率为例进行说明。图神经网络模型为能够对化合物分子的图结构信息进行处理,以预测化合物分子中的化学键的断裂概率的模型,也即能够分辨产物分子中的哪些化学键更容易断裂的模型。本申请实施例对图神经网络模型的模型结构不加以限定,示例性地,图神经网络模型可以是任一种基于图的深度学习网络模型,可以设计的简单,也可以设计的复杂。例如,图神经网络模型可以是指图卷积网络(Graph Convolutional Networks,GCN)模型、图注意力网络(Graph Attention Networks,GAT)模型、信息传递神经网络(Message Passing Neural Network,MPNN)模型等。
调用图神经网络模型基于图结构信息,预测产物分子中的化学键的断裂概率的过程为图神经网络模型的内部处理过程,与图神经网络模型的模型结构有关,本申请实施例对此不加以限定。示例性地,调用图神经网络模型基于图结构信息,预测产物分子中的化学键的断裂概率的过程包括:调用图神经网络模型基于图结构信息提取产物分子中的化学键的目标特征;基于产物分子中的化学键的目标特征预测产物分子中的化学键的断裂概率。化学键的目标特征为预测化学键的断裂概率所依据的特征。
示例性地,化学键的断裂概率可以为1或0,此种情况下,预测化学键的断裂概率的过程可视为对化学键进行二分类预测的过程,断裂概率为1表示化学键极有可能为在合成反应中形成的化学键,断裂概率为0表示化学键极小可能为合成反应中形成的化学键。当然,在示例性实施例中,化学键的断裂概率还可能为0~1之间的任一概率,本申请实施例对此不加 以限定。
在确定产物分子中的化学键的断裂概率后,能够从产物分子中的化学键中确定出断裂概率满足参考条件的化学键,将断裂概率满足参考条件的化学键确定为产物分子中的断裂化学键。断裂化学键即为从产物分子中获取待补全分子所依据的化学键。示例性地,断裂化学键还可以称为反应位点。
断裂概率满足参考条件的化学键是指在合成反应中形成的可能性较大的化学键,也即更容易断裂的化学键,根据此种化学键对产物分子进行断键处理的可靠性较高,从而提高获取的待补全分子的可靠性,进而提高预测的反应物分子的可靠性。
断裂概率满足参考条件根据经验设置,获取根据应用场景灵活调整,本申请实施例对此不加以限定。在示例性实施例中,断裂概率满足参考条件是指断裂概率不小于概率阈值,概率阈值根据经验设置,或者根据应用场景灵活调整,例如,概率阈值为0.5,或者,概率阈值为0.8等。在示例性实施例中,断裂概率满足参考条件还可以是指断裂概率为各个化学键的断裂概率中前L(L为不小于1的整数)大的断裂概率,L的取值根据经验设置,或者根据应用场景灵活调整,例如,L的取值为3,或者L的取值为2等。
在示例性实施例中,在调用图神经网络模型基于图结构信息,预测产物分子中的化学键的断裂概率之前,需要先训练图神经网络模型。在示例性实施例中,训练图神经网络模型的过程包括:获取训练化合物分子的图结构信息以及训练化合物分子中的化学键的标准断裂概率;调用图神经网络模型基于训练化合分子的图结构信息,预测训练化合物分子中的化学键的训练断裂概率;基于标准断裂概率和训练断裂概率之间的差异,确定参考损失;基于参考损失更新图神经网络模型的模型参数;响应于训练过程满足第一终止条件,将当前训练得到的模型确定为训练完成的图神经网络模型。
满足第一终止条件根据经验设置,或根据应用场景灵活调整,例如,满足第一终止条件是指参考损失收敛、参考损失小于第一损失阈值、模型参数的更新次数达到第一次数阈值等。第一损失阈值和第一次数阈值根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定。
训练化合物分子是指能够获知图结构信息以及化学键的标准断裂概率的化合物分子,训练化合物分子的数量可能为一个,也可能为多个,本申请实施例对此不加以限定。示例性地,训练化合物分子的图结构信息的获取原理与产物分子的图结构信息的获取原理相同,此处不再加以赘述。示例性地,训练化合物分子的图结构信息可以与训练化合物分子对应存储在数据库中,从而能够从数据库中直接提取训练化合物分子的图结构信息。
训练化合物分子中的化学键的标准断裂概率为训练化合物分子中的化学键的真实的断裂概率,用于为图神经网络模型的训练过程提供监督信息。示例性地,训练化合物分子中的化学键的标准断裂概率也可称为训练化合物分子中的化学键的断裂概率真值(ground-truth)标签。在示例性实施例中,训练化合物分子中的化学键的标准断裂概率与训练化合物分子对应存储在数据库中,从而能够从数据库中直接提取训练化合物分子中的化学键的标准断裂概率。
在示例性实施例中,训练化合物分子为一种通过已知的合成反应合成的分子,此种情况下,训练化合物分子中的化学键的标准断裂概率可以通过对训练化合物分子和合成该训练化合物分子的反应物分子进行比对得到。示例性地,通过对训练化合物分子和合成该训练化合物分子的反应物分子进行比对得到训练化合物分子中的化学键的标准断裂概率的过程可以为:若产物分子中的原子u和原子v之间通过类型为a的化学键buv相连(可以表示为BP[u,v,a]=1,对于某个a),而在合成该训练化合物分子的反应物分子中,原子u和原子v之间不存在任何类型的化学键连接(可以表示为BR[u,v,a]=0,对于任意a),则将标产物分子中的原子u和原子v之间的化学键buv的标准断裂概率记为1(可以表示为yuv=1);若在合成该训练化合物分子的反应物分子中,原子u和原子v之间存在某一类型的化学键连接,则将标产物分子中的原子u和原子v之间的化学键buv的标准断裂概率记为0(可以表示为yuv=0)。
调用图神经网络模型基于训练化合分子的图结构信息,预测训练化合物分子中的化学键 的训练断裂概率的实现原理与调用图神经网络模型基于产物分子的图结构信息,预测产物分子中的化学键的断裂概率的实现原理相同,此处不再加以赘述。
在获取训练化合物分子中的化学键的训练断裂概率之后,基于标准断裂概率和训练断裂概率之间的差异,确定参考损失。本申请实施例对标准断裂概率和训练断裂概率之间的差异的衡量方式不加以限定,示例性地,标准断裂概率和训练断裂概率之间的差异是指标准断裂概率和训练断裂概率之间的交叉熵差异,或者标准断裂概率和训练断裂概率之间的差异是指标准断裂概率和训练断裂概率之间的均方差异等。
示例性地,以标准断裂概率和训练断裂概率之间的差异是指标准断裂概率和训练断裂概率之间的交叉熵差异为例,参考损失可以基于公式1计算得到:
其中,L表示参考损失;K(K为不小于1的整数)表示训练化合物分子的数量;表示第k(k从1依次取值到K)个训练化合物分子的化学键连接信息;表示根据确定的第k个训练化合物分子中的化学键,通过综合考虑各个训练化合分子中的各个化学键来获取参考损失;yuv表示化学键buv的标准断裂概率;表示化学键buv的训练断裂概率。
步骤2023:计算机设备基于断裂化学键对产物分子进行断键,得到待补全分子。
示例性地,基于断裂化学键对产物分子进行断键是指断开产物分子中的断裂化学键,将断开产物分子中的断裂化学键后得到的各个分子作为各个待补全分子。需要说明的是,对产物分子进行断键可能得到一个待补全分子,也可能得到多个待补全分子,本申请实施例对此不加以限定。
示例性地,待补全分子表示为其中,H(H为不小于1的整数)表示待补全分子的数量,表示第h(h为1到H中的任一取值)个待补全分子。示例性地,在产物分子(表示为GP)的基础上获取待补全分子的过程可视为建模概率分布的过程。
示例性地,从产物分子中确定待补全分子的过程可以称为反应位点预测过程,该反应位点预测过程可视为反应物分子预测过程中的第一阶段。示例性地,该第一阶段可以如图4中的(1)所示,产物分子中的断裂化学键的数量为一个,断开该一个断裂化学键后,能够得到两个待补全分子。需要说明的是,图4中的(1)中的两个待补全分子上的虚线圆圈标记的是断裂化学键连接的两个原子。
需要说明的是,基于上述步骤2021至步骤2023对产物分子进行断键的过程仅为一种示例性实现过程,本申请实施例并不局限于此。在一些实施例中,还可以根据经验从产物分子的化学键中选择断裂化学键,然后基于选择的断裂化学键对产物分子进行断键。
在步骤203中,计算机设备调用分子补全模型对待补全分子进行补全,得到补全结果,基于补全结果确定产物分子的反应物分子,其中,分子补全模型基于样本化合物分子以及样本待补全分子训练得到,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到。
示例性地,对样本化合物分子中的子结构进行掩码是指隐藏样本化合物分子中的子结构,在对样本化合物分子中的子结构进行掩码后,无法得知样本化合物分子中的子结构的原本状态。对样本化合物分子中的子结构进行掩码的方式可以根据经验设置,也可以根据实际的应用场景灵活调整,只要能够实现对样本化合物分子中的子结构的隐藏即可。示例性地,对样本化合物分子中的子结构进行掩码的方式可以为利用特定结构替换样本化合物分子中的子结构,也可以为对样本化合分子中的子结构进行遮盖等。示例性地,特定结构是指与真实的化合物分子的结构不同的结构,以便于区分被掩码的部分以及未被掩码的部分。
在待补全分子的基础上预测反应物分子的过程通过调用分子补全模型实现。分子补全模型是基于样本化合物分子和通过对样本化合物分子中的子结构进行掩码得到的样本待补全分子训练得到的。样本待补全分子是在样本化合物分子本身的基础上获取的,基于样本化合物分子和样本待补全分子训练得到分子补全模型的过程可视为将样本化合物分子进行掩码操作,训练模型去重构被掩码的部分的过程。此种训练过程为采用自监督学习策略训练模型的过程,在采用自监督学习策略训练模型的过程中,自监督学习任务是在“Mask and Fill(掩码和填充)” 的思想下设计的任务,基本思路是将化合物分子的部分子结构进行掩码操作,然后训练一个分子补全模型去重构该子结构。
在采用自监督学习策略进行训练的过程中,模型训练所依据的数据为在样本化合物分子本身的基础上得到的数据,无论样本化合物是否为已知的合成反应中的化合物,均能够用作模型训练所依据的数据,也就是说,模型的训练过程并不会受已知的合成反应的限制,利用此种训练过程训练得到的分子补全模型的泛化能力较强,有利于有效适应与各种合成反应相关的反应物分子预测场景,扩展了适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
在本申请实施例中,假设待补全分子的数量与反应物分子的数量相同,也即一个待补全分子对应一个反应物分子。调用分子补全模型对待补全分子进行补全是指调用分子补全模型对每个待补全分子分别进行补全,在补全后,每个待补全分子均对应一个补全结果,根据每个待补全分子的补全结果均能确定产物分子的一个反应物分子。
调用分子补全模型对每个待补全分子进行补全的原理相同,本申请实施例以调用分子补全模型对待补全分子进行补全的过程为例进行说明。需要说明的是,对于待补全分子的数量为多个的情况,可以调用分子补全模型对多个待补全分子进行并行补全,以提高预测反应物分子的效率。
调用分子补全模型对待补全分子进行补全,能够得到待补全分子的补全结果。本申请实施例对待补全分子的补全结果的形式不加以限定,只要能够根据待补全分子的补全结果确定出反应物分子即可。该补全结果指示补全后的分子。
示例性地,待补全分子的补全结果可以为待补全分子的补全后的分子的图结构信息,此种情况下,可以根据图结构信息确定待补全分子的补全后的分子,将待补全分子的补全后的分子作为一个反应物分子。
示例性地,待补全分子的补全结果还可以为待补全分子中的缺失结构的指示信息,该指示信息包括用于表征待补全分子中的缺失结构的信息以及用于指示待补全分子中待与缺失结构连接的位置的信息,此种情况下,可以根据指示信息确定缺失结构,然后将缺失结构与该待补全分子进行连接,将连接后得到的分子作为一个反应物分子。示例性地,用于表征待补全分子中的缺失结构的信息可以为待补全分子中的缺失结构的图结构信息、分子式、分子字符串等,本申请实施例对此不加以限定。
示例性地,对于待补全分子中的缺失结构为单原子或单化学键级别的结构的情况,待补全分子的补全结果还可以为待补全分子中的缺失结构的分类结果以及连接位置预测结果,该分类结果包括待补全分子中的缺失结构与参考结构(如,参考原子或参考化学键)的匹配概率,连接位置预测结果指示待补全分子中待与缺失结构连接的位置。此种情况下,可以根据分类结果,确定待补全分子中的缺失结构的类别(也即单原子或单化学键的类别),然后在待补全分子中待与缺失结构连接的位置处将该类别的结构与待补全分子连接,将连接后得到的分子作为一个反应物分子。
调用分子补全模型对待补全分子进行补全,得到待补全分子的补全结果的过程为分子补全模型的内部处理过程,与分子补全模型的结构有关,本申请实施例对此不加以限定。
在示例性实施例中,分子补全模型为一种基于流的生成模型(Flow-based Generative Models),该基于流的生成模型为一种可逆模型。接下来,对基于流的生成模型进行介绍。
基于流的生成模型直接最大化似然估计(Maximum Likelihood Estimation)。同时,基于流的生成模型可以给出生成结果的似然估计,也即基于流的生成模型具备更强的解释性。基于流的生成模型的大致思路是,通过一连串的可逆变换(invertible transformations)将复杂的数据概率分布变成常见的简单分布(也可以称为隐变量分布)。假设数据来自于分布X~PX(X),隐变量分布为Z~PZ(Z)(一般是高斯分布),通过一连串可逆函数变换 使得Z=fθ(X),整个流程是L为不小于1的整数。基于流的生成模型的目标是最大化对数似然估计对数似然估计logpθ(x)的表达式可 以结合基于流的生成模型的可逆映射推导得到,如公式2所示:
其中,det(·)表示矩阵的行列式计算,在高维情况下是一个n×n的矩阵,被称为雅可比矩阵,n为z的维度,n为不小于1的整数;pθ(z)表示z在参数θ下的概率;pθ(x)表示x在参数θ下的概率。
基于流的生成模型通常采用耦合层(coupling layer)的网络层设计方案,以兼顾计算效率和模型表示能力,耦合层的输入x与输出z之间的关系如公式3和公式4所示:
z1:d=x1:d    (公式3)
其中,公式3表示将输入x的前d(d为不小于1且不大于n的整数)维进行复制,公式4表示将输入x的剩余维度(也即从第d+1维到第n维)进行变换。公式4中的Sθ(·)和Tθ(·)表示两个变换函数,这两个变换函数均用于输出与xd+1:n维度相同的变换信息。示例性地,Sθ(·)表示尺度(scale)函数,Tθ(·)表示转换(transformation)函数。⊙表示矩阵对应位置元素相乘。这样的设计使得能够通过简单的变换上述公式3和公式4进行耦合层的逆运算,从而实现从z到x的逆变换,逆运算基于下述公式5和公式6实现:
x1:d=z1:d   (公式5)
此耦合层的设计使得雅克比矩阵变成对角矩阵,对角矩阵的行列式计算是对角线元素的乘积,即其中,j表示Sθ(z1:d)中的任一元素。此种情况下,尺度函数和转换函数可以是任意复杂的神经网络而不会增加雅克比矩阵行列式的计算量,从而提高计算效率。
在示例性实施例中,对于分子补全模型为一种基于流的生成模型的情况,调用分子补全模型对待补全分子进行补全,得到待补全分子的补全结果的过程包括以下步骤2031至步骤2034。
步骤2031:计算机设备基于待补全分子的原子特征信息,确定目标原子特征隐变量,该目标原子特征隐变量为补全后的分子的原子特征隐变量,基于待补全分子的化学键连接信息,确定目标化学键连接隐变量,该目标化学键连接隐变量为补全后的分子的化学键连接隐变量。
待补全分子的原子特征信息用于表征待补全分子中的原子的特征,目标原子特征隐变量用于对待补全分子的补全后的分子中原子的特征进行假设。获取待补全分子的原子特征信息的原理与获取产物分子的原子特征信息的原理相同,此处不再加以赘述。
在示例性实施例中,基于待补全分子的原子特征信息,获取目标原子特征隐变量的过程包括:从已知概率分布中采样待补全分子的缺失结构的原子特征隐变量;基于待补全分子的原子特征信息和缺失结构的原子特征隐变量,确定目标原子特征隐变量。
待补全分子的缺失结构是指待补全分子待补全的结构,缺失结构的原子特征隐变量用于对缺失结构的中的原子的特征进行假设。缺失结构的原子特征隐变量是从已知概率分布中采样得到的,也就是说,缺失结构的原子特征隐变量为服从已知概率分布的变量。已知概率分布为能够确定服从该概率分布的变量的概率的任一分布,已知概率分布的类型可以根据经验设置,本申请实施例对此不加以限定,例如,已知概率分布可以是指高斯分布,也可以是指均匀分布等。示例性地,采样缺失结构的原子特征隐变量的方式可以为随机采样。
在示例性实施例中,缺失结构的原子特征隐变量和待补全分子的原子特征信息的形式均为矩阵,在缺失结构的原子特征隐变量的矩阵和待补全分子的原子特征信息的矩阵中,均是一行元素对应一个原子。示例性地,基于待补全分子的原子特征信息和缺失结构的原子特征隐变量,确定目标原子特征隐变量的方式可以为:对待补全分子的原子特征信息和缺失结构 的原子特征隐变量进行纵向拼接,基于拼接得到的矩阵,获取目标原子特征隐变量。
示例性地,基于拼接得到的矩阵,获取目标原子特征隐变量的方式可以为:将拼接得到的矩阵作为目标原子特征隐变量。
示例性地,基于拼接得到的矩阵,获取目标原子特征隐变量的方式还可以为:若拼接得到的矩阵的维度为第一参考维度,将拼接得到的矩阵作为目标原子特征隐变量;若拼接得到的矩阵的维度小于第一参考维度,将拼接得到的矩阵扩充为第一参考维度的矩阵,将扩充后得到的矩阵作为目标原子特征隐变量。此种方式能够保证目标原子特征隐变量的维度为第一参考维度,从而提高目标原子特征隐变量的规范性。
第一参考维度为预先设置的用于约束原子特征方面的信息的维度的参数,第一参考维度可认为是最大反应物分子的原子特征信息的维度,也即本申请实施例认为第一参考维度不小于拼接得到的矩阵的维度。示例性地,将拼接得到的矩阵扩充为第一参考维度的矩阵的过程可以是指:将拼接得到的矩阵置于左上角位置,在其他位置添加0元素,直至得到第一参考维度的矩阵。
待补全分子的化学键连接信息用于表征待补全分子中的原子之间的化学键连接情况,目标化学键连接隐变量用于对待补全分子的补全后的分子中的原子之间的化学键连接情况进行假设。获取待补全分子的化学键连接信息的原理与获取产物分子的化学键连接信息的原理相同,此处不再加以赘述。
在示例性实施例中,基于待补全分子的化学键连接信息,确定目标化学键连接隐变量的过程包括:从已知概率分布中采样待补全分子的缺失结构的化学键连接隐变量;基于待补全分子的化学键连接信息和缺失结构的化学键连接隐变量,获取目标化学键连接隐变量。
缺失结构的化学键连接隐变量用于对缺失结构的中的原子之间的化学键连接情况进行假设。缺失结构的化学键连接隐变量是从已知概率分布中采样得到的,也就是说,缺失结构的化学键连接隐变量为服从已知概率分布的变量。示例性地,采样缺失结构的化学键连接隐变量的方式可以为随机采样。
示例性地,缺失结构的化学键连接隐变量能够指示出缺失结构中的原子之间的化学键连接情况(假设的情况),待补全分子的化学键连接信息能够指示出待补全分子中的原子之间的化学键连接情况(真实的情况),基于缺失结构的化学键连接隐变量和待补全分子的化学键连接信息,能够获取一个用于指示缺失结构中原子以及待补全分子中的原子(也即补全后的分子中的原子)之间的化学键连接情况(假设的情况)的目标信息,基于该目标信息获取目标化学键连接隐变量。示例性地,目标信息和目标化学键连接隐变量的形式均为矩阵。
需要说明的是,补全后的分子中的原子之间的化学键连接情况除包括缺失结构中的原子之间的化学键连接情况以及待补全分子中的原子之间的化学键连接情况外,还包括缺失结果中的原子与补全分子中的原子之间的化学键连接情况。缺失结果中的原子与补全分子中的原子之间的化学键连接情况可以根据经验设置,如,默认缺失结果中的原子与补全分子中的原子之间不存在化学键连接。
示例性地,基于目标信息获取目标化学键连接隐变量的方式可以为:将目标信息作为目标化学键连接隐变量。
示例性地,基于目标信息获取目标化学键连接隐变量的方式还可以为:若标信息的维度为第二参考维度,将标信息作为目标化学键连接隐变量;若目标信息的维度小于第二参考维度,将目标信息扩充为第二参考维度的矩阵,将扩充后得到的矩阵作为目标化学键连接隐变量。此种方式能够保证目标化学键连接隐变量的维度为第二参考维度,从而提高目标化学键连接隐变量的规范性。
第二参考维度为预先设置的用于约束化学键连接方面的信息的维度的参数,第二参考维度可认为是最大反应物分子的化学键连接信息的维度,也即本申请实施例认为第二参考维度不小于目标信息的维度。示例性地,将目标信息扩充为第二参考维度的矩阵的过程可以是指:将目标信息置于左上角位置,在其他位置添加0元素,直至得到第二参考维度的矩阵。
步骤2032:计算机设备调用分子补全模型对目标化学键连接隐变量进行变换,得到目标化学键连接信息。
目标化学键连接信息用于表征待补全分子的补全后的分子中的原子之间的化学键连接情况,对目标化学键连接隐变量进行变换,得到目标化学键连接信息的过程是指根据待补全分子的补全后的分子中的原子之间的化学键连接情况的假设信息,预测待补全分子的补全后的分子中的原子之间的化学键连接情况的表征信息的过程。
示例性地,调用分子补全模型对目标化学键连接隐变量进行变换,得到目标化学键连接信息的实现过程包括:调用分子补全模型基于目标化学键连接隐变量,获取待补全分子的参考化学键连接信息和缺失结构的参考化学键连接隐变量;基于待补全分子的参考化学键连接信息对缺失结构的参考化学键连接隐变量进行变换,得到缺失结构的参考化学键连接信息;基于待补全分子的参考化学键连接信息和缺失结构的参考化学键连接信息,获取目标化学键连接信息。
在示例性实施例中,目标化学键连接隐变量为第二参考维度的矩阵,基于目标化学键连接隐变量获取待补全分子的参考化学键连接信息的方式为:将目标化学键连接隐变量中用于指示待补全分子中的原子之间的化学键连接情况的信息保持不变,将其他信息置为0,得到待补全分子的参考化学键连接信息。此种方式下获取的待补全分子的参考化学键连接信息同样为第二参考维度的矩阵。
在示例性实施例中,基于目标化学键连接隐变量获取缺失结构的参考化学键连接隐变量的方式为:将目标化学键连接隐变量中用于指示缺失结构中的原子之间的化学键连接情况的信息保持不变,将其他信息置为0,得到缺失结构的参考化学键连接隐变量。此种方式下获取的缺失结构的参考化学键连接隐变量同样为第二参考维度的矩阵。
示例性地,基于待补全分子的参考化学键连接信息对缺失结构的参考化学键连接隐变量进行变换,得到缺失结构的参考化学键连接信息的实现方式可以为:基于待补全分子的参考化学键连接信息,获取第一参考变换信息;利用第一参考变换信息对缺失结构的参考化学键连接隐变量进行变换,得到缺失结构的参考化学键连接信息。示例性地,第一参考变换信息可以基于至少一个变换函数(如,公式6中涉及的Sθ(·)变换函数和Tθ(·)变换函数)获取。
示例性地,利用第一参考变换信息对缺失结构的参考化学键连接隐变量进行变换,得到缺失结构的参考化学键连接信息的实现过程可以利用公式6表示,其中,利用xd+1:n表示缺失结构的参考化学键连接信息,利用Tθ(z1:d)和Sθ(z1:d)表示基于z1:d获取的第一参考变换信息,利用z1:d表示待补全分子的参考化学键连接信息,利用zd+1:n表示缺失结构的参考化学键连接隐变量。需要说明的是,变换不会改变信息的维度,也即缺失结构的参考化学键连接隐变量的维度与缺失结构的参考化学键连接信息的维度相同。
示例性地,待补全分子的参考化学键连接信息和缺失结构的参考化学键连接信息均为第二参考维度的矩阵,基于待补全分子的参考化学键连接信息和缺失结构的参考化学键连接信息,获取目标化学键连接信息的方式可以为:将待补全分子的参考化学键连接信息和缺失结构的参考化学键连接信息中的对应位置的元素相加,将相加后得到的矩阵作为目标化学键连接信息。示例性地,还可以将待补全分子的参考化学键连接信息和缺失结构的参考化学键连接信息之间的笛卡尔乘积作为目标化学键连接信息。
示例性地,分子补全模型包括化学键补全模型,该步骤2032可以通过调用分子补全模型中的化学键补全模型实现,也即调用分子补全模型中的化学键补全模型对目标化学键连接隐变量进行变换,得到目标化学键连接信息。
化学键补全模型能够将分子的化学键连接隐变量变换成该分子的化学键连接信息。其中,分子的化学键连接隐变量用于对该分子中的原子之间的化学键连接情况进行假设,分子的化学键连接信息用于表征该分子中的原子之间的化学键连接情况,也就是说,目标化学键补全模型的输入为分子中的原子之间的化学键连接情况的假设信息,输出为分子中的原子之间的化学键连接情况的表征信息。
示例性地,化学键补全模型为一种可逆模型,也就是说,化学键补全模型存在逆模型。化学键补全模型的逆模型能够将分子的化学键连接信息逆变换成该分子的化学键连接隐变量。也就是说,化学键补全模型的逆模型的输入为分子中的原子之间的化学键连接情况的表征信息,输出为分子中的原子之间的化学键连接情况的假设信息。
化学键补全模型是通过训练得到的,在训练过程中化学键补全模型的模型结构不变,化学键补全模型的模型结构可以参见图5所示的实施例中介绍的模型结构,此处暂不赘述。
步骤2033:计算机设备调用分子补全模型对目标原子特征隐变量进行变换,得到目标原子特征信息。
目标原子特征信息用于表征待补全分子的补全后的分子中的原子的特征,对目标原子特征隐变量进行变换,得到目标原子特征信息的过程是指根据待补全分子的补全后的分子中的原子的特征的假设信息,预测待补全分子的补全后的分子中的原子的特征的表征信息的过程。
示例性地,调用分子补全模型对目标原子特征隐变量进行变换,得到目标原子特征信息的实现过程包括:调用分子补全模型基于目标原子特征隐变量,获取待补全分子的参考原子特征信息和缺失结构的参考原子特征隐变量;基于待补全分子的参考原子特征信息对缺失结构的参考原子特征隐变量进行变换,得到缺失结构的参考原子特征信息;基于待补全分子的参考原子特征信息和缺失结构的参考原子特征信息,获取目标原子特征信息。
在示例性实施例中,目标原子特征隐变量为第一参考维度的矩阵,基于目标原子特征隐变量获取待补全分子的参考原子特征信息的方式为:将目标原子特征隐变量中用于指示待补全分子中的原子的特征的信息保持不变,将其他信息置为0,得到待补全分子的参考原子特征信息。此种方式下获取的待补全分子的参考原子特征信息同样为第一参考维度的矩阵。
在示例性实施例中,基于目标原子特征隐变量获取缺失结构的参考原子特征隐变量的方式为:将目标原子特征隐变量中用于指示缺失结构中的原子的特征的信息保持不变,将其他信息置为0,得到缺失结构的参考原子特征隐变量。此种方式下获取的缺失结构的参考原子特征隐变量同样为第一参考维度的矩阵。
示例性地,基于待补全分子的参考原子特征信息对缺失结构的参考原子特征隐变量进行变换,得到缺失结构的参考原子特征信息的实现方式可以为:基于待补全分子的参考原子特征信息,获取第二参考变换信息;利用第二参考变换信息对缺失结构的参考原子特征隐变量进行变换,得到缺失结构的参考原子特征信息。示例性地,第二参考变换信息可以基于至少一个变换函数(如,公式6中涉及的Sθ(·)变换函数和Tθ(·)变换函数)获取。
示例性地,利用第二参考变换信息对缺失结构的参考原子特征隐变量进行变换,得到缺失结构的参考原子特征信息的实现过程可以利用公式6表示,其中,利用xd+1:n表示缺失结构的参考原子特征信息,利用Tθ(z1:d)和Sθ(z1:d)表示基于z1:d获取的第二参考变换信息,利用z1:d表示待补全分子的参考原子特征信息,利用zd+1:n表示缺失结构的参考原子特征隐变量。需要说明的是,变换不会改变信息的维度,也即缺失结构的参考原子特征隐变量的维度与缺失结构的参考原子特征的维度相同。
示例性地,待补全分子的参考原子特征信息和缺失结构的参考原子特征信息均为第一参考维度的矩阵,基于待补全分子的参考原子特征信息和缺失结构的参考原子特征信息,获取目标原子特征信息的方式可以为:将待补全分子的参考原子特征信息和缺失结构的参考原子特征信息中的对应位置的元素相加,将相加后得到的矩阵作为目标原子特征信息。示例性地,还可以将待补全分子的参考原子特征信息和缺失结构的参考原子特征信息之间的笛卡尔乘积作为目标原子特征信息。
在示例性实施例中,对目标原子特征隐变量进行变换的过程需要考虑目标化学键连接信息的约束,以保证变换过程的可靠性。此种情况下,需要基于目标化学键连接信息,获取目标原子特征隐变量的约束信息,然后调用目标原子补全模型在约束信息的约束下对目标原子特征隐变量进行变换。
示例性地,在约束信息的约束下对目标原子特征隐变量进行变换的差异性过程体现在基 于待补全分子的参考原子特征信息对缺失结构的参考原子特征隐变量进行变换,得到缺失结构的参考原子特征信息的过程中,也即,对缺失结构的参考原子特征隐变量进行变换所依据的除了包括待补全分子的参考原子特征信息外,还包括约束信息。
本申请实施例对获取目标原子特征隐变量的约束信息的方式不加以限定,可以根据经验设置,也可以根据应用场景灵活调整。示例性地,获取目标原子特征隐变量的约束信息的方式可以为将目标化学键连接信息作为目标原子特征隐变量的约束信息。示例性地,获取目标原子特征隐变量的约束信息的方式还可以为对目标化学键连接信息进行标准化处理,得到目标原子特征隐变量的约束信息。对目标化学键连接信息进行标准化处理用于提高目标化学键连接信息的规范性,标准化处理的方式可以根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定。例如,标准化处理可以调用图标准化(Graph Normalization,简称Graphnorm)模块实现。
示例性地,分子补全模型包括原子补全模型,该步骤2033可以通过调用分子补全模型中的原子补全模型实现,也即调用分子补全模型中的原子补全模型对目标原子特征隐变量进行变换,得到目标原子特征信息。
原子补全模型能够将分子的原子特征隐变量变换成该分子的原子特征信息。其中,分子的原子特征隐变量用于对该分子中的原子的特征进行假设,分子的原子特征信息用于表征该分子中的原子的特征,也就是说,原子补全模型的输入为分子中的原子的特征的假设信息,输出为分子中的原子的特征的表征信息。
示例性地,原子补全模型为一种可逆模型,也就是说,原子补全模型存在逆模型。原子补全模型的逆模型能够将分子的原子特征信息逆变换成该分子的原子特征隐变量。也就是说,原子补全模型的逆模型的输入为分子中的原子的特征的表征信息,输出为分子中的原子的特征的假设信息。
原子补全模型是通过训练得到的,在训练过程中原子补全模型的模型结构不变,原子补全模型的模型结构可参见图5所示的实施例中的原子补全模型的模型结构,此处暂不赘述。
步骤2034:计算机设备基于目标化学键连接信息和目标原子特征信息,确定待补全分子的补全结果。
待补全分子的补全结果为能够确定待补全分子的补全后的分子的结果,目标化学键连接信息能够指示出待补全分子的补全后的分子中的原子之间的化学键连接情况,目标原子特征信息能够指示出待补全分子的补全后的分子中的原子的特征,根据补全后的分子中的原子之间的化学键连接情况以及补全后的分子中的原子的特征即可唯一确定一个补全后的分子。因此,能够基于目标化学键连接信息和目标原子特征信息,获取待补全分子的补全结果。
示例性地,基于目标化学键连接信息和目标原子特征信息,获取待补全分子的补全结果的方式可以为:将包括目标化学键连接信息和目标原子特征信息的信息作为待补全分子的补全结果。
示例性地,基于目标化学键连接信息和目标原子特征信息,获取待补全分子的补全结果的方式还可以为:基于目标化学键连接信息和目标原子特征信息,确定待补全分子的分子式或者图结构,将分子式或者图结构作为待补全分子的补全结果。
示例性地,基于步骤2031至步骤2034获取待补全分子的补全结果的过程可以表示为:
输入:目标原子补全模型的逆模型和目标化学键补全模型的逆模型其中,AS和BS为待补全分子Gs的原子特征信息和化学键连接信息,为待补全分子中的缺失结构的原子特征隐变量和化学键连接隐变量;为补全后的分子的目标原子特征隐变量;为补全后的分子的目标化学键连接隐变量;
1、//获取待处理信息该待处理信息包括补全后的分子的目标原子特征隐变量和目标化学键连接隐变量
2、//调用目标化学键补全模型对目标化学键连接隐变量进行变换,得到目标化学键连接信息BR
3、//基于Graphnorm模块对目标化学键连接信息BR进行标准化处理,得到约束信息
4、//调用目标原子补全模型在约束信息的约束下对目标原子特征隐变量进行变换,得到目标原子特征信息AR
输出:补全结果GR=(AR,BR)。
示例性地,基于流的生成模型的输入是待补全分子GS=(VS,BS,AS),输出是补全后的分子(也即反应物分子)GR=(VR,BR,AR)。VS、BS和AS分别表示待补全分子的原子集合、化学键连接信息和原子特征信息,VR、BR和AR分别表示反应物分子的原子集合、化学键连接信息和原子特征信息。示例性地,基于流的生成模型为一种非自回归的生成模型,能够一次性生成补全结果,相比于自回归式的生成模型,生成速度更快,用于提高反应物分子预测的效率。
需要说明的是,上述步骤2031至步骤2034仅以分子补全模型为一种基于流的生成模型为例,介绍了调用分子补全模型对待补全分子进行补全的实现过程,本申请实施例并不局限于此。在一些实施例中,分子补全模型也可以为其他类型的生成模型,如,变分自编码器(VAE)生成模型和生成对抗模型(GAN)等,本申请实施例对此不加以限定。当然,在一些实施例中,分子补全模型还可以卷积神经网络模型等。在分子补全模型的不同情况下,调用分子补全模型对待补充分子进行补全,得到该待补全分子的补全结果的过程也有所不同,本申请实施例在此不再一一介绍。
示例性地,调用分子补全模型对待补充分子进行补全,得到该待补全分子的补全结果的过程可以为:调用分子补全模型提取待补全分子的特征;基于待补全分子的特征预测待补全分子的补全后的分子的特征;基于待补全分子的补全后的分子的特征,获取该待补全分子的补全结果。
示例性地,对至少一个待补全分子进行补全的过程可视为反应物分子预测过程中的第二阶段。在该第二阶段中,在第一阶段得到的待补全分子的基础上,添加原子和化学键,将待补全分子补全(或者称为还原)为原来的反应物分子。这一阶段的操作可认为是在建模概率分布可以当成条件生成问题来解决。其中,GR表示反应物分子。如前文描述,本申请实施例假设了一个待补全分子对应一个反应物分子,则需要建立的概率分布可以表示为其中,表示的反应物分子,表示第h(h为不小于1的整数)个待补全分子。示例性地,该第二阶段如图4中的(2)所示,调用分子补全模型分别对两个待补全分子进行补全,得到产物分子的两个反应物分子。
本申请实施例提供的反应物分子的预测方法,反应物分子的预测过程依赖分子补全模型实现,分子补全模型是基于样本化合物分子和样本待补全分子训练得到的。由于样本待补全分子通过对样本化合物分子中的子结构进行掩码得到,也就是说,分子补全模型的训练过程所依据的数据是在样本化合物分子本身的基础上得到的数据,此种训练过程为一种基于样本化合物分子的自监督训练过程,此种自监督训练过程无需关注样本化合物分子是否为已知的合成反应中的化合物,因而此种自监督训练过程并不会受已知的合成反应的限制,利用该训练过程训练得到的分子补全模型的泛化能力较强,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
本申请实施例提供了一种分子补全模型的训练方法,该方法可应用于上述图1所示的实施环境。该分子补全模型的训练方法由计算机设备执行,该计算机设备可以为终端11,也可以为服务器12,本申请实施例对此不加以限定。如图5所示,本申请实施例提供的分子补全模型的训练方法包括如下步骤501和步骤502。
在步骤501中,计算机设备获取样本化合物分子以及样本待补全分子,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到。
样本化合物分子为对分子补全模型训练一次所依据的化合物分子,样本化合物分子的数 量可能为一个,也可能为多个,本申请实施例对此不加以限定。示例性地,样本化合物分子可以从包括化合物分子的任一数据集中提取得到。示例性地,包括化合物分子的数据集可以为包括已知的合成反应的数据集,也可以为不包括已知的合成反应的数据集,本申请实施例对此不加以限定。也就是说,样本化合物分子的获取灵活性较高,不局限于包括已知的合成反应的数据集,从而有利于提高训练得到的模型的泛化能力。示例性地,由于模型的训练过程无需依据样本化合物的标签,所以,包括化合物的数据集可以为无标签的数据集。
在获取样本化合物分子后,可以获取样本化合物分子的样本待补全分子,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到。示例性地,对样本化合物分子中的子结构进行掩码是指隐藏样本化合物分子中的子结构,在对样本化合物分子中的子结构进行掩码后,无法得知样本化合物分子中的子结构的原本状态。对样本化合物分子中的子结构进行掩码的方式可以根据经验设置,也可以根据实际的应用场景灵活调整,只要能够实现对样本化合物分子中的子结构的隐藏即可。示例性地,对样本化合物分子中的子结构进行掩码的方式可以为利用特定结构替换样本化合物分子中的子结构,也可以为对样本化合分子中的子结构进行遮盖等。示例性地,特定结构是指与真实的化合物分子的结构不同的结构,以便于区分被掩码的部分以及未被掩码的部分。
样本化合物分子中的子结构是指样本化合物分子中的一部分,本申请实施例对样本化合物分子中的被掩码的子结构的复杂程度不加以限定,示例性地,样本化合物分子中的被掩码的子结构可以为样本化合物分子中的一个原子,也可以为样本化合物分子中的一个化学键,也可以为样本化合物分子中的一个由至少一个原子和至少一个化学键构成的结构等。
示例性地,样本化合物分子中的被掩码的子结构为样本化合物分子中的一个原子的情况可以如图6中的(1)所示,样本化合物分子中的被掩码的子结构为样本化合物分子中的一个化学键的情况可以如图6中的(2)所示,样本化合物分子中的被掩码的子结构为样本化合物分子中的一个由至少一个原子和至少一个化学键构成的结构的情况可以如图6中的(3)所示。在图6中,问号标记的以及被遮挡的部分即为被掩码的子结构。
在示例性实施例中,对于样本化合物分子中的被掩码的子结构为样本化合物分子中的一个原子或样本化合物分子中的一个化学键的情况,对分子补全模型进行训练的过程可视为基于单原子或单化学键的重构任务对分子补全模型进行训练的过程,基于单原子或单化学键的重构任务为较为简单的任务,有利于提高模型训练的收敛速度。示例性地,基于单原子或单化学键的重构任务对分子补全模型进行训练可视为多分类问题,该多分类问题用于预测被掩码的单原子或单化学键的类别。
在示例性实施例中,由至少一个原子和至少一个化学键构成的结构可以称为子图,对于样本化合物分子中的被掩码的子结构为样本化合物分子中的一个由至少一个原子和至少一个化学键构成的结构的情况,对样本化合物进行掩码的级别可称为子图级别,此种级别的掩码有利于提高模型的补全能力。
示例性地,样本化合物分子中被掩码的子结构的可以根据经验从样本化合物分子中选定。例如,可以从样本化合物分子中任选一个中心原子,从该中心原子进行g(g为不小于0的整数)跳,将g跳覆盖到的子结构作为样本化合物分子中被掩码的子结构。此种选定样本化合物分子中被掩码的子结构的方式较为简单。
示例性地,样本化合物分子中被掩码的子结构还可以通过参考候选结构集从样本化合物分子中选定。此种情况下,样本化合物分子中被掩码的子结构为样本化合物中的属于候选结构集的结构。属于候选结构集的结构是指构成候选结构集的一个结构。候选结构集为可信度满足选取条件的结构的集合。可信度满足选取条件的结构是指可信度较高的结构,通过参考可信度较高的结构从样本化合物分子中选定被掩码的子结构,有利于提高被掩码的子结构的合理性,避免整体结构的崩坏,进而提高模型的补全性能。
可信度满足选取条件根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定。在示例性实施例中,可信度满足选取条件的结构可以包括已知的合成反应中的产 物分子与反应物分子之间的差异结构。已知的合成反应可以从逆合成数据集中提取,逆合成数据集可以根据经验选定,例如,逆合成数据集为USPTO-50K数据集,该数据集中包含5万个逆合成反应,每个逆合成反应均为一个已知的合成反应。
在示例性实施例中,可信度满足选取条件的结构还可以包括motif(可以称为基序,是构成任一种特征序列的基本结构)、出现频率大于频率阈值的官能团以及根据BRICS(一种拆分分子的算法)对参考分子进行拆分得到的结构等。示例性地,出现频率可以是在某些文章或期刊中出现的频率,也可以是在逆合成数据集中出现的频率等。参考分子可以根据经验选定。
示例性地,候选结构集还可以称为子结构词典。子结构词典的构建方式可以灵活选定,只要保证内部的结构为可信度满足选取条件的结构即可。不同的构建方式下,子结构词典中的结构的平均大小以及平均出现频率等可能有所不同。
示例性地,对于样本化合物分子中的被掩码的子结构为样本化合物中的属于候选结构集的结构的情况,获取样本待补全分子的实现过程可以为:从样本化合物分子中任选一个化学键,切割该化学键,得到两个结构,将两个结构中的原子数量较小的结构与候选结构集进行匹配,若匹配成功(也即该原子数量较小的结构属于候选结构集),则确定该原子数量较小的结构的切割方案合理,将该原子数量较小的结构作为样本化合物分子中的被掩码的子结构,对该被掩码的子结构进行掩码,得到样本待补全分子。示例性地,若匹配失败(也即该原子数量较小的结构不属于候选结构集),则重新选取化学键进行切割。此种实现过程可以称为基于分子切割(Molecular Decomposition)获取样本待补全分子的过程。此种分子切割能够保证被掩码的子结构以及剩下的部分都保留有意义的结构,降低补全(或称为重构)难度。
示例性地,选取同一样本化合物分子中不同的化学键进行切割,得到的原子数量较小的结构有所不同,例如,如图7所示,若选取样本化合物分子中的化学键1进行切割,则得到的原子数量较小的结构如701所示;若选取样本化合物分子中的化学键2进行切割,则得到的原子数量较小的结构如702所示。
需要说明的是,对于同一样本化合物分子,在不同的掩码方案下,可以得到不同的待补全分子,该样本化合物分子和每个待补全分子均可以构成一个数据对(pair),本申请实施例中的模型训练过程是在数据对的基础上进行的。也就是说,本申请实施例中的样本待补全分子是指样本化合物分子的任一待补全分子。
在步骤502中,计算机设备基于样本化合物分子、样本待补全分子和分子补全模型,确定训练损失;基于训练损失更新分子补全模型的模型参数,得到训练后的分子补全模型。
训练损失用于为分子补全模型的模型参数的更新提供监督信息,基于样本化合物分子、样本待补全分子和分子补全模型获取训练损失的实现方式与分子补全模型的类型有关,本申请实施例对此不加以限定。
在一种可能实现方式中,计算机设备基于样本化合物分子、样本待补全分子和分子补全模型确定训练损失的实现方式包括以下步骤5021至步骤5025。
步骤5021:计算机设备获取样本化合物分子的样本原子特征信息和样本化学键连接信息。
样本化合物分子的样本原子特征信息用于对样本化合物分子中的原子的特征进行表征,样本化合物分子的样本化学键连接信息用于对样本化合物分子中的原子之间的化学键连接情况进行表征。获取样本化合物分子的样本原子特征信息和样本化学键连接信息的原理与图2所示的实施例中获取产物分子的原子特征信息和化学键连接信息的原理相同,此处不再加以赘述。在示例性实施例中,样本化合物分子的样本化学键连接信息为第一参考维度的矩阵,样本化合物分子的样本原子特征信息为第二参考维度的矩阵。
步骤5022:计算机设备基于样本化合物分子和样本待补全分子之间的差异,确定原子掩码信息和化学键掩码信息。
原子掩码信息指示样本化合物分子中的原子的掩码情况,例如哪些原子被掩码,哪些原子没有被掩码,化学键掩码信息指示样本化合物分子中的原子之间的化学键的掩码情况,例如哪些化学键被掩码,哪些化学键没有被掩码。由于样本待补全分子是通过对样本化合物中 的子结构进行掩码得到的,所以,通过比对样本化合物分子和样本待补全分子之间的差异结构,可以确定原子掩码信息和化学键掩码信息。
示例性地,原子掩码信息与样本原子特征信息的维度相同,如,均为第二参考维度的矩阵。原子掩码信息中的任一位置的元素的取值指示样本原子特征信息中位于相同位置的元素的掩码情况,如,若原子掩码信息中的任一位置的元素的取值为0,则指示样本原子特征信息中位于相同位置的元素的掩码情况为被掩码;若原子掩码信息中的任一位置的元素的取值为1,则指示样本原子特征信息中位于相同位置的元素的掩码情况为没有被掩码。
示例性地,化学键掩码信息与样本化学键连接信息的维度相同,如,均为第一参考维度的矩阵。化学键掩码信息中的任一位置的元素的取值指示样本化学键连接信息中位于相同位置的元素的掩码情况,如,若化学键掩码信息中的任一位置的元素的取值为0,则指示样本化学键连接信息中位于相同位置的元素的掩码情况为被掩码;若化学键掩码信息中的任一位置的元素的取值为1,则指示样本化学键连接信息中位于相同位置的元素的掩码情况为没有被掩码。
步骤5023:计算机设备调用分子补全模型,基于化学键掩码信息对样本化学键连接信息进行逆变换,得到样本化学键连接隐变量。
样本化学键连接隐变量用于对样本化合物分子中的原子之间的化学键连接情况进行假设。分子补全模型为一种可逆模型,既能够将分子的化学键连接隐变量变换为分子的化学键连接信息,也能够将分子的化学键连接信息逆变换为分子的化学键连接隐向量。由于在模型的训练过程中,样本化合物分子的样本化学键连接信息为已知的较为准确的信息,所以,调用分子补全模型实现对样本化学键连接信息的逆变换。
在一种可能实现方式中,计算机设备调用分子补全模型基于化学键掩码信息对样本化学键连接信息进行逆变换,得到样本化学键连接隐变量的过程包括以下步骤50231至步骤50233。
步骤50231:计算机设备调用分子补全模型基于化学键掩码信息和样本化学键连接信息,确定第一化学键连接信息和第二化学键连接信息,第一化学键连接信息为样本待补全分子的化学键连接信息,第二化学键连接信息为子结构的化学键连接信息。
第一化学键连接信息表征样本待补全分子中的原子之间的化学键连接情况,第二化学键连接信息表征被掩码的子结构中的原子之间的化学键连接情况。由于化学键掩码信息能够指示出样本化合物分子中的原子之间的化学键掩码情况,被掩码的化学键即为被掩码的子结构中的原子之间的化学键,没有被掩码的化学键即为样本待补全分子中的原子之间的化学键,所以基于化学键掩码信息和样本化合物分子的样本化学键连接信息,能够获取样本待补全分子的第一化学键连接信息以及子结构的第二化学键连接信息。
示例性地,基于化学键掩码信息和样本化学键连接信息,确定第一化学键连接信息以及第二化学键连接信息的方式可以为:将样本化学键连接信息中与化学键掩码信息指示的没有被掩码的化学键相关的信息保留,将其他信息置为0,得到第一化学键连接信息;将样本化学键连接信息中与化学键掩码信息指示的被掩码的化学键相关的信息保留,将其他信息置为0,得到第二化学键连接信息。此种方式下,第一化学键连接信息的维度和第二化学键连接信息的维度均与样本化学键连接信息的矩阵相同,如,均为第一参考维度的矩阵。
步骤50232:计算机设备基于第一化学键连接信息对第二化学键连接信息进行逆变换,得到子结构的化学键连接隐变量。
子结构的化学键连接隐变量用于对子结构中的原子之间的化学键连接情况进行假设,示例性地,子结构的化学键连接隐变量为一种服从已知概率分布的变量,已知概率分布根据经验设置,或者根据应用场景灵活调整,例如,已知概率分布为高斯分布,或者为均匀分布等。
在示例性实施例中,基于第一化学键连接信息对第二化学键连接信息进行逆变换,得到子结构的化学键连接隐变量的实现过程包括:基于第一化学键连接信息,获取第一样本变换信息;利用第一样本变换信息对第二化学键连接信息进行逆变换,得到子结构的化学键连接隐变量。示例性地,第一样本变换信息可以基于至少一个变换函数(如,Sθ(·)变换函数和Tθ(·) 变换函数)获取。需要说明的是,逆变换不会改变信息的维度,也即子结构的化学键连接隐变量的维度与第二化学键连接信息的维度相同。
步骤50233:计算机设备基于第一化学键连接信息和子结构的化学键连接隐变量,确定样本化学键连接隐变量。
示例性地,第一化学键连接信息和子结构的化学键连接隐变量均为第一参考维度的矩阵,基于第一化学键连接信息和子结构的化学键连接隐变量,获取样本化学键连接隐变量的方式可以为:将第一化学键连接信息和子结构的化学键连接隐变量中的对应位置的元素相加,将相加后得到的矩阵作为样本化学键连接隐变量。示例性地,还可以将第一化学键连接信息和子结构的化学键连接隐变量之间的笛卡尔乘积作为样本化学键连接隐变量。
示例性地,获取样本化学键连接隐变量的过程可以基于下述公式7和公式8实现:

其中,表示第一化学键连接信息;表示第二化学键连接信息;表示第一样本变换信息;表示基于第一样本变换信息对第二化学键连接信息进行逆变换的运算方式;表示子结构的化学键连接隐变量。为样本化学键连接隐变量的两个组成部分。示例性地,公式7用于将样本待补全分子的第一化学键连接信息保持不变,公式8用于将被掩码的子结构的第二化学键连接信息变换到服从高斯分布的隐变量。Sθ和Tθ可以采用神经网络结构,如,图神经网络结构。
示例性地,分子补全模型包括化学键补全模型,化学键补全模型为一种基于流的生成模型,也即化学键补全模型用于基于流的生成方式补全化学键连接信息。示例性地,分子补全模型还可以称为Synthon Flow(合成子流)模型,化学键补全模型还可以称为Synthon Bond Flow(合成子化学键流,简称SB Flow)模型。
化学键补全模型为一种可逆模型,化学键补全模型的逆模型与化学键补全模型的关系为:化学键补全模型的逆模型的输入为化学键补全模型的输出,化学键补全模型的逆模型的输出为化学键补全模型的输入。示例性地,化学键补全模型的输入和输出的信息的维度相同。示例性地,化学键补全模型的输入为分子中的原子之间的化学键连接情况的假设信息,输出为分子中的原子之间的化学键连接情况的表征信息,也就是说,化学键模型的逆模型的输入为分子中的原子之间的化学键连接情况的表征信息,输出为分子中的原子之间的化学键连接情况的假设信息。
由于当前的输入信息为样本化学键连接信息(也即样本化合物分子中的原子之间的化学键连接情况的表征信息),所以,该步骤5023可以通过调用分子补全模型中的化学键补全模型的逆模型实现。也就是说,调用分子补全模型中的化学键补全模型的逆模型基于化学键掩码信息对样本化学键连接信息进行逆变换,得到样本化学键连接隐变量。
示例性地,化学键补全模型包括至少一个化学键补全模块,各个化学键补全模块的结构相同,本申请实施例以化学键补全模型包括一个化学键补全模块为例进行说明。
示例性地,化学键补全模型的结构可以如图8所示,原子补全模型包括挤压(Squeeze)模块、标准化处理(Actnorm)模块、可逆卷积(Invertible Convolution)模块、分裂/掩码(Split/Mask)模块、仿射耦合(Affine Coupling)模块以及至少一个变换信息获取模块。示例性地,变换信息获取模块包括卷积子模块、标准化处理(Batchnorm)子模块和激活(Relu)子模块。在图8中,变换信息获取模块的数量为l(l为不小于1的整数)个,可逆卷积模块的卷积核为1*1,变换信息获取模块中的卷积子模块的卷积核为3*3。示例性地,挤压模块用于将输入的化学键连接信息的维度进行变换;标准化处理模块用于对信息进行标准化处理;可逆卷积模块用于对化学键连接信息中的种类维度进行重排;分裂/掩码模块用于基于化学键掩码信息将样本化合物的样本化学键连接信息拆分成两个部分(第一化学键连接信息和第二化学键连接信息);仿射耦合模块用于实现对第二化学键连接信息的逆变换;l个变换信息获取模块用于获取第一样本变换信息。
示例性地,在图8所示的化学键补全模型的结构下,获取样本化学键连接隐变量的过程可以为:将化学键掩码信息MB和样本化学键连接信息BR输入化学键补全模型,依次经过挤压模块、标准化处理模块、可逆卷积模块和分裂/掩码模块的处理,得到第一化学键连接信息和第二化学键连接信息利用l个变换信息获取模块(每个变换信息获取模块均包括卷积子模块、标准化处理子模块和激活子模块)对第一化学键连接信息进行处理,能够得到第一样本变换信息然后由仿射耦合模块利用第一样本变换信息对第二化学键连接信息进行逆变换,得到子结构的化学键连接隐变量基于第一化学键连接信息和子结构的化学键连接隐变量得到样本化学键连接隐变量
需要说明的是,图8所示的化学键补全模型的结构仅为一种示例性举例,本申请实施例并不局限于此。也就是说,化学键补全模型的结构还可以包括更多或更少的模块。
在示例性实施例中,基于化学键掩码信息对样本化学键连接信息进行逆变换可以是直接基于化学键掩码信息对样本化学键连接信息进行逆变换,也可以是指基于化学键掩码信息对处理后的化学键连接信息进行逆变换,其中,处理后的化学键连接信息可以通过调用GLOW模块(一种信息处理模块)对样本化学键连接信息进行处理得到。
步骤5024:计算机设备调用分子补全模型,基于原子掩码信息对样本原子特征信息进行逆变换,得到样本原子特征隐变量。
样本原子特征隐变量用于对样本化合物分子中的原子的特征进行假设。分子补全模型为一种可逆模型,既能够将分子的原子特征隐变量变换为分子的原子特征信息,也能够将分子的原子特征信息逆变换为分子的原子特征隐向量。由于在模型的训练过程中,样本化合物分子的样本原子特征信息为已知的较为准确的信息,所以,调用分子补全模型实现对样本原子特征信息的逆变换。
在一种可能实现方式中,计算机设备调用分子补全模型基于原子掩码信息对样本原子特征隐变量进行逆变换,得到样本原子特征隐变量的过程包括以下步骤50241至步骤50243。
步骤50241:计算机设备调用分子补全模型,基于原子掩码信息和样本原子特征信息,确定第一原子特征信息和第二原子特征信息,该第一原子特征信息为样本待补全分子的原子特征信息,该第二原子特征信息为子结构的原子特征信息。
第一原子特征信息用于表征样本待补全分子中的原子的特征,第二化学键连接信息用于表征被掩码的子结构中的原子的特征。由于原子掩码信息能够指示出样本化合物分子中的原子的掩码情况,被掩码的原子即为被掩码的子结构中的原子,没有被掩码的原子即为样本待补全分子中的原子,所以基于原子掩码信息和样本化合物分子的样本原子特征,能够获取样本待补全分子的第一原子特征信息以及子结构的第二原子特征信息。
示例性地,基于原子掩码信息和样本原子特征信息,能够获取样本待补全分子的第一原子特征信息以及子结构的第二原子特征信息的方式可以为:将样本原子特征信息中的与原子掩码信息指示的没有被掩码的原子相关的信息保留,将其他信息置为0,得到第一原子特征信息;将样本原子特征信息中的与原子掩码信息指示的被掩码的原子相关的信息保留,将其他信息置为0,得到第二原子特征信息。此种方式下,第一原子特征信息的维度和第二原子特征信息的维度均与样本原子特征信息的矩阵相同,如,均为第二参考维度的矩阵。
步骤50242:计算机设备基于第一原子特征信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量。
子结构的原子特征隐变量用于对子结构中的原子的特征进行假设,示例性地,子结构的原子特征隐变量为一种服从已知概率分布的变量,已知概率分布根据经验设置,或者根据应用场景灵活调整,例如,已知概率分布为高斯分布,或者为均匀分布等。
在示例性实施例中,基于第一原子特征信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量的实现过程包括:基于第一原子特征信息,获取第二样本变换信息;利用第二样本变换信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量。示例性地,第二样本变换信息可以基于至少一个变换函数(如,Sθ(·)变换函数和Tθ(·)变换函数) 获取。需要说明的是,逆变换不会改变信息的维度,也即子结构的原子特征隐变量的维度与第二原子特征信息的维度相同。
在示例性实施例中,基于第一原子特征信息对第二原子特征信息进行逆变换的过程需要考虑样本化学键连接信息的约束,以保证逆变换过程的可靠性。此种情况下,需要基于样本化学键连接信息,获取样本约束信息,然后基于第一原子特征信息和样本约束信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量。此种情况下,第二样本变换信息通过综合考虑第一原子特征信息和样本约束信息得到。
本申请实施例对获取样本约束信息的方式不加以限定,可以根据经验设置,也可以根据应用场景灵活调整。示例性地,获取样本约束信息的方式可以为将样本化学键连接信息作为样本约束信息。示例性地,获取样本约束信息的方式还可以为对样本化学键连接信息进行标准化处理,得到样本约束信息。对样本化学键连接信息进行标准化处理用于提高样本化学键连接信息的规范性,标准化处理的方式可以根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加以限定。例如,标准化处理可以调用图标准化模块实现。
步骤50243:计算机设备基于第一原子特征信息和子结构的原子特征隐变量,获取样本原子特征隐变量。
示例性地,第一原子特征信息和子结构的原子特征隐变量均为第二参考维度的矩阵,基于第一原子特征信息和子结构的原子特征隐变量,获取样本原子特征隐变量的方式可以为:将第一原子特征信息和子结构的原子特征隐变量中的对应位置的元素相加,将相加后得到的矩阵作为样本原子特征隐变量。示例性地,还可以将第一原子特征信息和子结构的原子特征隐变量之间的笛卡尔乘积作为样本原子特征隐变量。
示例性地,获取样本原子特征隐变量的过程可以基于下述公式9和公式10实现:

其中,表示第一原子特征信息;表示样本约束信息;表示第二原子特征信息;表示第二样本变换信息; 表示基于第二样本变换信息对第二原子特征信息进行逆变换的运算方式;示子结构的原子特征隐变量。为样本原子特征隐变量的两个组成部分。
示例性地,公式9用于将样本待补全分子的第一原子特征信息保持不变,公式10用于将被掩码的子结构的第二原子特征信息变换到服从高斯分布的隐变量,公式10将第一原子特征信息和基于样本化学键连接信息BR获取的样本约束信息作为输入条件。Sθ和Tθ可以采用神经网络结构,如,图神经网络结构,Sθ和Tθ的输出维度均与第二原子特征信息的维度相同,Sθ和Tθ的处理逻辑可以利用公式11表示:
其中,hA表示Sθ或Tθ函数的输出结构(也即);MA∈{0,1}表示原子掩码信息,用于将非样本待补全分子(也即被掩码的子结构)的原子特征信息掩码成零,将样本待补全分子的原子特征信息保持不变。MA的维度与样本原子特征信息的维度相同,MA中包括样本原子特征信息中的每个元素的掩码值,示例性地,对于所有的j(j为任一原子的子特征信息的各个维度中的任一维度),如果va∈VS,则MA[a,j]=1,否则MA[a,j]=0。其中,va∈VS表示原子a为样本待补全分子中的原子构成的原子集VS中的一个原子,M[a,j]表示原子a的子特征信息中第j个维度的元素的掩码值。
表示第一原子特征信息,能够在MA的基础上,通过MA⊙AR运算得到。表示基于样本化学键连接信息获取的样本约束信息中与化学键类别i相关的信息,i为不小于1且不大于C的整数,C(C为不小于1的整数)为候选化学键类型的数量。Graphconv()表示图神经网络结构;Wi和W0表示图神经网络结构的参数。
示例性地,分子补全模型包括原子补全模型,原子补全模型为一种基于流的生成模型, 也即原子补全模型用于基于流的生成方式补全原子特征信息。示例性地,原子补全模型还可以称为Synthon Graph Flow(合成子图结构流,简称SG Flow)模型。
原子补全模型为一种可逆模型,原子补全模型的逆模型与原子补全模型的关系为:原子补全模型的逆模型的输入为原子补全模型的输出,原子补全模型的逆模型的输出为原子补全模型的输入。示例性地,原子补全模型的输入和输出的信息的维度相同。示例性地,原子补全模型的输入为分子中的原子的特征的假设信息,输出为分子中的原子的特征表征信息,也就是说,化学键模型的逆模型的输入为分子中的原子的特征的表征信息,输出为分子中的原子的假设信息。
由于当前的输入信息为样本原子特征信息(也即样本化合物分子中的原子的特征的表征信息),所以,该步骤5024可以通过调用分子补全模型中的原子补全模型的逆模型实现。也就是说,调用分子补全模型中的原子补全模型的逆模型基于原子掩码信息对样本原子特征信息进行逆变换,得到样本原子特征隐变量。
示例性地,原子补全模型包括至少一个原子补全模块,各个原子补全模块的结构相同,本申请实施例以原子补全模型包括一个原子补全模块为例进行说明。
示例性地,原子补全模型的结构可以如图9所示,原子补全模型包括标准化处理(Actnorm)模块、分裂/掩码(Split/Mask)模块、仿射耦合(Affine Coupling)模块、图标准化(Graphnorm)模块以及变换信息获取模块。示例性地,变换信息获取模块包括至少一个参考处理模块和一个多层感知机(MLP)模块,任一参考处理模块包括图卷积子模块、标准化处理(Batchnorm)子模块和激活(Relu)子模块。在图9中,参考处理模块的数量为l(l为不小于1的整数)个,分裂/掩码模块用于基于原子掩码信息将样本化合物的样本原子特征信息拆分成两个部分(第一原子特征信息和第二原子特征信息);标准化处理模块用于在每一个批次(batch)内针对矩阵的每一行进行标准化操作;仿射耦合模块用于实现对第二原子特征信息的逆变换;变换信息获取模块用于获取第二样本变换信息。
示例性地,在图9所示的原子补全模型的结构下,获取样本原子特征隐变量的过程可以为:将原子掩码信息MA和样本原子特征信息AR输入原子补全模型,依次经过标准化处理模块和分裂/掩码模块的处理,得到第一原子特征信息和第二原子特征信息通过图标准化模块对样本化学键连接信息BR进行标准化处理,得到样本约束信息利用包括l个参考处理模块(每个参考处理模块均包括图卷积子模块、标准化处理子模块和激活子模块)和一个MLP模块的变换信息获取模块对第一原子特征信息和样本约束信息进行处理,能够得到第二样本变换信息然后由仿射耦合模块利用第二样本变换信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量基于第一原子特征信息和子结构的原子特征隐变量得到样本原子特征隐变量
需要说明的是,图9所示的原子补全模型的结构仅为一种示例性举例,本申请实施例并不局限于此。也就是说,原子补全模型的结构还可以包括更多或更少的模块。
步骤5025:计算机设备基于样本化学键连接隐变量和样本原子特征隐变量,确定训练损失。
训练损失用于衡量样本化学键连接隐变量和样本原子特征隐变量的预测质量。例如训练损失为数值,训练损失越大,说明样本化学键连接隐变量和样本原子特征隐变量的预测质量越差,也即说明分子补全模型的性能越差;训练损失越小,说明样本化学键连接隐变量和样本原子特征隐变量的预测质量越好,也即说明分子补全模型的性能越好。
在示例性实施例中,基于样本化学键连接隐变量和样本原子特征隐变量,确定训练损失的实现过程包括:基于样本化学键连接隐变量和第一样本变换信息,确定第一似然函数值;基于样本原子特征隐变量和第二样本变换信息,确定第二似然函数值;基于第一似然函数值和第二似然函数值,确定目标似然函数值;将与目标似然函数值呈负相关关系的数值确定为训练损失。例如。将目标似然函数值的相反数作为训练损失,或者,将目标似然函数值的对 数值的相反数作为训练损失等。
第一似然函数值用于衡量在给定样本化学键连接隐变量的基础上,调用化学键补全模型得到样本化学键连接信息的概率,示例性地,第一似然函数值可以基于下述公式12计算得到:
其中,表示第一似然函数值;BR表示样本化学键连接信息;表示样本化学键连接隐变量;表示样本化学键连接隐变量的概率;是一个行列式,表示第一似然函数值的对数值与样本化学键连接隐变量的概率的对数值之间的差异,基于第一样本变换信息计算得到。
第二似然函数值用于衡量在给定样本原子特征隐变量的基础上,调用化学键补全模型得到样本原子特征信息的概率,示例性地,第二似然函数值可以基于下述公式13计算得到:
其中,表示第二似然函数值;表示样本原子特征隐变量的概率;表示在的约束下得到的样本原子特征信息;表示在的约束下得到的样本原子特征隐变量;表示样本约束信息;是一个行列式,表示第二似然函数值的对数值与样本原子特征隐变量的概率的对数值之间的差异,基于第二样本变换信息计算得到。
示例性地,基于第一似然函数值和第二似然函数值,获取目标似然函数值的过程如公式14所示:
其中,GR表示样本化合物分子;BR表示样本化学键连接信息;表示样本原子特征信息;表示第一似然函数值;表示第二似然函数值;表示目标似然函数值。
示例性地,获取目标似然函数值的过程可以如下所示:
输入:样本化合物分子的信息GR=(AR,BR)和其掩码信息M(包括原子掩码信息和化学键掩码信息),原子补全模型的逆模型和化学键补全模型的逆模型其中,AR表示样本原子特征信息,BR表示样本化学键连接信息;
1、//利用GLOW模块(一种信息处理模块)对样本化学键连接信息BR进行处理,得到处理后的化学键连接信息
2、//调用化学键补全模型的逆模型基于处理后的化学键连接信息获取化学键连接隐变量此过程考虑了掩码信息M中化学键掩码信息
3、//基于化学键连接信息以及行列式获取第一似然函数值
4、//利用Graphnorm模块(一种图标准化模块)对样本化学键连接信息BR进行标准化处理,得到样本约束信息
5、//调用原子补全模型的逆模型基于样本原子特征信息AR和样本约束信息获取原子特征隐变量此过程考虑了掩码信息M中的原子掩码信息
6、//基于原子特征隐变量以及行列式获取第二似然函数值
7、//将包括原子特征隐变量和化学键连接隐变量的信息作为样本隐变量
8、//基于第一似然函数值和第二似然函数值获取目标似然函数值的对数值
输出:
需要说明的是,基于步骤5021至步骤5025确定训练损失的过程仅为一种示例性实现过程,本申请实施例并不局限于此。在示例性实施例中,基于样本化合物分子、样本待补全分子和分子补全模型获取训练损失的实现方式还可以为:计算机设备调用分子补全模型对样本待补全分子进行补全,得到补全结果,基于补全结果确定预测补全分子;基于预测补全分子和样本化合物分子之间的差异,确定训练损失。
调用分子补全模型对样本待补全分子进行补全的实现过程参见图2所示的实施例中步骤203,此处不再赘述。基于调用分子补全模型对样本待补全分子进行补全得到的补全结果,能够得到预测补全分子,预测补全分子为分子补全模型预测的样本待补全分子的补全后的分子。样本化合物分子为样本待补全分子的真实的补全后的分子,基于预测补全分子和样本化合物分子之间的差异,能够获取用于为分子补全模型的模型参数更新提供监督信息的训练损失。
本申请实施例对衡量两个分子之间的差异的方式不加以限定,示例性地,基于两个分子中的原子的差异(原子数量的差异、原子类型的差异、原子特征的差异等)和化学键之间的差异(化学键数量的差异,化学键类型的差异、化学键特征的差异等),确定两个分子之间的差异。示例性地,提取两个分子的分子特征,将两个分子特征之间的差异作为两个分子之间的差异。示例性地,可以通过调用分子特征提取模型提取两个分子的分子特征。示例性地,两个分子特征为相同维度的向量或矩阵,可以基于两个分子特征中的对应位置的元素之间的差异确定两个分子特征之间的差异。示例性地,也可以计算两个分子特征之间的相似度,将与相似度呈负相关关系的数值(如,相似度的相反数)作为两个分子特征之间的差异。
无论哪种情况,在确定训练损失后,基于训练损失更新分子补全模型的模型参数。示例性地,基于训练损失更新分子补全模型的模型参数的过程可以基于梯度下降法实现,也即,基于训练损失获取分子补全模型的模型参数的更新梯度,基于更新梯度更新分子补全模型的模型参数。示例性地,对于基于步骤5021至步骤5025确定训练损失的情况,基于训练损失更新分子补全模型的模型参数的过程也可以是指基于训练损失更新分子补全模型的逆模型的过程。
示例性地,模型训练过程为迭代过程,在基于训练损失更新分子补全模型的模型参数后,得到训练一次的分子补全模型;判断当前训练过程是否满足目标终止条件,若当前训练过程满足目标终止条件,则可以将训练一次的分子补全模型作为训练后的分子补全模型;若当前训练过程不满足目标终止条件,则可以参考步骤501和步骤502的方式获取新的训练损失,然后利用新的训练损失对当前得到的分子补全模型的模型参数进行更新,以此类推,直至当前训练过程满足目标终止条件,将满足目标终止条件时得到的分子补全模型作为训练后的分子补全模型。需要说明的是,获取新的训练损失所依据的样本化合物分子以及样本待补全分子可以部分或完全改变,也可以不改变。
满足目标终止条件根据经验设置,或者根据应用场景灵活调整,本申请实施例对此不加 以限定。在示例性实施例中,满足目标终止条件可以是指训练损失收敛、训练损失小于第二损失阈值、模型参数的更新次数达到第二次数阈值等。第二损失阈值和第二次数阈值可以根据经验设置,也可以根据应用场景灵活调整,本申请实施例对此不加以限定。
在示例性实施例中,训练得到分子补全模型的过程可以为一种课程学习(curriculum learning)的过程,此种情况下,可以构造从易到难的任务逐步对分子补全模型进行训练。示例性地,任务的难度可以基于样本化合物分子中被掩码的子结构的复杂程度或者样本化合物分子中被掩码的子结构掩码比例(masking ratio)来衡量。若样本化合物分子中被掩码的子结构的复杂程度较低或者样本化合物分子中被掩码的子结构掩码比例较低,则模型能够接收到的信息较多,分子补全任务较容易;若样本化合物分子中被掩码的子结构的复杂程度较高或者样本化合物分子中被掩码的子结构掩码比例较高,则模型能够接收到的信息较少,分子补全任务较困难。通过从易到难的任务逐步对分子补全模型进行训练,可以逐步提高分子补全模型的训练效果。当然,在一些实施例中,也可以直接利用复杂的任务(如,在子图级别的掩码方案下构建的任务)对分子补全模型进行训练,本申请实施例对此不加以限定。
示例性地,对于通过从易到难的任务逐步对分子补全模型进行训练的情况,满足目标终止条件还可以是指基于难度最大的任务对分子补全模型训练完毕。示例性地,基于任一难度的任务对分子补全模型训练是指基于该任一难度的样本数据对分子补全模型训练,例如样本数据是与该任一难度匹配的样本化合物分子及其样本待补全分子)。基于任一难度的任务对分子补全模型训练完毕可以是指在基于该任一难度的样本数据对分子补全模型进行训练的过程中,损失达到收敛,或者损失小于某一损失阈值,再或者训练次数达到某一次数阈值等。
在示例性实施例中,训练得到分子补全模型的过程还可以为先预训练再微调的过程,其中,预训练所依据的样本数据基于任一个无标签的化合物分子数据集中的化合物分子得到,微调所依据的样本数据基于已知的合成反应中的产物分子得到。此种情况下,满足目标终止条件可以是指分子补全模型微调完毕,如,在基于微调所依据的样本数据对分子补全模型进行训练的过程中,损失达到收敛,或者损失小于某一损失阈值,再或者训练次数达到某一次数阈值等。
需要说明的是,模型训练过程中所利用的一些常见参数,如,学习率、迭代周期(epoch)以及批次大小(batch size)等均可以根据经验设置,或者根据计算机设备的计算能力、应用场景等灵活调整,本申请实施例对此不加以限定。
本申请实施例提供的分子补全模型的训练过程,可以先在一个无标签的大数据集上进行预训练,该操作间接地进行了数据增广,扩充了更多可利用的有效信息,从而能够提高分子补全模型的泛化能力,能够适应更广泛的反应物分子预测场景,从而提高反应物分子的预测可靠性和预测准确性。由于预训练的分子补全任务与反应物分子预测过程中的待补全分子的补全任务非常相关,所以在预训练后,可以在包括已知的合成反应的数据集(如,逆合成数据集)的基础上对分子补全模型进行微调,以使模型会具备更强的泛化能力。示例性地,反应物分子预测过程中的待补全分子的补全任务可以认为是一般性的分子补全任务的一些特例。
示例性地,本申请实施例使用基于图表示的分子重构作为自监督学习任务,在方案能够扩展以往的分子自监督学习任务,将“Mask and Fill”的思想更好的应用到图结构数据领域。将自监督学习策略应用于反应物分子的预测领域,将模型在分子大数据集上进行自监督任务(如分子结构补全)的学习,进而在逆合成数据集进行微调,最终可以通过给定的产物分子直接预测出反应物分子,进而推导合成路径。此种方式训练得到的模型可提升反应物分子的预测能力,突破数据瓶颈,泛化能力更强。此外,本申请实施例可以使用基于流的生成模型实现分子的补全,基于流的生成模型是非自回归的生成模型,可以一次性生成,相比自回归模型生成效率更高,推断速度更快,能够达到相近或更高的预测准确率,并且基于流的生成模型可以给出预测结果的似然函数值,有更好的可解释性。
本申请实施例提供的分子补全模型的训练方法,基于样本化合物分子和样本待补全分子对分子补全模型进行训练。由于样本待补全分子通过对样本化合物分子中的子结构进行掩码 得到,也就是说,分子补全模型的训练过程所依据的数据是在样本化合物分子本身的基础上得到的数据,此种训练过程为一种基于样本化合物分子的自监督训练过程,此种自监督训练过程无需关注样本化合物分子是否为已知的合成反应中的化合物,因而此种自监督训练过程并不会受已知的合成反应的限制,利用该训练过程训练得到的分子补全模型的泛化能力较强,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
参见图10,本申请实施例提供了一种反应物分子的预测装置,设置于计算机设备中,该装置包括:
第一获取单元1001,用于获取产物分子,对产物分子进行断键,得到待补全分子,产物分子是指待预测反应物分子的任一化合物分子;
补全单元1002,用于调用分子补全模型对待补全分子进行补全,得到补全结果,基于补全结果确定产物分子的反应物分子;
其中,分子补全模型基于样本化合物分子以及样本待补全分子训练得到,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到。
在一种可能实现方式中,补全单元1002,用于基于待补全分子的原子特征信息,确定目标原子特征隐变量,目标原子特征隐变量为补全后的分子的原子特征隐变量;基于待补全分子的化学键连接信息,确定目标化学键连接隐变量,目标化学键连接隐变量为补全后的分子的化学键连接隐变量;调用分子补全模型对目标化学键连接隐变量进行变换,得到目标化学键连接信息,对目标原子特征隐变量进行变换,得到目标原子特征信息;基于目标化学键连接信息和目标原子特征信息,确定待补全分子的补全结果。
在一种可能实现方式中,第一获取单元1001,用于获取产物分子的图结构信息;基于图结构信息,预测产物分子中化学键的断裂概率,将断裂概率满足参考条件的化学键确定为产物分子中的断裂化学键;基于断裂化学键对产物分子进行断键,得到待补全分子。
本申请实施例提供的反应物分子的预测装置,反应物分子的预测过程依赖分子补全模型实现,分子补全模型是基于样本化合物分子和样本待补全分子训练得到的。由于样本待补全分子通过对样本化合物分子中的子结构进行掩码得到,也就是说,分子补全模型的训练过程所依据的数据是在样本化合物分子本身的基础上得到的数据,此种训练过程为一种基于样本化合物分子的自监督训练过程,此种自监督训练过程无需关注样本化合物分子是否为已知的合成反应中的化合物,因而此种自监督训练过程并不会受已知的合成反应的限制,利用该训练过程训练得到的分子补全模型的泛化能力较强,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
参见图11,本申请实施例提供了一种分子补全模型的训练装置,该装置包括:
第二获取单元1101,用于获取样本化合物分子以及样本待补全分子,样本待补全分子通过对样本化合物分子中的子结构进行掩码得到;
第三获取单元1102,用于基于样本化合物分子、样本待补全分子和分子补全模型,确定训练损失;
更新单元1103,用于基于训练损失更新分子补全模型的模型参数,得到训练后的分子补全模型。
在一种可能实现方式中,第三获取单元1102,用于获取样本化合物分子的样本原子特征信息和样本化学键连接信息;基于样本化合物分子和样本待补全分子之间的差异,确定原子掩码信息和化学键掩码信息;调用分子补全模型基于化学键掩码信息对样本化学键连接信息进行逆变换,得到样本化学键连接隐变量,基于原子掩码信息对样本原子特征信息进行逆变换,得到样本原子特征隐变量;基于样本化学键连接隐向量和样本原子特征隐向量,确定训练损失。
在一种可能实现方式中,第三获取单元1102,用于调用分子补全模型基于化学键掩码信 息和样本化学键连接信息,确定第一化学键连接信息和第二化学键连接信息,第一化学键连接信息为样本待补全分子的化学键连接信息,第二化学键连接信息为子结构的化学键连接信息;基于第一化学键连接信息对第二化学键连接信息进行逆变换,得到子结构的化学键连接隐变量;基于第一化学键连接信息和子结构的化学键连接隐变量,确定样本化学键连接隐变量。
在一种可能实现方式中,第三获取单元1102,用于基于原子掩码信息和样本原子特征信息,确定第一原子特征信息和第二原子特征信息,第一原子特征信息为样本待补全分子的原子特征信息,第二原子特征信息为子结构的原子特征信息;基于第一原子特征信息对第二原子特征信息进行逆变换,得到子结构的原子特征隐变量;基于第一原子特征信息和子结构的原子特征隐变量,确定样本原子特征隐变量。
在一种可能实现方式中,第三获取单元1102,用于调用分子补全模型对样本待补全分子进行补全,得到补全结果,基于补全结果确定预测补全分子;基于预测补全分子和样本化合物分子之间的差异,确定训练损失。
在一种可能实现方式中,子结构为样本化合物分子中的属于候选结构集的结构,候选结构集为可信度满足选取条件的结构的集合。
本申请实施例提供的分子补全模型的训练装置,基于样本化合物分子和样本待补全分子对分子补全模型进行训练。由于样本待补全分子通过对样本化合物分子中的子结构进行掩码得到,也就是说,分子补全模型的训练过程所依据的数据是在样本化合物分子本身的基础上得到的数据,此种训练过程为一种基于样本化合物分子的自监督训练过程,此种自监督训练过程无需关注样本化合物分子是否为已知的合成反应中的化合物,因而此种自监督训练过程并不会受已知的合成反应的限制,利用该训练过程训练得到的目标分子补全模型的泛化能力较强,有利于扩展适应场景,从而有利于提高反应物分子的预测可靠性和预测准确性。
需要说明的是,上述实施例提供的装置在实现其功能时,仅以上述各功能单元的划分进行举例说明,实际应用中,可根据需要而将上述功能分配由不同的功能单元完成,即将设备的内部结构划分成不同的功能单元,以完成以上描述的全部或者部分功能。另外,上述实施例提供的装置与方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
在示例性实施例中,还提供了一种计算机设备,该计算机设备包括处理器和存储器,该存储器中存储有至少一条计算机程序。该至少一条计算机程序由一个或者一个以上处理器加载并执行,以使该计算机设备实现上述任一种反应物分子的预测方法或分子补全模型的训练方法。该计算机设备可以为服务器,也可以为终端。接下来,对服务器和终端的结构分别进行介绍。
图12是本申请实施例提供的一种服务器的结构示意图,该服务器可因配置或性能不同而产生比较大的差异,可以包括一个或多个处理器(Central Processing Units,CPU)1201和一个或多个存储器1202,其中,该一个或多个存储器1202中存储有至少一条计算机程序,该至少一条计算机程序由该一个或多个处理器1201加载并执行,以使该服务器实现上述各个方法实施例提供的反应物分子的预测方法或分子补全模型的训练方法。当然,该服务器还可以具有有线或无线网络接口、键盘以及输入输出接口等部件,以便进行输入输出,该服务器还可以包括其他用于实现设备功能的部件,在此不做赘述。
图13是本申请实施例提供的一种终端的结构示意图。该终端可以是:PC、手机、智能手机、PDA、可穿戴设备、PPC、平板电脑、智能车机、智能电视、智能音箱、智能语音交互设备、智能家电、车载终端、VR设备、AR设备。终端还可能被称为用户设备、便携式终端、膝上型终端、台式终端等其他名称。通常,终端包括有:处理器1301和存储器1302。
处理器1301可以包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器1301可以采用DSP(Digital Signal Processing,数字信号处理)、FPGA(Field-Programmable Gate Array,现场可编程门阵列)、PLA(Programmable Logic Array,可编程逻辑阵列)中的至 少一种硬件形式来实现。
存储器1302可以包括一个或多个计算机可读存储介质,该计算机可读存储介质可以是非暂态的。存储器1302还可包括高速随机存取存储器,以及非易失性存储器,比如一个或多个磁盘存储设备、闪存存储设备。在一些实施例中,存储器1302中的非暂态的计算机可读存储介质用于存储至少一个指令,该至少一个指令用于被处理器1301所执行,以使该终端实现本申请中方法实施例提供的反应物分子的预测方法或分子补全模型的训练方法。
在一些实施例中,终端还可选包括有:外围设备接口1303和至少一个外围设备。处理器1301、存储器1302和外围设备接口1303之间可以通过总线或信号线相连。各个外围设备可以通过总线、信号线或电路板与外围设备接口1303相连。具体地,外围设备包括:射频电路1304或显示屏1305中的至少一种。
外围设备接口1303可被用于将I/O(Input/Output,输入/输出)相关的至少一个外围设备连接到处理器1301和存储器1302。在一些实施例中,处理器1301、存储器1302和外围设备接口1303被集成在同一芯片或电路板上;在一些其他实施例中,处理器1301、存储器1302和外围设备接口1303中的任意一个或两个可以在单独的芯片或电路板上实现,本实施例对此不加以限定。
射频电路1304用于接收和发射RF(Radio Frequency,射频)信号,也称电磁信号。射频电路1304通过电磁信号与通信网络以及其他通信设备进行通信。射频电路1304将电信号转换为电磁信号进行发送,或者,将接收到的电磁信号转换为电信号。
显示屏1305用于显示UI(User Interface,用户界面)。该UI可以包括图形、文本、图标、视频及其它们的任意组合。当显示屏1305是触摸显示屏时,显示屏1305还具有采集在显示屏1305的表面或表面上方的触摸信号的能力。该触摸信号可以作为控制信号输入至处理器1301进行处理。此时,显示屏1305还可以用于提供虚拟按钮和/或虚拟键盘,也称软按钮和/或软键盘。
本领域技术人员可以理解,图13中示出的结构并不构成对终端的限定,可以包括比图示更多或更少的组件,或者组合某些组件,或者采用不同的组件布置。
在示例性实施例中,还提供了一种计算机可读存储介质,该计算机可读存储介质中存储有至少一条计算机程序,该至少一条计算机程序由计算机设备的处理器加载并执行,以使计算机实现上述任一种反应物分子的预测方法或分子补全模型的训练方法。
在一种可能实现方式中,上述计算机可读存储介质可以是只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、只读光盘(Compact Disc Read-Only Memory,CD-ROM)、磁带、软盘和光数据存储设备等。
在示例性实施例中,还提供了一种计算机程序产品,该计算机程序产品包括计算机程序或计算机指令,该计算机程序或计算机指令由处理器加载并执行,以使计算机实现上述任一种反应物分子的预测方法或分子补全模型的训练方法。
需要说明的是,本申请所涉及的信息(包括但不限于用户设备信息、用户个人信息等)、数据(包括但不限于用于分析的数据、存储的数据、展示的数据等)以及信号,均为经用户授权或者经过各方充分授权的,且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准。例如,本申请中涉及到的产物分子等是在充分授权的情况下获取的。
应当理解的是,在本文中提及的“多个”是指两个或两个以上。“和/或”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。字符“/”一般表示前后关联对象是一种“或”的关系。
以上所述仅为本申请的示例性实施例,并不用以限制本申请,凡在本申请的原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。

Claims (22)

  1. 一种反应物分子的预测方法,所述方法包括:
    计算机设备获取产物分子,对所述产物分子进行断键,得到待补全分子,所述产物分子是指待预测反应物分子的任一化合物分子;
    所述计算机设备调用分子补全模型对所述待补全分子进行补全,得到补全结果,基于所述补全结果确定所述产物分子的反应物分子;
    其中,所述分子补全模型基于样本化合物分子以及样本待补全分子训练得到,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到。
  2. 根据权利要求1所述的方法,其中,所述计算机设备调用分子补全模型对所述待补全分子进行补全,得到补全结果,包括:
    所述计算机设备基于所述待补全分子的原子特征信息,确定目标原子特征隐变量,所述目标原子特征隐变量为补全后的分子的原子特征隐变量;
    所述计算机设备基于所述待补全分子的化学键连接信息,确定目标化学键连接隐变量,所述目标化学键连接隐变量为补全后的分子的化学键连接隐变量;
    所述计算机设备调用所述分子补全模型,对所述目标化学键连接隐变量进行变换,得到目标化学键连接信息,对所述目标原子特征隐变量进行变换,得到目标原子特征信息;
    所述计算机设备基于所述目标化学键连接信息和所述目标原子特征信息,确定所述待补全分子的补全结果。
  3. 根据权利要求2所述的方法,其中,所述对所述目标化学键连接隐变量进行变换,得到目标化学键连接信息,包括:
    基于所述目标化学键连接隐变量,获取所述待补全分子的参考化学键连接信息和缺失结构的参考化学键连接隐变量;
    基于所述待补全分子的参考化学键连接信息对所述缺失结构的参考化学键连接隐变量进行变换,得到所述缺失结构的参考化学键连接信息;
    基于所述待补全分子的参考化学键连接信息和所述缺失结构的参考化学键连接信息,确定所述目标化学键连接信息。
  4. 根据权利要求2所述的方法,其中,所述对所述目标原子特征隐变量进行变换,得到目标原子特征信息,包括:
    基于所述目标原子特征隐变量,获取所述待补全分子的参考原子特征信息和缺失结构的参考原子特征隐变量;
    基于所述待补全分子的参考原子特征信息对所述缺失结构的参考原子特征隐变量进行变换,得到所述缺失结构的参考原子特征信息;
    基于所述待补全分子的参考原子特征信息和所述缺失结构的参考原子特征信息,确定所述目标原子特征信息。
  5. 根据权利要求1-4任一项所述的方法,其中,所述对所述产物分子进行断键,得到待补全分子,包括:
    获取所述产物分子的图结构信息;
    基于所述图结构信息,预测所述产物分子中化学键的断裂概率;
    将所述断裂概率满足参考条件的化学键确定为所述产物分子中的断裂化学键;
    基于所述断裂化学键对所述产物分子进行断键,得到所述待补全分子。
  6. 根据权利要求5所述的方法,其中,所述产物分子的图结构信息包括所述产物分子的原子特征信息和化学键连接信息,所述原子特征信息包括所述产物分子中每个原子的子特征信息,所述化学键连接信息包括所述产物分子中每个化学键的子特征信息;
    所述获取所述产物分子的图结构信息,包括:
    获取所述产物分子中每个原子的属性信息,对每个原子的属性信息进行特征提取,得到每个原子的子特征信息;
    获取所述产物分子中每个化学键的属性信息,对每个化学键的属性信息进行特征提取,得到每个化学键的子特征信息。
  7. 根据权利要求5所述的方法,其中,所述基于所述图结构信息,预测所述产物分子中化学键的断裂概率,包括:
    调用图神经网络模型基于所述图结构信息提取所述产物分子中化学键的目标特征,基于所述目标特征预测所述产物分子中化学键的断裂概率,所述目标特征为预测所述化学键的断裂概率所依据的特征。
  8. 根据权利要求7所述的方法,其中,所述图神经网络模型的训练过程,包括:
    所述计算机设备获取训练化合物分子的图结构信息和所述训练化合物分子中化学键的标准断裂概率;
    所述计算机设备调用所述图神经网络模型基于所述训练化合分子的图结构信息,预测所述训练化合物分子中化学键的训练断裂概率;
    所述计算机设备基于所述标准断裂概率和所述训练断裂概率之间的差异,确定参考损失;
    所述计算机设备基于所述参考损失更新所述图神经网络模型的模型参数;
    所述计算机设备响应于训练过程满足第一终止条件,将当前训练得到的图神经网络模型确定为训练完成的图神经网络模型。
  9. 根据权利要求1-4任一项所述的方法,其中,所述基于所述补全结果确定所述产物分子的反应物分子,包括以下任一项:
    所述补全结果为补全后的分子的图结构信息,将所述图结构信息指示的分子确定为所述反应物分子;
    所述补全结果为所述待补全分子中缺失结构的指示信息,根据所述指示信息确定所述缺失结构,将所述缺失结构与所述待补全分子进行连接,将连接后得到的分子确定为所述反应物分子。
  10. 根据权利要求9所述的方法,其中,所述指示信息包括所述待补全分子中缺失结构的分类结果以及连接位置预测结果,所述分子结果包括所述待补全分子中缺失结构与参考结构的匹配概率,所述参考结构包括单原子级别结构和单化学键级别结构,所述连接位置预测结果指示所述待补全分子中待与缺失结构连接的位置;
    所述根据所述指示信息确定所述缺失结构,将所述缺失结构与所述待补全分子进行连接,将连接后得到的分子确定为所述反应物分子,包括:
    根据所述分类结果,确定所述待补全分子中缺失结构的类别,根据所述连接位置预测结果确定所述待补全分子中待与缺失结构连接的位置;
    在所述待补全分子中待与缺失结构连接的位置处,将所述类别的结构与所述待补全分子连接,将连接后得到的分子确定为所述反应物分子。
  11. 一种分子补全模型的训练方法,所述方法包括:
    计算机设备获取样本化合物分子以及样本待补全分子,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到;
    所述计算机设备基于所述样本化合物分子、所述样本待补全分子和分子补全模型,确定训练损失;
    所述计算机设备基于所述训练损失更新所述分子补全模型的模型参数,得到训练后的分子补全模型。
  12. 根据权利要求11所述的方法,其中,所述计算机设备基于所述样本化合物分子、所述样本待补全分子和分子补全模型,确定训练损失,包括:
    所述计算机设备获取所述样本化合物分子的样本原子特征信息和样本化学键连接信息;
    所述计算机设备基于所述样本化合物分子和所述样本待补全分子之间的差异,确定原子掩码信息和化学键掩码信息;
    所述计算机设备调用所述分子补全模型,基于所述化学键掩码信息对所述样本化学键连接信息进行逆变换,得到样本化学键连接隐变量;
    所述计算机设备调用所述分子补全模型,基于所述原子掩码信息对所述样本原子特征信息进行逆变换,得到样本原子特征隐变量;
    所述计算机设备基于所述样本化学键连接隐向量和所述样本原子特征隐向量,确定所述训练损失。
  13. 根据权利要求12所述的方法,其中,所述计算机设备调用所述分子补全模型,基于所述化学键掩码信息对所述样本化学键连接信息进行逆变换,得到样本化学键连接隐变量,包括:
    所述计算机设备调用所述分子补全模型基于所述化学键掩码信息和所述样本化学键连接信息,确定第一化学键连接信息和第二化学键连接信息,所述第一化学键连接信息为所述样本待补全分子的化学键连接信息,所述第二化学键连接信息为所述子结构的化学键连接信息;
    所述计算机设备基于所述第一化学键连接信息对所述第二化学键连接信息进行逆变换,得到所述子结构的化学键连接隐变量;
    所述计算机设备基于所述第一化学键连接信息和所述化学键连接隐变量,确定所述样本化学键连接隐变量。
  14. 根据权利要求12所述的方法,其中,所述计算机设备调用所述分子补全模型,基于所述原子掩码信息对所述样本原子特征信息进行逆变换,得到样本原子特征隐变量,包括:
    所述计算机设备调用所述分子补全模型基于所述原子掩码信息和所述样本原子特征信息,确定第一原子特征信息和第二原子特征信息,所述第一原子特征信息为所述样本待补全分子的原子特征信息,所述第二原子特征信息为所述子结构的原子特征信息;
    所述计算机设备基于所述第一原子特征信息对所述第二原子特征信息进行逆变换,得到所述子结构的原子特征隐变量;
    所述计算机设备基于所述第一原子特征信息和所述原子特征隐变量,确定所述样本原子特征隐变量。
  15. 根据权利要求12所述的方法,其中,所述计算机设备基于所述样本化学键连接隐向量和所述样本原子特征隐向量,确定所述训练损失,包括:
    所述计算机设备基于所述样本化学键连接隐变量和第一样本变换信息,确定第一似然函数值;
    所述计算机设备基于所述样本原子特征隐变量和第二样本变换信息,确定第二似然函数值;
    所述计算机设备基于所述第一似然函数值和所述第二似然函数值,确定目标似然函数值;
    所述计算机设备将与所述目标似然函数值呈负相关关系的数值确定为所述训练损失。
  16. 根据权利要求11所述的方法,其中,所述计算机设备基于所述样本化合物分子、所述样本待补全分子和分子补全模型,确定训练损失,包括:
    所述计算机设备调用所述分子补全模型对所述样本待补全分子进行补全,得到补全结果,基于所述补全结果,确定预测补全分子;
    所述计算机设备基于所述预测补全分子和所述样本化合物分子之间的差异,确定所述训练损失。
  17. 根据权利要求11-16任一所述的方法,其中,所述子结构为所述样本化合物分子中的属于候选结构集的结构,所述候选结构集为可信度满足选取条件的结构的集合。
  18. 一种反应物分子的预测装置,设置于计算机设备中,所述装置包括:
    第一获取单元,用于获取产物分子,对所述产物分子进行断键,得到待补全分子,所述产物分子是指待预测反应物分子的任一化合物分子;
    补全单元,用于调用分子补全模型对所述待补全分子进行补全,得到补全结果,基于所述补全结果确定所述产物分子的反应物分子;
    其中,所述分子补全模型基于样本化合物分子以及样本待补全分子训练得到,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到。
  19. 一种分子补全模型的训练装置,设置于计算机设备中,所述装置包括:
    第二获取单元,用于获取样本化合物分子以及样本待补全分子,所述样本待补全分子通过对所述样本化合物分子中的子结构进行掩码得到;
    第三获取单元,用于基于所述样本化合物分子、所述样本待补全分子和分子补全模型,确定训练损失;
    更新单元,用于基于所述训练损失更新所述分子补全模型的模型参数,得到训练后的分子补全模型。
  20. 一种计算机设备,所述计算机设备包括处理器和存储器,所述存储器中存储有至少一条计算机程序,所述至少一条计算机程序由所述处理器加载并执行,以使所述计算机设备实现如权利要求1至10任一所述的反应物分子的预测方法,或者如权利要求11至17任一所述的分子补全模型的训练方法。
  21. 一种计算机可读存储介质,所述计算机可读存储介质中存储有至少一条计算机程序,所述至少一条计算机程序由处理器加载并执行,以使计算机设备实现如权利要求1至10任一所述的反应物分子的预测方法,或者如权利要求11至17任一所述的分子补全模型的训练方法。
  22. 一种计算机程序产品,所述计算机程序产品包括计算机程序或计算机指令,所述计算机程序或所述计算机指令由处理器加载并执行,以使计算机设备实现如权利要求1至10任一所述的反应物分子的预测方法,或者如权利要求11至17任一所述的分子补全模型的训练方法。
PCT/CN2023/092036 2022-07-14 2023-05-04 反应物分子的预测、模型的训练方法、装置、设备及介质 Ceased WO2024012017A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/597,636 US20240212796A1 (en) 2022-07-14 2024-03-06 Reactant molecule prediction using molecule completion model

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202210830979.XA CN115206451B (zh) 2022-07-14 2022-07-14 反应物分子的预测、模型的训练方法、装置、设备及介质
CN202210830979.X 2022-07-14

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/597,636 Continuation US20240212796A1 (en) 2022-07-14 2024-03-06 Reactant molecule prediction using molecule completion model

Publications (1)

Publication Number Publication Date
WO2024012017A1 true WO2024012017A1 (zh) 2024-01-18

Family

ID=83581953

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/092036 Ceased WO2024012017A1 (zh) 2022-07-14 2023-05-04 反应物分子的预测、模型的训练方法、装置、设备及介质

Country Status (3)

Country Link
US (1) US20240212796A1 (zh)
CN (1) CN115206451B (zh)
WO (1) WO2024012017A1 (zh)

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115206451B (zh) * 2022-07-14 2025-10-14 腾讯科技(深圳)有限公司 反应物分子的预测、模型的训练方法、装置、设备及介质
CN116010875B (zh) * 2022-12-01 2026-01-23 国网重庆市电力公司营销服务中心 电表故障的分类方法、装置、电子设备及计算机存储介质
CN115966263B (zh) * 2022-12-21 2025-07-01 西北工业大学 一种基于原子特征传递网络的小分子单步逆合成预测方法
CN116864019A (zh) * 2023-07-26 2023-10-10 蔚泓智能信息科技(上海)有限公司 一种基于ai预测的化合物合成路线预测系统
CN119069034B (zh) * 2024-07-22 2026-02-06 中山大学 一种基于深度强化学习和课程学习的片段连接小分子化合物的优化方法
CN118782168A (zh) * 2024-09-10 2024-10-15 烟台国工智能科技有限公司 一种基于多步预测的合成路线排序方法及装置
JP7652479B1 (ja) 2025-01-24 2025-03-27 Mi-6株式会社 逆合成解析システム、逆合成解析装置、逆合成解析方法、及び、逆合成解析プログラム
CN120954537B (zh) * 2025-10-16 2026-03-24 南京大学 一种化学反应大语言模型训练方法及合成路径规划方法

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111524557A (zh) * 2020-04-24 2020-08-11 腾讯科技(深圳)有限公司 基于人工智能的逆合成预测方法、装置、设备及存储介质
CN113850801A (zh) * 2021-10-18 2021-12-28 深圳晶泰科技有限公司 晶型预测方法、装置及电子设备
CN114300065A (zh) * 2021-12-10 2022-04-08 深圳晶泰科技有限公司 分子设计方案的确定方法、装置、设备及存储介质
CN115206451A (zh) * 2022-07-14 2022-10-18 腾讯科技(深圳)有限公司 反应物分子的预测、模型的训练方法、装置、设备及介质

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11354582B1 (en) * 2020-12-16 2022-06-07 Ro5 Inc. System and method for automated retrosynthesis
CN113990405B (zh) * 2021-10-19 2024-05-31 上海药明康德新药开发有限公司 试剂化合物预测模型的构建方法、化学反应试剂自动预测补全的方法与装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111524557A (zh) * 2020-04-24 2020-08-11 腾讯科技(深圳)有限公司 基于人工智能的逆合成预测方法、装置、设备及存储介质
CN113850801A (zh) * 2021-10-18 2021-12-28 深圳晶泰科技有限公司 晶型预测方法、装置及电子设备
CN114300065A (zh) * 2021-12-10 2022-04-08 深圳晶泰科技有限公司 分子设计方案的确定方法、装置、设备及存储介质
CN115206451A (zh) * 2022-07-14 2022-10-18 腾讯科技(深圳)有限公司 反应物分子的预测、模型的训练方法、装置、设备及介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
CHAOCHAO YAN, QIANGGANG DING, PEILIN ZHAO, SHUANGJIA ZHENG, JINYU YANG, YANG YU, JUNZHOU HUANG: "Interpretable Retrosynthesis Prediction in Two Steps", CHEMRXIV, AMERICAN CHEMICAL SOCIETY (ACS), 21 February 2020 (2020-02-21), XP093129006, DOI: 10.26434/chemrxiv.11869692.v1 *

Also Published As

Publication number Publication date
CN115206451B (zh) 2025-10-14
CN115206451A (zh) 2022-10-18
US20240212796A1 (en) 2024-06-27

Similar Documents

Publication Publication Date Title
WO2024012017A1 (zh) 反应物分子的预测、模型的训练方法、装置、设备及介质
CN111259142B (zh) 基于注意力编码和图卷积网络的特定目标情感分类方法
EP4592866B1 (en) Data processing method and apparatus
CN111951805B (zh) 一种文本数据处理方法及装置
US20220044767A1 (en) Compound property analysis method, model training method, apparatuses, and storage medium
CN116720004B (zh) 推荐理由生成方法、装置、设备及存储介质
CN113361593B (zh) 生成图像分类模型的方法、路侧设备及云控平台
CN118246537B (zh) 基于大模型的问答方法、装置、设备及存储介质
CN117094395B (zh) 对知识图谱进行补全的方法、装置和计算机存储介质
WO2020147369A1 (zh) 自然语言处理方法、训练方法及数据处理设备
CN110019952B (zh) 视频描述方法、系统及装置
CN110781302A (zh) 文本中事件角色的处理方法、装置、设备及存储介质
CN117874234A (zh) 基于语义的文本分类方法、装置、计算机设备及存储介质
CN116186295A (zh) 基于注意力的知识图谱链接预测方法、装置、设备及介质
CN113434683A (zh) 文本分类方法、装置、介质及电子设备
CN120494107A (zh) 基于人工智能的样本生成方法、训练方法以及问答方法
CN116306612A (zh) 一种词句生成方法及相关设备
CN115062136B (zh) 基于图神经网络的事件消歧方法及其相关设备
US20250246191A1 (en) Interactive method based on large model, training method, intelligent agent, device, and medium
US20250299485A1 (en) Multi-object tracking using hierarchical graph neural networks
CN116680392A (zh) 一种关系三元组的抽取方法和装置
CN121074897B (zh) 多模态意图识别方法、装置、计算机设备和存储介质
US20260038190A1 (en) Filtering three-dimensional shape data for training text to 3d generative ai systems and applications
US20260038191A1 (en) Automatic annotation of three-dimensional shape data for training text to 3d generative ai systems and applications
CN117037930B (zh) 属性模型的训练方法、装置、设备、存储介质及程序产品

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23838515

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 03-06-2025)

122 Ep: pct application non-entry in european phase

Ref document number: 23838515

Country of ref document: EP

Kind code of ref document: A1