EP3642398A1 - Procédé et dispositif de sélection d'un sous-ensemble de molécules destinées à être utilisées pour prédire au moins une propriété d'une structure moléculaire - Google Patents
Procédé et dispositif de sélection d'un sous-ensemble de molécules destinées à être utilisées pour prédire au moins une propriété d'une structure moléculaireInfo
- Publication number
- EP3642398A1 EP3642398A1 EP18749450.5A EP18749450A EP3642398A1 EP 3642398 A1 EP3642398 A1 EP 3642398A1 EP 18749450 A EP18749450 A EP 18749450A EP 3642398 A1 EP3642398 A1 EP 3642398A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- molecules
- molecule
- descriptor
- value
- selection
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/40—Searching chemical structures or physicochemical data
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C20/00—Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
- G16C20/30—Prediction of properties of chemical compounds, compositions or mixtures
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16C—COMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
- G16C60/00—Computational materials science, i.e. ICT specially adapted for investigating the physical or chemical properties of materials or phenomena associated with their design, synthesis, processing, characterisation or utilisation
Definitions
- the invention relates to the general field of chemical molecules.
- It relates more particularly to the prediction of properties of a molecule having a molecular structure.
- the invention thus has a preferred but non-limiting application in the prediction of the toxicity of compounds, inert or energetic materials, or even highly energetic materials, which, in known manner, are capable of releasing energy in a very short time. Because of the energy released, such energetic materials are of interest to both military and civilian domains. They are nowadays commonly used in the manufacture of military machines, enter the constitution of gas (ex.propergol) necessary for the propulsion of missiles and space launchers, or are still used in the automobile industry for the manufacture of airbags, etc.
- gas ex.propergol
- QSAR Quantitative Structure Activity Relationship
- a QSAR is applied directly to the entire database.
- the database on which the QSAR is applied may contain molecules that are too different from the molecular substance whose biological activity is to be predicted, so that the resulting prediction can be made. prove wrong.
- Each MACCS structural key 166 is more precisely a vector comprising 166 components or characteristics, having positive or zero values and reflecting the presence or absence of one of the 166 molecular fragments in the molecule in question: thus, a zero value reflects the absence of the corresponding fragment in the structure of the molecule, while a positive value indicates the number of times that the corresponding fragment is present within the molecule, or simply its presence within the molecule.
- a metric conventionally used in combination with the structural keys MACCS 166 is the Tanimoto metric defined by:
- - Xi ⁇ Vf is equal to 1 if the components X t and ⁇ are both positive, and to 0 otherwise; and - Xi v Y t is equal to 1 if at least one of the components X t and ( is non-zero, and at 0 otherwise.
- this metric is applied by simplifying the MACCS structural key 166 of each molecule so as to obtain a binary vector, a zero component value reflecting the absence of the corresponding molecular fragment, while a component value equal to 1 translated the presence of this fragment.
- the Tanimoto metric thus calculated thus provides the ratio between the number of components of the keys X and Y common to the two molecular structures on the total number of components of the keys X and Y expressed (ie to which a non-zero value has been assigned in the keys ) for these two molecular structures.
- a structural similarity search is performed on the database, leading to the identification of a subset of molecules in the database having a minimal similarity to the molecular substance whose properties are to be predicted. Then a QSAR is applied on the subset of molecules thus identified. It is therefore clear that, depending on the similarity threshold that one sets to select the subset of molecules, it is possible to obtain a subset that does not contain enough molecules to apply the QSAR of in a relevant way, or on the contrary a subset which contains molecules too different from the molecular substance whose properties one seeks to predict. This can result in a false prediction.
- One known strategy for improving the performance of the aforementioned strategy is to identify a subset of molecules in the database from another known subset of molecules (eg subset of high energy molecules used). by an industrialist), and to select the molecules of the database which have a minimum similarity with each of the molecules of the known subset.
- a QSAR is then applied to the subset of the database thus identified from the known subset of molecules.
- the invention proposes a strategy for predicting the properties of an alternative molecular substance to the strategies proposed in the state of the art and making it possible to obtain a better quality prediction.
- the invention proposes an iterative method of selecting a subset of said reference molecules intended to be used for predicting at least one property of a so-called target molecular structure, the iterative process of selection comprising an initialization step associating with a so-called current molecule a value of a predetermined molecule descriptor, associated with the target molecular structure, and during each iteration of the selection method:
- An evaluation step for each molecule of a base comprising a plurality of molecules each associated with a value of the descriptor, of a so-called overall similarity measure between the value of the descriptor associated with said molecule and the value of the associated descriptor; to the current molecule;
- the invention is directed to a device for selecting a subset of said reference molecules intended to be used for predicting at least one property of a so-called target molecular structure, the selection device comprising a configured initialization module. for associating with a so-called current molecule a value of a predetermined molecule descriptor associated with the target molecular structure, said selection device being further configured to activate, during a plurality of successive iterations:
- An evaluation module configured to evaluate, for each molecule of a base comprising a plurality of molecules each associated with a value of the descriptor, a so-called global similarity measure between the value of the descriptor associated with said molecule and the value of the descriptor associated with the current molecule;
- a selection module configured to select molecules of the base having a global similarity measurement greater than a predetermined threshold, the selected molecules being added by said selection module to the reference subset;
- An update module configured to update the value of the descriptor associated with the current molecule from the values of the descriptors associated with at least a part of the molecules belonging to the reference subset.
- the invention also provides, according to a second aspect, a method for predicting at least one property of a so-called target molecular substance comprising: A selection step, by means of an iterative selection process according to the invention, of a subset of said reference molecules in a database comprising a plurality of molecules each associated with a value of a descriptor predetermined molecules;
- the invention also relates to a prediction device configured to predict at least one property of a target molecular substance comprising:
- a selection device configured to select a subset of said reference molecules in a database comprising a plurality of molecules each associated with a value of a predetermined molecule descriptor
- a prediction module configured to predict at least one property of said target molecular substance from the subset of reference molecules selected.
- This descriptor may be a descriptor comprising a plurality N of characteristics or components, N denoting an integer greater than or equal to 1, in which case the value of the descriptor is defined by the value of each of its N characteristics.
- N characteristics can be, for example, structural characteristics making it possible to characterize each molecule and if possible to discriminate between them.
- the values of the N characteristics of the molecule descriptor may reflect the presence or absence of N molecular fragments considered in the definition of a structural key MACCS 166.
- descriptors may be envisaged, such as other known two-dimensional (or fingerprints) descriptors such as MolPrint2D fingerprints, BCI, or those defined by the companies Tripos and Scitegic. These fingerprints are in the form of bit vectors, each bit encoding the presence (bit equal to 1) or the absence (bit equal to 0) of certain predefined structural fragments in the molecule or other characteristics.
- the invention also applies to other types of descriptors than 2D fingerprints.
- a descriptor having the form of a simple variable that is, comprising a single component / characteristic
- whose value can be a quantitative or qualitative numerical value
- descriptors having more complex forms such as vector, matrix or even graphic forms.
- a descriptor is for example a connectivity matrix between a plurality of predetermined atoms indicating for each pair of atoms the presence or absence of a bond in the molecule in question (the descriptor then comprises a plurality of characteristics or characteristics given by the components of the matrix).
- the invention therefore proposes a new way of selecting the molecules of the initial database used to predict the properties of a molecular substance, and which makes it possible to select a larger subset of molecules similar to the molecular substance and relevant for the prediction of its properties.
- This new way of selecting molecules is based on an iterative process of similarity search, initialized first with the target molecular substance whose properties are to be predicted. Then, over the iterations, "virtual" molecules are constructed from the descriptors of the molecules selected in the initial database during the iterations, and a new similarity search is performed from these virtual molecules.
- the invention thus leads, thanks to this recursive selection and to the taking into account of the similarities with the molecules of the database, to a more complete and more careful selection of the molecules of the base intended to be used for predicting the biological properties. of the target molecular substance.
- the prediction produced by the invention is advantageously adaptive. It can easily use public databases, regularly updated, and listing the properties of different molecules with regard to different tests performed on these molecules.
- the number of iterations considered for selecting the subset of reference molecules can be fixed by means of a parameterizable stopping criterion.
- the evaluation, selection and updating steps are then repeated until a predetermined stopping criterion is verified.
- Different stopping criteria can be envisaged, for example:
- the number of iterations and / or molecules of the reference subset can be calibrated empirically.
- the choice of one or the other of the aforementioned criteria may depend on several parameters, such as, for example, the type of target molecular substance considered, a compromise between the number of molecules selected and the quality. prediction, the method that will be used to predict the properties of the target molecular substance from the properties of the selected molecules, etc.
- the evaluation step comprises, for each molecule of the base, a step of calculating, for each of the N characteristics of the descriptor, a so-called local similarity measure between the value of this characteristic of the descriptor associated with said molecule and the value of this characteristic of the descriptor associated with the current molecule, the global similarity measure evaluated for said molecule being obtained from the local similarity measurements calculated for this molecule.
- the calculation step includes for each descriptor feature:
- Such a calculation step advantageously makes it possible to obtain a measurement of similarity that is more precise than in the state of the art. It can be easily applied to numerical values (eg integers) of descriptor characteristics that are positive or null, and not just binary. This gives an assessment of the similarity between two molecular substances more precise and more generic than in the state of the art.
- the calculated distance denoted d, can verify:
- a measure of similarity is defined as a real number between 0 and 1, taking conventionally the value 0 when the two molecules are considered totally different (ie not similar), and the value 1 when they are considered. as totally identical (ie similar). Intermediate values can be considered, representing shades of similarity between these two extremes. To comply with this definition, different conversion functions may be considered.
- the conversion function can verify:
- d denotes the distance to be converted and ⁇ a predetermined real number.
- the overall similarity measure evaluated for said molecule is the ratio between:
- This definition of the global similarity measure makes it possible to take into account several levels of expression of the same descriptor characteristic in the compared molecules: it is not limited to discerning only two levels of binary expression (absence or presence of the characteristic of the descriptor) unlike in particular the metric Tanimoto described above and considered in the state of the art.
- this global similarity measure advantageously considers that the common non-expression of the same descriptor (i.e. null value for this descriptor for the two compared molecules) is a mark of similarity between the two compared molecules.
- This current molecule is in a way the representative of the molecules of the reference subset used at the next iteration to complete the reference subset.
- said at least part of the molecules belonging to the reference subset used for the update comprises the molecules selected during the selection step of this iteration that did not already belong to the reference set before this selection step.
- this first variant only the newly selected molecules are taken into account during the current iteration.
- This first variant may, however, lead to the selection in the reference set of molecules a little too far in terms of similarity of the target molecular structure.
- the said at least part of the molecules belonging to the reference subset used for the update comprises the molecules selected during the step of selecting this iteration.
- the said at least part of the molecules belonging to the reference subset used for the update all comprise the molecules belonging to the reference subset at the end of the step of selecting this iteration.
- the inventors have found that the second and third variants above have a fairly similar behavior and lead to comparable results in terms of prediction. They also give better results than the first variant.
- the value associated with the current molecule of each descriptor characteristic is updated with an arithmetic or weighted average of the values of this characteristic of the descriptor associated with the molecules of said descriptor. least part of the molecules belonging to the reference subset.
- This first variant leads to values of the characteristics of the descriptor which are in some way "artificial", and do not correspond to characteristic values present in said at least part of the molecules of the subset used for the update.
- the value associated with the current molecule of each feature of the descriptor is updated with the most frequent value of this characteristic of the descriptor among the values of this feature of the descriptor associated with the molecules of said at least a portion of the molecules belonging to the reference subset, or if a plurality of distinct values satisfy this condition, with the highest value among this plurality of distinct values.
- the various steps of the selection method and / or the prediction method are determined by computer program instructions.
- the invention also relates to a computer program on an information carrier, this program being capable of being implemented in a device of selection, respectively in a prediction device, or more generally in a computer, this program comprising instructions adapted to the implementation of the steps of a selection method, respectively of a prediction method, as described above .
- This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other form desirable shape.
- the invention also relates to a computer readable information or recording medium, and comprising instructions of a computer program as mentioned above.
- the information or recording medium may be any entity or device capable of storing the program.
- the medium may comprise storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a hard disk.
- the information or recording medium may be a transmissive medium such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means.
- the program according to the invention can be downloaded in particular on an Internet type network.
- the information or recording medium may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the method in question.
- the selection method, the prediction method, the selection device and the prediction device according to the invention present in combination all or part of the aforementioned characteristics.
- FIG. 1 shows, schematically, a prediction device according to the invention, in a particular embodiment
- FIG. 2 represents the hardware architecture of the prediction device of FIG. 1, in a particular embodiment
- FIG. 3 illustrates the different steps of a selection method according to the invention
- FIG. 4 illustrates the different steps of a prediction method according to the invention.
- Annexes 1 to 6 show the performances achieved by the prediction method according to the invention. Detailed description of the invention
- FIG. 1 represents, in its environment, a prediction device 1 according to the invention, in a particular embodiment.
- the prediction device 1 is configured to predict at least one property of a substance called TARGm target unknown. It is assumed that this target substance has a mono-molecular structure from which it is possible to extract the value of a descriptor comprising a predetermined number N of (structural) characteristics for characterizing the target substance.
- HEM high energy molecule
- TARGm target substance By “prediction of at least one property of TARGm target substance” is meant here the prediction of its biological activity.
- a property that we are trying to predict may be for example a toxicological property of TARGm target substance, in particular to meet the requirements of the European REACH Regulation.
- the invention also applies to the prediction of other types of properties of a molecule, such as, for example, physico-chemical properties (logP or molecular weight), structural properties, absorption properties, distribution properties, of Metabolism, or Elimination (ADMET), therapeutic properties, etc.
- the prediction device 1 comprises:
- a selection device 2 according to the invention.
- the prediction device 1 has the hardware architecture of a computer as represented in FIG. 2, and the selection device 2 and the prediction module 3 are software modules installed in a memory of the prediction device 1. More particularly, the prediction device 1 comprises in particular a processor 4, a random access memory 5, a read-only memory 6, a non-volatile flash memory 7, input / output interfaces 8 (such as a screen, a keyboard, etc. .), as well as means of communication 9.
- the prediction device 1 comprises in particular a processor 4, a random access memory 5, a read-only memory 6, a non-volatile flash memory 7, input / output interfaces 8 (such as a screen, a keyboard, etc. .), as well as means of communication 9.
- the databases 10 may be hosted on remote servers or stored in a memory of the prediction device 1 (for example in its non-volatile memory 7).
- the communication means 9 of the prediction device 1 allow it to access or download them via a telecommunications network, or to obtain these databases via a recording medium such as a USB key (Universal Serial). Bus) or a CDROM. They can include for this purpose a USB port, a network card, a WIFI interface (WIreless FIdelity), etc.
- the read-only memory 6 of the prediction device 1 constitutes a recording medium in accordance with the invention, readable by the processor 4 and on which is recorded here a computer program PROG according to the invention.
- the computer program PROG defines functional modules (and software here), configured to implement the steps of the selection method and the prediction method according to the invention.
- the two aforementioned methods can be defined by instructions from two different programs.
- the functional modules defined by the program PROG rely on and / or control the hardware elements 4-9 of the prediction device 1 mentioned above. They include in particular here, as illustrated in FIG.
- An initialization module 2A configured to associate with a current so-called CURm molecule updated during the selection process according to the invention, the value of the MACCS descriptor 166 associated with the target molecule TARGm (the value of the descriptor comprising here N characteristics) ;
- An evaluation module 2B configured to evaluate so-called "global" similarity measurements between the values of the descriptors associated with a predetermined set of molecules (typically the molecules of a database 10) and the value of the descriptor associated with the current molecule CURm;
- a selection module 2C configured to select molecules of the predetermined set considered having a global similarity measurement greater than a predetermined threshold, and to add the molecules thus selected to a so-called reference subset designated CREF;
- a 2D update module configured to update the value of the descriptor associated with the current molecule CURm from the values of the descriptors associated with at least a part of the molecules belonging to the CREF reference subset.
- Evaluation module 2B, selection module 2C and 2D update module are modules of selection device 2, and are configured for the implementation of a selection method according to the invention. They are activated by the selection device 2 repeatedly during a plurality of iterations, and more specifically in the embodiment described here, as long as a predetermined criterion (parameterizable) is not verified.
- the program PROG here also defines the prediction module 3 of the prediction device 1.
- the prediction module 3 is configured to predict at least one property of the target molecular substance TARGm from the molecules of the reference subset CREF selected by the selection device 2.
- No limitation is attached to the prediction technique implemented by the prediction module 3. It may be for example a QSAR type relationship, a neural network, a prediction technique by principal component analysis, etc.
- This prediction technique uses the experimental results achieved by the molecules of the reference subset CREF listed in the database 10 whose subset CREF was extracted.
- the prediction device 3 predicts at least one property of the molecular substance TARGm from the properties listed in the databases 10 for a plurality of molecules.
- the prediction made by the prediction device 3 is based on a prior selection by the selection device 2 of a reference subset CREF comprising a plurality of molecules extracted from the database 10.
- FIG. 3 illustrates the main steps of the selection method according to the invention implemented by the selection device 2 in order to make this selection of the CREF reference subset.
- the selection method is an iterative method, comprising an initialization step (step E10) and implementing a plurality of iterations.
- the iterations are linked as long as a predetermined stop criterion CRU is not checked.
- the different stopping criteria envisaged are described in more detail later.
- the initialization module 2A of the selection device 2 initializes the reference subset CREF to an empty set.
- MACCS (CURm, l),..., MACCS (CURm, N) are subsequently denoted as the values of the N MACCS characteristics associated with the current molecule CURm.
- the selection device 2 then starts the iterations of the selection process (step E20 of incrementing the index iter).
- S global similarity metric
- This global similarity metric is more precisely calculated here between the N values of the N characteristics of the associated MACCS descriptor 166 in the base 10 of the MOLk molecule and the N values of the N characteristics of the MACCS descriptor 166 associated with the current molecule CURm (step E30 ).
- local similarity measures are defined here from a local similarity function Is which at any pair of integer characteristic values (x, y) associates a real number ls (x, y) (denoted here ls (CURm, MOLk, n) for the nth characteristic), between 0 and 1 and satisfying the following properties:
- ls (x, y) ls (y, x) for x and y any natural integers.
- the function Is results from the composition of a function d comparable to a geometric distance between the values x and y, and a function f of converting the distance between x and y into a measurement local similarity, ie:
- the evaluation module 2B uses the distance d thus defined:
- evaluation module 2B uses as conversion function f, a standardized Gauss function defined
- ⁇ is a predetermined real number
- a conversion function is preferably chosen, associating with any number of the real straight line a real value between 0 and 1 such that:
- ls (CURm, MOLk, n) f (d (MACCS (CURm, n), MACCS (MOLk, n)) where MACCS (CURm, n) and MACCS (MOLk, n) respectively denote the value of the nth characteristic of the MACCS descriptor of the current molecule CURm and the value of the nth characteristic of the MACCS descriptor of the molecule of the MOLk molecule.
- this expression of the overall similarity results from a search by the inventors of a similarity measure which, unlike the Tanimoto metric commonly used in the techniques of the prior art, makes it possible to take into account different levels of expression of the same characteristic of the descriptor (ie different values of the same characteristic) between two compared molecules, and which also considers the common non-expression of the same descriptor characteristic (ie null value of this characteristic) as a mark of similarity between the two compared molecules.
- considering that the pairs of MOL-A and MOL-B molecules corresponding to different MACCS descriptor characteristics have empty intersections, and where w n , n 1, ..., N denote real weights. Then asking:
- different real weights of 1 can be applied by the evaluation module 2.
- these weights can be determined by expertise from a business knowledge of the relevance of each feature of the descriptor given the type of TARGm target molecule whose property is to be predicted.
- These weights can also be determined using statistical methods, in particular classification methods such as Linear Discriminant Analysis (LDA), which makes it possible to determine weights leading to a better discrimination between the experimentally positive molecules. (ie who are considered to have responded positively to the toxicity test considered) and negative (ie who are considered to have responded negatively to the toxicity test considered).
- LDA Linear Discriminant Analysis
- the selection device 2 determines which molecules of the base 10 have measuring overall similarity greater than a predetermined threshold THRmin (or equivalently greater than or equal to a predetermined threshold THRmin ') and selects them (step E40).
- the molecules thus selected form a set C (iter) of molecules considered to be similar to the current molecule CURm.
- the threshold THRmin is a constant parameter here during the iterations of the selection process, and between 0 and 1. It may depend in particular on the type of target molecule TARGm whose properties are to be determined (eg high energy molecule, solvent, plasticizers, liquid, etc.). This threshold can be determined experimentally beforehand.
- the inventors have determined by experimentation that a threshold
- THRmin 0.85 (or greater than or equal to 0.85) leads to good predictions for different categories of molecules (fillers, plasticizers, liquids, etc.).
- the THRmin threshold may change over the iterations.
- the set of molecules C (iter) selected during the current iteration iter is then added by the selection module 2C to the set of reference CREF (step E50). It should be noted that certain molecules contained in the set C (iter) may already be present in the reference set CREF, in which case the addition of the molecules of the set C (iter) to the set of reference CREF is limited to add only the new molecules not already present in the CREF reference set.
- the selection device 2 via its 2D update module, updates the value of the MACCS descriptor associated with the current molecule (step E60).
- this update is carried out from the descriptor values of at least a part of the molecules present in the CREF reference subset at the end of step E50.
- the update of the MACCS descriptor characteristic values of the current molecule CURm is based on the values of the characteristics of the MACCS descriptor of the molecules selected during the iterative iteration, ie on the molecules contained in the set C (iter).
- the update of the MACCS descriptor characteristic values of the current molecule CURm is based on the MACCS descriptor characteristic values of all the molecules belonging to the CREF reference set at the end of step E50.
- the updating of the MACCS descriptor characteristic values of the current molecule CURm is based solely on the values of the characteristics of the MACCS descriptor of the newly selected molecules during the selection step E40 implemented. during the current iteration iter, in other words on the values of the characteristics of the MACCS descriptor of the molecules belonging to the set C (iter) but which do not already belong to the set of reference CREF before the step E50.
- the update module 2D uses the most frequent value of each characteristic among the values of this characteristic associated with the molecules considered for the update. In case of ambiguity, that is, if several distinct values satisfy this frequency condition, the 2D update module uses the highest value among this plurality of distinct values.
- the 2D update module may use an average of the values of this characteristic associated with the molecules considered for updating (or the integer value closest to this average to obtain integer characteristics), this average possibly being an arithmetic or weighted average.
- a new current molecule CURm is thus obtained on which a new search for similarity in the base 10 can be performed during the next iteration.
- the selection device 2 verifies, at the end of step E60, whether the CRU stop criterion is verified (test step E70).
- Different stopping criteria can be envisaged, for example:
- This stopping criterion can be parameterizable.
- the numbers ITERMAX and KMAX are also parameterizable, and depend in particular on the type of molecules considered.
- step E20 a new iteration of the selection method is implemented (incrementation step E20), this iteration comprising the repetition of the steps E30 to E70 for the new current molecule CURm obtained during step E60.
- the iterations of the selection method are interrupted and the reference set CREF is supplied to the prediction module 3 for the prediction of the properties of the molecule substance TARGm target.
- the reference set CREF considered is preferably that obtained at the end of the iteration making it possible not to exceed the KMAX number.
- FIG. 4 illustrates the different steps of the prediction method implemented by the prediction device 1.
- step F10 repeats the steps of the selection method of the reference subset CREF previously described with reference to FIG. 3 and implemented by the selection device 2 of the prediction device 1.
- the reference set CREF obtained by the selection device 2 is then supplied to the prediction module 3.
- the latter is configured to predict at least one property of the target molecular substance TARGm from the molecules of the reference set CREF selected by the selection device 2 (step F20).
- prediction technique implemented by the prediction module 3 for this purpose. It can in particular use a QSAR type relationship as described above and commonly used in the state of the art, or a neural network, a prediction technique by principal component analysis, etc.
- This prediction technique uses the experimental results achieved by the molecules of the CREF reference set and listed in the database whose CREF set has been extracted. The use of such prediction techniques is known per se and is not described in more detail here.
- the prediction device 1 then obtains at the end of step F20 a prediction of at least one biological property of TARGm target molecular substance. Other predictions can be made by the prediction device 1 from other databases corresponding to other biological tests.
- the invention via the proposed new selection method, makes it possible to obtain a reliable prediction of the properties of a molecular substance from the properties of molecules of the same type listed in public databases in particular.
- the inventors have observed an improvement in the predictions obtained with respect to the state of the art prediction techniques for different categories of molecules (fillers, plasticizers, oxidizers, liquids, stabilizers, pyrotechnic components, etc.) and for various regulatory tests known to those skilled in the art (eg AMES mutagenicity test, chromosome aberration test, UDS unscheduled DNA synthesis test, carcinogenicity test, etc.).
- Some results are provided in Annexes 1 to 6 to illustrate the performance of the selection and prediction methods according to the invention.
- Appendix 1 illustrates prediction results obtained for the AMES test using five different prediction methods.
- the AMES test is, in a known manner, a mutagenicity test carried out on different bacterial cultures and aimed at determining whether a molecule has a mutagenic property (indicated in the table in Appendix 1 by a "+” symbol, a "-” symbol indicating that the molecule does not exhibit mutagenic property).
- the table presents in its first column data which were obtained experimentally from the molecules tested. These data were validated at European level and were used as a reference to determine the relevance of the predictions made using the different prediction methods tested. For each of these methods, when a result obtained is between 0 and 0.4, it is considered negative, that is to say as reflecting the absence of mutagenic property in the molecule tested; when this result is between 0.4 and 0.6, it is considered doubtful; and when this result is greater than 0.6, it is considered to be negative, that is to say as reflecting the presence of the mutagenic property in the molecule tested.
- the table given in Appendix 1 provides the prediction results obtained via the five methods tested for different charge-type molecules: the five prediction methods were each applied on a starting data base comprising 7723 reference molecules. More precisely :
- the column of the table bearing the reference (1) corresponds to the application of a QSAR relation on the initial data base;
- the column of the table bearing the reference (2) corresponds to the application of a QSAR relation on a database obtained by selecting in the starting database the molecules presenting a similarity metric (Tanimoto metric) of 0.8 ;
- the column of the table bearing the reference (3) corresponds to the application of a QSAR relation on a database obtained by selecting in the starting database the molecules presenting a similarity metric (Tanimoto metric) of 0.8 ;
- the column of the table bearing the reference (4) corresponds to the application of a QSAR relation on a database obtained by the iterative selection method according to the invention and applied on the basis of initial data (MACCS structural descriptors) 166).
- the stopping criteria considered for the iterative process are a maximum of 5 iterations or 600 selected molecules in the starting base.
- the local and global metrics described in the previously detailed embodiment have been used; and
- the column of the table bearing reference (6) corresponds to the application of an automatic learning algorithm also commonly referred to as a "machine learning” algorithm on a database obtained by means of the iterative selection method according to the invention and applied on the basis of initial data (MACCS 166 structural descriptors).
- the stopping criteria considered for the iterative process are a maximum of 5 iterations or 600 selected molecules in the starting base.
- the local and global metrics described in the previously detailed embodiment have been used.
- Appendix 2 reflects other prediction results obtained for the AMES test, for different categories of molecules (fillers, plasticizers, oxidants, liquids, stabilizers and pyrotechnic molecules), with the selection and prediction methods according to the invention ( column "prediction" of the different tables in Appendix 2).
- column "prediction" of the different tables in Appendix 2 The same assumptions as those used in Annex 1 were considered (maximum number of iterations equal to 5, 600 molecules selected at most, local and global metrics detailed previously, MACCS 166 structural descriptors); the actual prediction step was carried out on the basis of molecules selected by the selection method according to the invention by applying a machine learning type algorithm.
- the prediction method has led to a correct prediction for all the test molecules of the charge type (ie all the percentages reported are greater than 60%), for all the liquid-tested molecules, and for the whole tested molecules of the stabilizing type;
- Annexes 3 to 5 reflect prediction results obtained via the prediction method according to the invention for other known regulatory tests (chromosome aberration test in Annex 3, UDS test in Annex 4, carcinogenicity test in Annex 5). .
- the same assumptions as those used in Appendix 2 were considered for the implementation of the processes according to the invention and the interpretation of the results presented.
- Appendix 6 compares the results obtained via the prediction method according to the invention and via another prior art prediction method known as ACD (Advanced Chemistry Development) Percepta (described in more detail on the web page https://www.acdlabs.com/products/percepta/).
- ACD Advanced Chemistry Development
- the results concerning the prediction method according to the invention were obtained from two different starting bases (referenced by "first test base” and "second test base”).
- the first test basis is the one already used to generate the results reported in Appendices 2 to 5.
- the first column of results in the table presented in Appendix 6 gives the rate of good predictions obtained via the prediction method according to the invention with respect to different molecules tested for the different tests considered.
- This first column lists the different results shown in Appendices 2 to 6 for all categories of molecules considered together, and supplements these results for other known regulatory tests (Mouse Lymphoma Test (M LA), DLT, and Reprotoxicity Test).
- AMES mutagenicity test Charge type molecules
Landscapes
- Chemical & Material Sciences (AREA)
- Crystallography & Structural Chemistry (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Computing Systems (AREA)
- Theoretical Computer Science (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Computation (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR1700668A FR3068047B1 (fr) | 2017-06-22 | 2017-06-22 | Procede et dispositif de selection d'un sous-ensemble de molecules destinees a etre utilisees pour predire au moins une propriete d'une structure moleculaire |
| PCT/FR2018/051529 WO2018234718A1 (fr) | 2017-06-22 | 2018-06-22 | Procédé et dispositif de sélection d'un sous-ensemble de molécules destinées à être utilisées pour prédire au moins une propriété d'une structure moléculaire |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP3642398A1 true EP3642398A1 (fr) | 2020-04-29 |
Family
ID=60382254
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP18749450.5A Pending EP3642398A1 (fr) | 2017-06-22 | 2018-06-22 | Procédé et dispositif de sélection d'un sous-ensemble de molécules destinées à être utilisées pour prédire au moins une propriété d'une structure moléculaire |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20230154571A1 (fr) |
| EP (1) | EP3642398A1 (fr) |
| FR (1) | FR3068047B1 (fr) |
| WO (1) | WO2018234718A1 (fr) |
Families Citing this family (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2020051714A1 (fr) * | 2018-09-13 | 2020-03-19 | Cyclica Inc. | Procédé et système de prédiction de propriétés de structures chimiques |
| CN115116553B (zh) * | 2021-03-19 | 2024-12-13 | 本源量子计算科技(合肥)股份有限公司 | 分子的参数配置方法、装置、介质及电子装置 |
| JP2022150078A (ja) * | 2021-03-26 | 2022-10-07 | 富士通株式会社 | 情報処理プログラム、情報処理装置、及び情報処理方法 |
| CN114300065B (zh) * | 2021-12-10 | 2024-12-27 | 深圳晶泰科技有限公司 | 分子设计方案的确定方法、装置、设备及存储介质 |
| WO2023102923A1 (fr) * | 2021-12-10 | 2023-06-15 | 深圳晶泰科技有限公司 | Procédé et appareil de détermination pour un programme de conception moléculaire, dispositif et support de stockage |
| US12587274B2 (en) | 2023-03-28 | 2026-03-24 | Quantum Generative Materials Llc | Satellite optimization management system based on natural language input and artificial intelligence |
| FR3147900A1 (fr) | 2023-04-17 | 2024-10-18 | Arianegroup Sas | Procédé d'évaluation d'une prédiction d'un résultat d'une propriété d'une molécule, et dispositif associé |
| US12368503B2 (en) | 2023-12-27 | 2025-07-22 | Quantum Generative Materials Llc | Intent-based satellite transmit management based on preexisting historical location and machine learning |
| US12603701B2 (en) | 2023-12-27 | 2026-04-14 | Quantum Generative Materials Llc | Distributed satellite constellation management and control system |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6185506B1 (en) * | 1996-01-26 | 2001-02-06 | Tripos, Inc. | Method for selecting an optimally diverse library of small molecules based on validated molecular structural descriptors |
| US7219020B1 (en) * | 1999-04-09 | 2007-05-15 | Axontologic, Inc. | Chemical structure similarity ranking system and computer-implemented method for same |
| US7472121B2 (en) * | 2005-12-15 | 2008-12-30 | International Business Machines Corporation | Document comparison using multiple similarity measures |
| WO2013192110A2 (fr) * | 2012-06-17 | 2013-12-27 | Openeye Scientific Software, Inc. | Calculs de similarité moléculaire sécurisés |
-
2017
- 2017-06-22 FR FR1700668A patent/FR3068047B1/fr active Active
-
2018
- 2018-06-22 WO PCT/FR2018/051529 patent/WO2018234718A1/fr not_active Ceased
- 2018-06-22 US US16/624,700 patent/US20230154571A1/en active Pending
- 2018-06-22 EP EP18749450.5A patent/EP3642398A1/fr active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| FR3068047A1 (fr) | 2018-12-28 |
| US20230154571A1 (en) | 2023-05-18 |
| WO2018234718A1 (fr) | 2018-12-27 |
| FR3068047B1 (fr) | 2021-02-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3642398A1 (fr) | Procédé et dispositif de sélection d'un sous-ensemble de molécules destinées à être utilisées pour prédire au moins une propriété d'une structure moléculaire | |
| Chen et al. | Context-aware transcript quantification from long-read RNA-seq data with Bambu | |
| Huang et al. | MSstatsTMT: statistical detection of differentially abundant proteins in experiments with isobaric labeling and multiple mixtures | |
| Lin et al. | A combinatorial approach to the peptide feature matching problem for label-free quantification | |
| Goodman | The antibody horror show: an introductory guide for the perplexed | |
| Rogers et al. | CScape-somatic: distinguishing driver and passenger point mutations in the cancer genome | |
| US20220099566A1 (en) | Deformulation techniques for deducing the composition of a material from a spectrogram | |
| Chen et al. | RankMotif++: a motif-search algorithm that accounts for relative ranks of K-mers in binding transcription factors | |
| Kavak et al. | Discovery and genotyping of novel sequence insertions in many sequenced individuals | |
| CN109994155A (zh) | 一种基因变异识别方法、装置和存储介质 | |
| Wang et al. | MMFPh: a maximal motif finder for phosphoproteomics datasets | |
| Annala et al. | A linear model for transcription factor binding affinity prediction in protein binding microarrays | |
| Kuijpers et al. | Split Pool Ligation-based Single-cell Transcriptome sequencing (SPLiT-seq) data processing pipeline comparison | |
| Vukovic et al. | Methodology of aiQSAR: a group-specific approach to QSAR modelling | |
| WO2020199337A1 (fr) | Procédé et dispositif d'identification de génovariation et support de stockage | |
| Majidian et al. | Hap10: reconstructing accurate and long polyploid haplotypes using linked reads | |
| Szucs et al. | Impact of structural similarity on the accuracy of retention time prediction | |
| Pakkir Shah et al. | The hitchhiker’s guide to statistical analysis of feature-based molecular networks from non-targeted metabolomics data | |
| Arturi et al. | MLinvitroTox reloaded for high-throughput hazard-based prioritization of high-resolution mass spectrometry data | |
| Klingberg et al. | Towards compound identification of synthetic opioids in nontargeted screening using machine learning techniques | |
| Keil et al. | SQANTI-reads: a tool for the quality assessment of long read data in multi-sample lrRNA-seq experiments | |
| Zhang et al. | VEF: a variant filtering tool based on ensemble methods | |
| Madej et al. | PyViscount: Validating False Discovery Rate Estimation Methods Via Random Search Space Partition | |
| EP3149638A1 (fr) | Procede d'identification d'une relation entre des elements biologiques | |
| Sun et al. | BPDA2d—a 2D global optimization-based Bayesian peptide detection algorithm for liquid chromatograph–mass spectrometry |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20200113 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: BA ME |
|
| RIN1 | Information on inventor provided before grant (corrected) |
Inventor name: ALLIOD, CHARLOTTE Inventor name: DENIS, ROLAND Inventor name: TERREUX, RAPHAEL Inventor name: JACOB, GUY |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| RAP3 | Party data changed (applicant data changed or rights of an application transferred) |
Owner name: UNIVERSITE CLAUDE BERNARD LYON I Owner name: CENTRE NATIONAL DE LA RECHERCHE SCIENTIFIQUE Owner name: ARIANEGROUP SAS |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20230802 |