EP4591308A1 - Bioinformatikverfahren zur bestimmung therapeutischer zielregionen - Google Patents
Bioinformatikverfahren zur bestimmung therapeutischer zielregionenInfo
- Publication number
- EP4591308A1 EP4591308A1 EP23790046.9A EP23790046A EP4591308A1 EP 4591308 A1 EP4591308 A1 EP 4591308A1 EP 23790046 A EP23790046 A EP 23790046A EP 4591308 A1 EP4591308 A1 EP 4591308A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- residues
- sequences
- protein
- target
- bonds
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B15/00—ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
- G16B15/30—Drug targeting using structural data; Docking or binding prediction
Definitions
- TITLE BIOINFORMATICAL METHOD FOR DETERMINING THERAPEUTIC TARGET REGIONS TECHNICAL FIELD
- the present application concerns a bioinformatics method for identifying reliable and durable therapeutic target regions, in order to optimize the search for new drugs, particularly antivirals.
- STATE OF THE ART RNA viruses still represent a serious public health problem. Indeed, their high mutation rate allows them to quickly acquire resistance to these treatments. To prevent the emergence of resistance, it is recommended to prioritize invariant amino acids. Indeed, mutations in highly conserved positions lead to deterioration or alteration of biological functions and could thus render the virus non-viable. However, due to their very small number, the invariant positions alone cannot constitute binding sites for a drug.
- Lao J. et al. also propose to identify pairs of covariant mutations called “lethal synthetics” (SL) [Brouillet et al., Petitjean et al.]. SLs represent mutations that are not lethal but, when combined, render the virus nonviable. These SLs have already been studied in the search for anticancer drugs [Kuiken HJ, and Beijersbergen RL] and anti-HIV agents. Lao et al. proposes a series of computational steps to identify the best target residues that, once mutated or blocked by a drug, could substantially affect the biological function of the targeted pathogen. The method proposed in Lao et al. however, has disadvantages.
- the targets identified by this method are not described in a manner sufficiently precise, which penalizes the user in his development program. For example, this method does not reveal whether the target proposed by the software has a large number of invariant residues, a determined volume or is unlikely to mutate in the future. Furthermore, it does not take into account the fact that the batch of initial sequences may not be very usable or, on the contrary, very reliable and therefore highly predictive. Finally, this method lacks crucial information to establish the relevance of the identified targets: the nature of the chemical bonds that may exist between the residues of the candidate protein and more particularly in the target region. This biochemical information would make it possible to strengthen the validity of the targets identified by the genetic approach of Lao et al. and select the most relevant ones.
- the invention aims to improve the method of Lao et al. For this, the present inventors propose to integrate into this method steps based on the 3D structure and the intramolecular chemical interactions of the target protein.
- a method of this type must be able to generate results quickly, whether on proteins from different microorganisms, or on proteins from different variants of the same microorganism. It is therefore important to be able to have a rapid and safe method for processing several proteins in parallel, making it possible in particular to take into account the 3D structures of proteins from known variants. All of these improvements make it possible to improve the reliability and interest of the method described in Lao et al., which was theoretical but difficult to use (because it was insufficiently described and poorly documented).
- the method of the present invention is more effective and more reliable than that described in Lao et al., to the extent that it has been scientifically enriched and supplemented by new criteria making it possible to describe finely, both statistically and biochemically, the targets detected. Thanks to these improvements, the method of the invention becomes essential for identifying important target regions on pathogenic organisms, with a view to allowing researchers to design the drugs of tomorrow.
- a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Identification, in a set of sequences nucleotides and polypeptides characteristic of said candidate protein, previously aligned, with invariant residues and/or pairs of lethal synthetic residues called target residues; b) Identification of at least one candidate region consisting of at least one pair of target residues identified in step a), said at least one pair comprising target residues located at a determined distance in space and being exposed to the surface of the candidate protein, preferably in a pocket; c) Determination of the advantageous chemical interactions between each residue and/or between each pair of residues within the candidate protein, from the 2D and/or 3D structure of the candidate protein; said chemical interactions advantageous being hydrophobic bonds and/or hydrogen bonds and/or salt bridges and/or bonds with negative repulsions and/or positive repulsions; d) Selection of residues linked by said advantageous chemical interactions
- the step of identifying the advantageous chemical interactions consists of a determination of the chemical interactivity of each residue within the protein and/or a determination of the chemical interactivity of each pair of residues spaced a determined distance; - the distance determined in step b) is between 2 and 8 Angstroms, preferably 5 Angstroms; - the method comprises an evaluation of the chemical interactivity of the determined bonds comprising a calculation of at least one score among: a percentage of hydrophobic bonds and/or a percentage of hydrogen bonds and/or, a percentage of salt bridges and/or a percentage of bonds with negative repulsions and/or a percentage of bonds with positive repulsions and/or a percentage of chemical bonds; - the target region is a pocket identified from a set of aligned 3D structures of the candidate protein and in which the chemical interactions are determined on a set of aligned 3D structures of the candidate protein; - the
- the invention aims to identify a therapeutic target region on the sole basis of the chemical interactions which exist between the residues of the region.
- a computer-implemented method for determining at least one therapeutic target region on a candidate protein of a target pathogenic organism, said method comprising the following steps: a) Determination of the advantageous chemical interactions between each residue and/or between each pair of residues within the candidate protein, from the 2D and/or 3D structure of the candidate protein; said advantageous chemical interactions are hydrophobic bonds and/or hydrogen bonds and/or salt bridges and/or with negative repulsions and/or positive repulsions; b) Selection of residues linked by said advantageous chemical interactions, said residues being at a distance of at most 10 Angstroms, said residues forming a therapeutic target region.
- FIG. 1 illustrates a diagram of the stages of a A computer-implemented method of determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism.
- the highlighted bonds are those corresponding to couples already defined by the technique described in Lao et al and therefore confirmed by the chemical method of the invention, connecting pairs or groups of amino acids at a suitable distance (highlighted in gray); these couples will therefore be selected within the framework of the method of the invention.
- Figure 3 highlights the amino acids of interest selected in Figure 2 within the amino acids identified after implementation of steps E1 and E2 of the method of the invention (for more details, see Figure 3 de Lao et al.) DETAILED DESCRIPTION OF THE INVENTION
- the present invention relates to a computer-implemented method for determining at least one therapeutic target region on the surface of a candidate protein of a pathogenic organism.
- target pathogenic organism we mean here any type of organism that can cause disease in a “host” (such as a human being, a plant or an animal). This organism is preferably a microorganism such as a virus, a bacteria, a parasite, a fungus, a protozoa (amoeba, sporozoa or flagellates), etc.
- a harmful macro-organism such as a worm (belonging for example to the group of helminths, platyhelminths such as trematodes or cestodes, or nemathelminthes such as Ascaris, Toxocara, or Trichuris nematodes) or a insect.
- helminths such as trematodes or cestodes
- nemathelminthes such as Ascaris, Toxocara, or Trichuris nematodes
- target pathogenic organism of tumor cells or cells infected by a pathogenic microorganism such as a virus, a bacteria, etc.
- these cells often express on their surface or in their cytoplasm (or even in their nuclei) proteins involved in maintaining the proliferation signals of these cells, resistance to cell death, escape from the immune system, angiogenesis, activating invasion and metastasis, replicative immortality, escape from growth factor suppressors, reprogramming of energy metabolism, which amplifies the disease (cancer or infection). It is therefore entirely logical to use the method of the invention to identify target regions suitable for therapeutic research, on proteins which are also involved in these disorders.
- Pathogenic organisms are essentially made up of proteins, some of which are necessary for their development, their infectivity and/or their pathogenicity. For each known pathogen, numerous proteins of this type have been identified and constitute a preferred target for researchers in the pharmaceutical industry.
- these proteins will be called “candidate proteins”, in that they have been previously identified as candidates potentially having an impact on the development, infectivity and/or pathogenicity of an organism. pathogen of interest.
- a molecule to be selected as a “drug” it must have a substantial influence, in the short, medium or long term, on the function of at least one of these candidate proteins. To do this, it must first be able to come into contact with the candidate protein (if possible, thanks to several contact zones).
- a “therapeutic target region” is therefore called a privileged contact zone between a drug and a candidate protein, said zone having been selected in such a way that the interaction between these two elements (the drug of on the one hand and the protein on the other hand) can chemically be strong, stable and significantly influence the biological function of the protein.
- a candidate protein which is known to act on the development, infectivity and/or pathogenicity of the pathogenic organism that is being studied , has been identified.
- this candidate protein must be known and well described in the literature.
- polypeptide sequences and 3D sequences must have already been characterized in the art, and be easily accessible.
- Nucleotide sequences encoding these polypeptide sequences must also be known. All these sequences are generally provided in official and openly accessible databases. The sequences of these bases, hereinafter called “BDD1”, are generally anonymized, freely accessible and uploaded by various international research units. These are for example the bases Genbank, NCBI, INSDC, EMBL, HIVDB, LANL, FLUDB, GISAID, etc., well known to those skilled in the art. In a step prior to the process of the invention, a polypeptide sequence and a reference nucleotide sequence must be selected (step E0).
- reference sequences can be chosen from a phylogenetic tree (in this case, the reference sequence is at the root of the tree representing the sequences studied) or, if the tree does not exist, by recalculating a ancestral sequence, reconstructed by one of the following three methods: parsimony, maximum likelihood or Bayesian method.
- This reference sequence can also be a consensus sequence from the batch of sequences studied, but in this case it only serves to be compared to other sequences, because a sequence Computed consensus may be a sequence that never existed, so it is not necessarily functional. Subsequently, the other known/listed nucleotide and polypeptide sequences for this candidate protein are identified in the “BDD1” databases. These additional sequences will hereinafter be called “initial sequences”.
- a protein must perform a certain number of functions which are only achievable if it adopts a certain structure in space and has the chemical radicals in the appropriate location. This is called the structure-function link.
- certain mutations change the structure of the protein and thus cause it to lose one or more functions. If these functions are essential, it follows that the virus becomes non-replicative and can therefore no longer develop.
- These mutations are mainly of three kinds: invariant positions, pairs of lethal synthetics and pairs of compensatory mutations.
- a single mutated invariant position renders the protein non-functional, when to achieve the same goal, two positions must be mutated in the case of pairs of lethal synthetics.
- a compensatory mutation pair is defined as follows: a first mutation renders the protein non-functional but a second could appear and restore the function of the protein studied. In these three cases, a strong constraint is imposed on the protein.
- these initial sequences are first processed to select invariant residues and/or pairs of lethal synthetic residues (step E1). Once these invariant residues and/or pairs of synthetic lethal residues are known, at least one candidate region is identified, based on other structural criteria (step E2). Steps E1 and E2 were described by Lao et al.
- a region containing at least two or three amino acids of interest which have been selected to be invariant residues and/or lethal synthetic covariant residues (SL) is called “candidate region from E2”. not deriving from a common ancestor, close to each other, exposed on the surface of the candidate protein and possibly in a pocket, thanks to the steps E1 and E2 of the present application, that is to say according to the method described in Lao et al.
- the method of the invention provides for determining the chemical bonds involved globally between the residues of the protein, and in particular between the residues located in the candidate region (steps E3, E4, E5 in Figure 1).
- At least one candidate region most likely to be an effective therapeutic target is identified.
- the method of the invention thus makes it possible to select, among the candidate regions obtained by following the indications of Lao et al., the therapeutic target region(s) most likely to allow the identification of effective drugs.
- one or more scores can advantageously be calculated to verify that the information resulting from each step of the method of the invention is relevant for the rest of the process and in relation to the expected results. These scores provide researchers with information on the quality and therefore reliability of the results obtained. For the sake of readability of the description, the detailed expressions of the scores are given at the end of the description.
- Step E1 consists of identifying, in a set of nucleotide and polypeptide sequences, characteristic of said candidate protein, and previously cleaned and aligned, the invariant residues and/or pairs of synthetic covariant residues lethal, which will be called, in the context of the invention, “target residues”.
- the databases used to store the sequences of pathogenic organisms may contain erroneous sequences, which, if not eliminated, could generate false results. This is why, in the method of the invention, the initial sequences must first be “cleaned” (step E11).
- This “cleaning” consists of filtering the initial sequences, for example in the following way: ⁇ Initial sequences having a length different from the sequence chosen as reference can be excluded (if the batch of initial sequences contains sufficient sequences); ⁇ If the batch of initial sequences does not contain many sequences and if the alignment seems possible (because the sequences do not contain hyper variable regions, and few poorly described regions), an alignment can be carried out to homogenize the lengths. In practice, from sequences of variable lengths, we obtain an alignment of a single length, which is greater than the length of the sequence having the greatest length. ⁇ Initial sequences with a number of mutations greater than three standard deviations from the average number of mutations per sequence are excluded.
- N-terminal and C-terminal ends of the initial sequences are identified and the sequences are oriented in the same direction, so that their ends can be superposed (for example all the sequences can be oriented from the N-end terminal towards the C-terminus).
- This cleaning step E11 also makes it possible to identify, among all the initial sequences, those which have heterogeneous lengths and/or makes it possible to exclude sequences which present unacceptable anomalies.
- anomalies are, for example, an aberrant number of mutations, an aberrant number of poorly defined amino acids, an aberrant number of missing amino acids, etc.
- This function of score f(S1QS) reflects the impact of the sequenced data on the quality of the result obtained at the end of the process. It is also recommended to interrupt the process when the number of mutations belonging to the batches of initial sequences retained after “cleaning” is too low (typically, a number of mutations not allowing statistically correct results, for example which does not allow the calculation of ⁇ ⁇ (covariants having less than 5 representatives)). In this case, it is preferable to select a new set of sequences from the BDD1 database, enrich it by downloading new sequences, or change the candidate protein. Conversely, we consider that we have a good set of initial sequences when at least 1000 sequences of acceptable quality have been identified, these sequences carrying sufficient mutations.
- acceptable quality we mean here initial sequences presenting a number of anomalies less than three standard deviations from the average number of these anomalies per sequence.
- the S7 score can be calculated, to measure the total number of initial sequences not containing an aberrant number of mutations. Such an S7 score depends on the S2, S3 and S4 scores presented above. Then, an alignment of the sequences obtained after cleaning is advantageously carried out (step E12).
- the alignment of the sequences can be implemented according to two possibilities: i) either the sequences all have the same length: in this case the alignment is generated by aligning all the sequences on their N-terminal end, ii) or they n do not all have the same length: in this case it is difficult to use multiple alignment methods, which are generally used for a smaller number of sequences. In this case, it is possible to select a sample representative of the population of sequences studied and to use the HMMER suite (http://hmmer.org) which makes it possible to generate a profile on which all the sequences can then align. one by one (using the hmmbuild and hmmcalibrate functions). This method makes it possible to align hundreds of thousands of sequences in a time of the order of a minute.
- the quality of the alignment of the sequences can be evaluated (step E12') via the calculation of one or more scores (detailed later) relating to the impact of gaps in the alignment (score S8), the redundancy of the batch of sequence (score S10), the impact of hypervariable regions (score S11), the impact of deletions and insertions (score S12), the impact of post-translational modifications (score S13), and/or the impact of the existence of different subtypes (score S14).
- the S11, S12, and S13 scores are defined based on data found in the literature.
- One or more score functions can be calculated at this stage to assess whether the set of aligned sequences after cleaning is sufficiently complete and robust to effectively predict the existence of a therapeutic target region within the selected candidate protein.
- the score function f(S3) detailed below reflects the quality of the alignment and the possible prediction. Some terms of this function describe the precision with which the initial sequences are described and have an impact on the statistical results (f(S3QS) and other terms of this function show the heterogeneity of the batch of sequences studied and have an impact on the description of the target itself f(S3SC). More precisely, the score function f(S3QS) translates the impact of the alignment on the statistical prediction of the target region and the score function f(S3SC) translates the impact of the alignment on the prediction of the target.
- the process can be interrupted to be started again from new initial sequences.
- sequences to test the user has a set of cleaned and aligned sequences called “sequences to test”.
- sequences to test the invariant residues are then identified (step E13).
- the “invariant residues” are by definition amino acids which do not change (or almost not) position within the protein, in all the test sequences studied. As sequencing methods are not 100% reliable, an average error rate of 0.3% can be applied (Cheng C. et al, 2022). Thus, in the context of the present invention, a residue is defined as “invariant” if it is present at the same position on at least 99.7% of the sequences to be tested.
- the quality of the selection of invariant residues can advantageously be evaluated by calculating the mutational richness of the sequences (step E13').
- the S18 score can be calculated (see below). Thanks to this score, it is possible to verify that the sequences on which the invariant residues have been detected are sufficiently heterogeneous. Indeed, to be able to affirm that a residue is invariant for functional reasons (mutations appeared at this position, but were not selected because they were lethal) it is necessary to be able to show that mutations appeared elsewhere on the sequence. Once the number of invariant residues is known, it is advantageous to determine the percentage of these residues by calculating the S19 score. This score makes it possible to evaluate the impact of invariant positions on the final result.
- the targets most likely to be stable in the long term are those which are the most invariant, therefore made up of the greatest number of invariant residues. It is also possible to calculate the score function ⁇ ( ⁇ ⁇ ) which makes it possible to evaluate the degree of invariance of the batch of sequences studied, and gives an idea of its long-term mutational incapacity term.
- noisy's ⁇ ⁇ defined rather in the form (where A and B are the specific amino acids found at positions i and j), which allows to take into account each residue among 20 and not only the mutated or non-mutated state of the original residue.
- a and B are the specific amino acids found at positions i and j
- the p-values can be readjusted using a method known by its English name (“false discovery rate”). The residuals are considered “dependent” or “covariant” if their p-value is below 0.05 (Noivirt O. et al.).
- step E16 it is preferable to eliminate the pairs of covariant residues which share common ancestors. This can be achieved in particular by studying the DNA sequences coding the initial sequences: after alignment of these DNA sequences, the mutated codons causing non-synonymous mutations are identified and selected. Indeed, non-synonymous mutations induce the appearance of different amino acids on a physicochemical level; while, on the contrary, synonymous mutations code for the same residue.
- a linkage disequilibrium coefficient D' known as Lewontin coefficient
- Lewontin coefficient can then be calculated with all of the recoded DNA sequence data (differentiating between synonymous and non-synonymous positions).
- this coefficient couples sharing the same ancestor, whose covariation is therefore not the consequence of functional interdependencies, can be identified. Thanks to these steps, it is possible to determine whether the covariation of the two residues identified as “covariants” in step E14 arises from the coevolution of these two residues or if it is due to the fact that they are phylogenetically linked to an ancestor. common.
- CM compensatory mutations
- SL synthetic lethals
- the S23 score can also be calculated, because it reflects the impact of the number of CMs and their strength on the result.
- the strength of a pair of covariants is defined by its ⁇ ⁇ ( ⁇ , ⁇ ). Indeed, the higher the ⁇ ⁇ ( ⁇ , ⁇ ), the higher the number of couples observed compared to the number of couples expected if there was no covariation. S22 and S23 therefore define the strength of the covariation for this precise couple.
- invariant residues and pairs of lethal synthetics are selected as part of at least one “candidate region” of the candidate protein studied. It is here possible to calculate the S9 score which reflects the impact of the variance at the 5' and 3' ends of the sequences.
- the step which consists of aligning the DNA sequences to identify the covariant residues which share a common ancestor one of the two techniques used consists of aligning all the sequences on their 5' end.
- the variance of this end reduces the chance of correctly aligning the sequences.
- the candidate region which will be selected as the best therapeutic target must also satisfy a number of other advantageous conditions.
- the amino acids that compose it may be subject to constraints due to the 3D structure of the candidate protein.
- the method of the invention contains a step E2 which evaluates the quality of the candidate regions obtained in step E1 with respect to the position of the amino acids which compose it, in relation to the 3D structure of the candidate protein.
- Step E2 therefore consists of identifying, among the regions and residues identified in step E1, at least one pair of target residues located at a close distance in space, and being exposed to the surface of the candidate protein, preferably in a pocket.
- a therapeutic target must be composed of residues that are close to each other in space.
- the target residues of the candidate region will preferably be separated by a maximum of 10 Angstroms, preferably a maximum of 5 Angstroms.
- the method of the invention advantageously requires access to the three-dimensional structures known and identified for the candidate protein. These structures are known and described in dedicated 3D structure databases, which we will call here “BDD2”. As for nucleotide and polypeptide sequences, these 3D sequences must often, prior to the method of the invention, have been processed (cleaning step E6, alignment step E7).
- 3D structures are often in pdb format. However, this format is not applied in the same way by the entire scientific community (format of the file itself, names of subunits or numbering of residues, etc.).
- a cleaning (step E6) of the structures is preferably implemented, making it possible to define a single and generalized format for all the pdb files used (standardization of the file format, the numbering of residues and atoms, the names of the substructures). units and their respective positions among others).
- dozens or even hundreds of structures of the same protein are available.
- the 3D structures identified for the same protein are preferably aligned. These are structural alignments which allow us to know to what extent the structures identified for this protein are different.
- the alignment is implemented from a reference model. We can then calculate the average difference between all these structures, which we call “RMSD”.
- RMSD average difference between all these structures
- a three-dimensional structure is defined by the position in space of each of the atoms that constitute it. If many of these positions are missing (missing data) the 3D structure is shaky in the sense that only part of these atoms have a specific place in the structure. If a lot of data is missing for each of the structures studied (knowing that the missing data is not at the same position in the protein space), it becomes difficult to make a structural alignment and even to compare the residues. found in the same position. The fact that a protein has several subunits further complicates its structure.
- S15 is, for structures, the equivalent of Sseq for sequences. We look at whether the number of 3D structures known for the candidate protein is high or not. If there are fewer than 200 known 3D structures, this number is insufficient to have a consistent average (this figure may be decreased, however, as 3D structures are increasingly reliable).
- S16 gives an idea of the variation in the resolutions of the structures. Indeed, 3D structures are determined using different techniques. The two most commonly used are X-ray diffraction and electron microscopy. A resolution threshold is defined for each of these structures, depending on the technique used. S16 evaluates this variation in resolution.
- S17 shows the structural heterogeneity of the batch of structures studied. Each of the atoms of the structure is aligned with the corresponding atom of the following structure.
- this function describes the precision with which 3D structures are described and have an impact on the statistical results ⁇ ⁇ ⁇ ⁇ ⁇ and other terms of this function show the heterogeneity of the batch of sequences studied and have a impact on the description of the target itself More precisely, the score function ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ reflects the impact of the alignment on the statistical prediction of the target region and the score function reflects the impact of alignment on target prediction. From the 3D coordinates of all the atoms of the residues in space, the pairs of residues close in space, that is to say separated by less than 5 or 10 Angstroms, are selected (step E21). Furthermore, the accessibility of target residues and/or their exposure to the protein surface must be taken into account.
- an effective therapeutic target must not be buried in the 3D structure of the protein, otherwise the drug will not be able to reach it.
- the residues which are buried in the protein are distinguished from those which are exposed on its surface and therefore accessible (to do this, it is possible to use the ASA program for example ).
- the ASA program for example.
- candidate regions containing at least two, preferably at least three accessible residues are selected (step E22).
- scores reflecting accessibility to evaluate the possibility of obtaining sufficient candidate region(s) (step E22').
- the S32 score evaluates the percentage of accessible positions and the S33 score evaluates the percentage of accessible residues.
- step E23 it may be advantageous to evaluate whether the candidate region selected previously has a 3D structure comparable to a “pocket” (step E23).
- structure prediction software for example the Fpocket software (Le Guilloux et al.), in which, to take into account the existence of small and large pockets, it is preferable to reduce the minimum and maximum radii of the alpha spheres to 2.5 ⁇ and 4 ⁇ respectively.
- the Fpocket software (Le Guilloux et al., 2009) makes it possible to determine all the pockets, whose cardinality is ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ h ⁇ ⁇ ⁇ on the surface of a protein (by providing a three-dimensional structure in entrance).
- the volume of a pocket capable of housing a small drug molecule preferably meets the following constraints: 60 ⁇ ⁇ ⁇ ⁇ h ⁇ ⁇ 500 ⁇ ⁇ ( ⁇ h ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ )
- We can calculate the percentage of pocket meeting this criterion for the protein studied ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ h ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ h ⁇ ⁇ ⁇ .
- the networks of invariant or covariant SL residues close in space, which are exposed on the surface of the candidate protein and preferably in a pocket having a volume between 60 ⁇ 3 and 500 ⁇ 3, will finally be retained .
- Networks of at least five target residues will preferably be retained.
- the invariant or covariant residues SL can be represented in the form of graphs (step E8). Indeed, from the detected networks, groups of interdependent residues are formed using mathematical graphs. In these graphs are integrated the invariant residues and the pairs of lethal synthetic residues which form invariance groups.
- the nodes of the graph are the residues, the edges define the link between the residues (lethal or invariant synthetics, in this case they are linked to all) and only exists if the two residues are positioned less than 10 Angstroms and on the surface of the candidate protein.
- this type of graph can be visualized using the free Graphviz software. These graphs are a way to evaluate the quality of the identified regions.
- Step E3 consists of determining whether there are advantageous chemical interactions between the residues and/or between the pairs of residues within the candidate protein, from the 3D structure of the candidate protein.
- This step E3 can alternatively be limited to the analysis of the advantageous chemical interactions existing between the residues and/or between the pairs of residues selected within the candidate region obtained in step E2, from the 3D structure of this region. candidate. Genetics makes it possible to functionally detect the impact of microscopic physicochemical changes occurring at the residue level. Thus the fact of being invariant (or of being part of an invariance group) is the consequence of the physico-chemical quality of one or more residues (the selection pressure imposes the maintenance of this (s) residue ( s) at this position). Conversely, if the physico-chemistry of a residue is changed (for example by a mutation or by the external environment), the function and/or structure of the region may be affected.
- the additional steps E3 and E4 therefore make it possible to highlight pairs of amino acids having a chemical interaction influencing their function, or likely to evolve during modifications in the environment of the protein (pH, temperature, ionic strength, etc.).
- Protein stability is in fact often dependent on pH, and varies depending on the subtype and/or origin of the host organism. For example, human influenza virus hemagglutinin (HA) is more stable than avian virus hemagglutinin (Galloway SE et al.).
- mutations can lead to changes in the chemical interaction network, and stabilize or destabilize this protein (Byrd-Leotis L. et al.).
- the intraprotein physicochemical network consists of several types of interactions, such as hydrophobic, hydrogen interactions, as well as salt bridges, negative repulsive and positive repulsive interactions, among others (Dyson HJ et al.; Hubbard RE & Kamran Haider M.; Sticke DF et al.; Barlow DJ & Thornton J. M; Harrison JS et al.). These interactions allow the protein to have a very particular shape: this is why they are important. If they were not there, the structure of the protein and surely its function would be modified. This is why these regions rich in these chemical bonds are important to determine, in the context of the method of the invention.
- the process described here therefore contains a step of characterizing the chemical interactions existing between all the atoms of the protein, or at least those present in the candidate region selected previously. This step aims in particular to identify the following interactions (Hubbard RE & Kamran Haider M.; Barlow DJ & Thornton J. M; Harrison JS et al.; Donald JE et al.; Freitas RF de & Schapira M.; Onofrio A. et al.
- - Hydrogen bonds consisting of a proton donor atom (OG, OG1, NE2, ND2, ND1, NE2, NZ, NE, NH1, NH2, OH, NE1) belonging to a proton donor amino acid (SER, THR, GLN, ASN, HIS, LYS, ARG, TYR, TRP) and a proton acceptor atom (OG, OG1, OE1, OE2, OD1, OD2, ND1, NE2, OH) belonging to a proton acceptor amino acid ( SER, THR, GLU, ASP, GLN, ASN, HIS, TYR).
- the atoms presenting this type of bond will be selected, when their distance is less than 3.5 ⁇ (Hubbard RE & Kamran Haider M.; Sticke DF et al.).
- These hydrogen bonds can consist of interaction between two side chains or between a side chain and the main chain of the protein.
- the proton donors (N) of the main chain or the proton acceptors of this same chain can belong to any amino acid of the main chain.
- Electrostatic bonds of two kinds: o Attractive bonds (salt bridges or ionic interaction) when a cation (NE, NH1, NH2, NZ, NE2, ND1) of a positively charged residue (ARG, LYS, HIS) , especially when it is at a distance less than 4 ⁇ from an anion (OD1, OD2, OE1, OE2) carried by a negatively charged residue (ASP, GLU).
- o Repulsive bonds made up of two atoms with identical charges (-/-) or (+/+), especially when they are separated by a distance of less than 5 ⁇ (Barlow DJ & Thornton JM; Harrison JS et al.) .
- the method of the invention therefore advantageously contains a step of calculating the distance between each atom of the protein, for example from a pdb, SwissProt, uniprot or Modbase file.
- the distances between atoms belonging to the same position are not calculated unless they belong to different protomers.
- atoms separated by a distance of at most 5 ⁇ are selected.
- This step is preferably carried out before listing the chemical interactions between the atoms.
- only chemical interactions linking nearby atoms need to be taken into consideration.
- step E3' It is in particular possible to calculate the percentage of hydrophobic bonds and/or the percentage of hydrogen bonds and/or the percentage of salt bridges and/or the percentage of negative repulsions and/or the percentage of positive repulsions and/or the percentage of bonds chemicals (step E3'). It is possible to also calculate a score function ⁇ ( ⁇ ⁇ ) which gives an image of the overall chemical interactions of the protein studied and is calculated (step E3'). Through all these steps, the residues linked by said advantageous chemical interactions, and being at a distance such that these bonds influence the function of these residues, are selected. This is step E4 in Figure 1. At this stage, it is possible to represent the different chemical bonds in the form of graphs (step E8'). Indeed, from the previously detected connections, mathematical graphs are formed.
- Step E5 makes it possible to identify the best therapeutic target region(s) making it possible to generate a potential drug by crossing the results obtained at the end of step E2 with those obtained at the end of step E4 .
- the quality of the selected targets can advantageously be evaluated by calculating several scoring functions. For example, it is possible to calculate the invariance of each target region by the function ⁇ ( ⁇ ⁇ ) . The invariance of each target region ensures that it will be stable over time.
- the function ⁇ ( ⁇ ⁇ ) makes it possible to obtain a score between 0 and 1 for each target S
- these two score functions make it possible to evaluate the effectiveness of the target region.
- the target regions can advantageously be represented in the form of graphs (step E8''). Indeed, from the previously detected pairs, groups of interdependent residues by means of mathematical graphs are formed. In these graphs are integrated the invariant residues and the pairs of lethal synthetic residues which form invariance groups, and which are linked by advantageous chemical bonds.
- the nodes of the graph are the selected positions/residues, the edges define the link between the positions (lethal, invariant synthetics, or advantageous chemical interaction) and only exists if the two residues at this position are positioned at less than 10 Angstroms and on the surface of the candidate protein. Additionally, this type of graph can be visualized using the free Graphviz software. These graphs are an additional way to evaluate the quality of the identified regions.
- the complexity of the graphs can be calculated (step E8''') by the score S26 as well as the number of connected graphs by the score S27. Such complexity is useful: if the graph is complex, this reflects a significant number of targets. If it is very complex, there may be intersections between non-empty targets.
- ⁇ ⁇ Sequence length heterogeneity.
- ⁇ the mean and ⁇ the standard deviation of the amino acid lengths of the sequences of the sample studied.
- ⁇ ⁇ > ⁇ ⁇ ⁇ ⁇ 0
- ⁇ ⁇ 1 ⁇ ( ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ )
- S4: ⁇ ⁇ Sequences with an aberrant number of gaps.
- ⁇ ⁇ 1 ⁇ ( ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ) The closer the S4 score is to 1, the more aberrant the number of gaps.
- ⁇ ⁇ ⁇ allows you to evaluate whether the input data is sufficient and robust for the prediction to be made:
- ⁇ ⁇ ⁇ ⁇ ⁇ as the total number of files studied and ⁇ the mean and ⁇ the standard deviation of the number of amino acids per file of the sample studied, in comparison to a reference file. This reference file can be chosen in two ways.
- ⁇ ⁇ is the total number of sequences in the sample excluding sequences with an aberrant number of mutations ⁇ ⁇ ⁇ ⁇ N, ⁇ ( ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ) > 1000, ⁇ ( ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ) ⁇ 1000, S8: ⁇ ⁇ : evaluates the impact of gaps in the alignment. Let us consider ⁇ the mean and ⁇ the standard deviation of the number of gaps calculated by sequences of the sample studied after alignment, in comparison to the same sequence before the alignment.
- ⁇ ⁇ and ⁇ ⁇ impact of the variance of the 5' and 3' ends.
- 5 amino acids (AA) from the 5' end and 5 AA from the 3' end are studied.
- the polypeptide sequences corresponding to the candidate protein can be very heterogeneous or not very heterogeneous in terms of the mutations they carry. An extreme situation may be that out of X sequences, they would all have the same AA sequence. We would then admit that this batch is very homogeneous, that the redundancy is absolute, and therefore would have a score equal to 0. Conversely, when the polypeptide sequences are very different taken two by two, then there is very little redundancy , and the score tends towards 1.
- ⁇ ⁇ 1 ⁇ ⁇ ⁇ ⁇ ( ⁇ ⁇ ): quality of the alignment of the primary sequences + ⁇ ⁇ + ( ⁇ ⁇ + ⁇ ⁇ + ⁇ ⁇ ) ⁇ 3 ⁇ 2) ⁇ 17 ⁇ ⁇ ⁇ : quality of the statistical prediction of alignment ⁇ ⁇ ⁇ : Impact of alignment on target prediction S18: ⁇ ⁇ : Mutational richness of sequences. Let us consider ⁇ the mean and ⁇ the standard deviation of the number of mutations calculated by sequences of the sample of sequences having a number of mutations ⁇ ⁇ ⁇ , in comparison to a reference sequence.
- a graph G is defined by a couple (S,A) with S a finite set of vertices, and A a finite set of pairs of vertices (si, sj) in S2.
- a couple is therefore a pair of vertices connected by an edge. This involves studying the invariant and SL residues involved in several pair relationships (close in space and located on the surface of the protein, being either invariant or involved in a relationship of lethal synthetics). From the link graph, we define this score as the average of the number of edges (a) per node ( ⁇ ) ⁇ S28: ⁇ ⁇ : Number of connected graphs.
- a graph is connected if each pair of vertices is connected by an edge.
- S28 ⁇ ⁇ : residues close in space (5 angstroms). This score corresponds to the proportion of pairs of residues located less than 5 angstroms from each other ( ⁇ ). From a reference pdb file, the calculation of the distance between the 2 ⁇ of a pair of residues is carried out for all pairs of residues. The number of pairs being less than 5 angstroms apart is evaluated and noted as ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
- S29 ⁇ ⁇ : residues close in space (10 angstroms). This score corresponds to the proportion of pairs of residues located less than 10 angstroms from each other ( ⁇ ).
- the number of pairs being separated by less than 5 angstroms is evaluated and noted as ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
- ⁇ ⁇ ⁇ : invariant and close SL in space (10 angstroms) From a reference pdb file, the calculation of the distance between the 2 ⁇ of a residue pair is performed for all residue pairs.
- the number of pairs separated by less than 10 angstroms is denoted ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ .
- ⁇ ( ⁇ ⁇ ) score between 0 and 1, for each target ⁇ ( ⁇ ⁇ ⁇ ( ⁇ ⁇ ⁇ 1 ) ) ⁇ 2))) ⁇ 4
- This scoring function allows us to give a score to each of the targets determined by our software and defines the “druggability” (in English “druggability”) of this target.
- “Druggability” means being able to effectively bind a small molecule which would have therapeutic capabilities and therefore represent a future drug. To this definition is added that this target would leave little or no possibility of therapeutic escape.
- the method of the invention was implemented on the hemagglutinin (HA) protein of the influenza virus.
- Galloway, SE, Reed, ML, Russell, CJ & Steinhauer, DA Influenza HA subtypes demonstrate divergent phenotypes for cleavage activation and pH of fusion: implications for host range and adaptation.
- PKAD a database of experimentally measured pKa values of ionizable groups in proteins. Database 2019, apel024 (2019). Petitjean M., Badel A., Veitia RA, Vanet A., Synthetic lethals in HIV: ways to avoid drug resistance: Running title: Preventing HIV resistance. Biol Direct. 2015 Apr 17;10:17 Sticke, DF, Presta, LG, Dill, KA & Rose, GD Hydrogen bonding in globular proteins. Journal of Molecular Biology 226, 1143–1159 (1992).
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Engineering & Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Biophysics (AREA)
- Pharmacology & Pharmacy (AREA)
- Crystallography & Structural Chemistry (AREA)
- Medicinal Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Theoretical Computer Science (AREA)
- Investigating Or Analysing Biological Materials (AREA)
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR2209473A FR3139935A1 (fr) | 2022-09-20 | 2022-09-20 | Procédé bio-informatique de détermination de régions cibles thérapeutiques |
| PCT/FR2023/051429 WO2024062188A1 (fr) | 2022-09-20 | 2023-09-19 | Procede bio-informatique de determination de regions cibles therapeutiques |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4591308A1 true EP4591308A1 (de) | 2025-07-30 |
Family
ID=84887621
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23790046.9A Pending EP4591308A1 (de) | 2022-09-20 | 2023-09-19 | Bioinformatikverfahren zur bestimmung therapeutischer zielregionen |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4591308A1 (de) |
| JP (1) | JP2025536116A (de) |
| FR (1) | FR3139935A1 (de) |
| WO (1) | WO2024062188A1 (de) |
-
2022
- 2022-09-20 FR FR2209473A patent/FR3139935A1/fr active Pending
-
2023
- 2023-09-19 JP JP2025517544A patent/JP2025536116A/ja active Pending
- 2023-09-19 WO PCT/FR2023/051429 patent/WO2024062188A1/fr not_active Ceased
- 2023-09-19 EP EP23790046.9A patent/EP4591308A1/de active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| FR3139935A1 (fr) | 2024-03-22 |
| JP2025536116A (ja) | 2025-10-31 |
| WO2024062188A1 (fr) | 2024-03-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Da et al. | Structural protein–ligand interaction fingerprints (SPLIF) for structure-based virtual screening: method and benchmark study | |
| Mongan et al. | Generalized Born model with a simple, robust molecular volume correction | |
| Verdonk et al. | Virtual screening using protein− ligand docking: avoiding artificial enrichment | |
| Nguyen et al. | ATG4 family proteins drive phagophore growth independently of the LC3/GABARAP lipidation system | |
| Sun et al. | A proactive genotype-to-patient-phenotype map for cystathionine beta-synthase | |
| McGaughey et al. | Comparison of topological, shape, and docking methods in virtual screening | |
| Bissantz et al. | Protein-based virtual screening of chemical databases. 1. Evaluation of different docking/scoring combinations | |
| Hannenhalli et al. | Analysis and prediction of functional sub-types from protein sequence alignments | |
| Lin et al. | Short carboxylic acid–carboxylate hydrogen bonds can have fully localized protons | |
| Çınaroğlu et al. | Comparative assessment of seven docking programs on a nonredundant metalloprotein subset of the PDBbind refined | |
| Greenidge et al. | Improving docking results via reranking of ensembles of ligand poses in multiple X-ray protein conformations with MM-GBSA | |
| Horvatovich et al. | Quest for missing proteins: update 2015 on chromosome-centric human proteome project | |
| Rosta et al. | Artificial reaction coordinate “tunneling” in free‐energy calculations: The catalytic reaction of RNase H | |
| Cascella et al. | Evolutionarily conserved functional mechanics across pepsin-like and retroviral aspartic proteases | |
| Rivas et al. | Population genomics of parallel evolution in gene expression and gene sequence during ecological adaptation | |
| Huang et al. | Efficient estimation of binding free energies between peptides and an MHC class II molecule using coarse‐grained molecular dynamics simulations with a weighted histogram analysis method | |
| Du et al. | The effect of alignment uncertainty, substitution models and priors in building and dating the mammal tree of life | |
| Jacobsen et al. | Inhibition mechanism of antimalarial drugs targeting the cytochrome bc1 complex | |
| Díez et al. | Integration of proteomics and transcriptomics data sets for the analysis of a lymphoma B-cell line in the context of the chromosome-centric human proteome project | |
| Rahimian et al. | SARS2Mutant: SARS-CoV-2 amino-acid mutation atlas database | |
| Schwans et al. | Experimental and computational mutagenesis to investigate the positioning of a general base within an enzyme active site | |
| Coimbra et al. | Different Enzyme Conformations Induce Different Mechanistic Traits in HIV‐1 Protease | |
| Akgüller et al. | Network-Medicine-Guided Drug Repurposing for Alzheimer’s Disease: A Multi-Dimensional Systems Pharmacology Approach | |
| Liebermann et al. | From microstates to macrostates in the conformational dynamics of GroEL: a single-molecule Förster resonance energy transfer study | |
| Gurumoorthy et al. | Disordered domain shifts the conformational ensemble of the folded regulatory domain of the multidomain oncoprotein c-Src |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250410 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) |