WO2014104401A1 - 結合部位同定装置、結合部位同定方法および結合部位同定プログラム - Google Patents

結合部位同定装置、結合部位同定方法および結合部位同定プログラム Download PDF

Info

Publication number
WO2014104401A1
WO2014104401A1 PCT/JP2013/085330 JP2013085330W WO2014104401A1 WO 2014104401 A1 WO2014104401 A1 WO 2014104401A1 JP 2013085330 W JP2013085330 W JP 2013085330W WO 2014104401 A1 WO2014104401 A1 WO 2014104401A1
Authority
WO
WIPO (PCT)
Prior art keywords
binding site
binding
substance
search
target substance
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2013/085330
Other languages
English (en)
French (fr)
Inventor
優哉 小玉
恒 竹内
高橋 栄夫
嶋田 一夫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ajinomoto Co Inc
National Institute of Advanced Industrial Science and Technology AIST
Original Assignee
Ajinomoto Co Inc
National Institute of Advanced Industrial Science and Technology AIST
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ajinomoto Co Inc, National Institute of Advanced Industrial Science and Technology AIST filed Critical Ajinomoto Co Inc
Priority to JP2014554633A priority Critical patent/JP6344768B2/ja
Publication of WO2014104401A1 publication Critical patent/WO2014104401A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B15/00ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
    • G16B15/30Drug targeting using structural data; Docking or binding prediction
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6803General methods of protein analysis not limited to specific proteins or families of proteins
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N24/00Investigating or analyzing materials by the use of nuclear magnetic resonance, electron paramagnetic resonance or other spin effects
    • G01N24/08Investigating or analyzing materials by the use of nuclear magnetic resonance, electron paramagnetic resonance or other spin effects by using nuclear magnetic resonance
    • G01N24/088Assessment or manipulation of a chemical or biochemical reaction, e.g. verification whether a chemical reaction occurred or whether a ligand binds to a receptor in drug screening or assessing reaction kinetics
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01RMEASURING ELECTRIC VARIABLES; MEASURING MAGNETIC VARIABLES
    • G01R33/00Arrangements or instruments for measuring magnetic variables
    • G01R33/20Arrangements or instruments for measuring magnetic variables involving magnetic resonance
    • G01R33/44Arrangements or instruments for measuring magnetic variables involving magnetic resonance using nuclear magnetic resonance [NMR]
    • G01R33/46NMR spectroscopy
    • G01R33/465NMR spectroscopy applied to biological material, e.g. in vitro testing
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B15/00ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01RMEASURING ELECTRIC VARIABLES; MEASURING MAGNETIC VARIABLES
    • G01R33/00Arrangements or instruments for measuring magnetic variables
    • G01R33/20Arrangements or instruments for measuring magnetic variables involving magnetic resonance
    • G01R33/44Arrangements or instruments for measuring magnetic variables involving magnetic resonance using nuclear magnetic resonance [NMR]
    • G01R33/46NMR spectroscopy
    • G01R33/4633Sequences for multi-dimensional NMR

Definitions

  • the present invention relates to a binding site identification device, a binding site identification method, and a binding site identification program.
  • Non-Patent Document 1 discloses a technique for identifying a binding site of the ligand in the protein when the ligand is bound to the protein. Specifically, Non-Patent Document 1 describes the three-dimensional structure of a protein and the experimental results regarding the number of presumed multiple amino acid residues that are considered to be involved in binding. A technique is disclosed in which a part that matches the experimental result is searched from the surface for the number of the plurality of types of amino acid residues, and the searched part is identified as a binding site.
  • Patent Document 1 can be cited as another prior art document.
  • paragraph 0077 of Patent Document 1 it is described that the position and sequence of a binding domain are determined by comparing a known amino acid sequence with an amino acid identified by NMR (Nuclear Magnetic Resonance).
  • the small molecule when targeting a small molecule as a ligand, the small molecule often binds to an amino acid residue in the middle of a protein pocket. Therefore, the amino acid residues that bind to the small molecule are not necessarily located at close positions in the distance along the surface of the protein.
  • Non-Patent Document 1 a binding site is searched along the surface of the protein, that is, only the surface of the protein is used as a search range. Therefore, when targeting a small molecule, an appropriate binding site can be identified. There was a problem that it was difficult.
  • the present invention has been made in view of the above problems, and a binding site identification device, a binding site identification method, and a binding site identification program capable of identifying an appropriate binding site even for a binding substance such as a small molecule.
  • the purpose is to provide.
  • a binding site identification apparatus is a binding site identification apparatus including a control unit that identifies a binding site of a binding substance in a target substance, The control unit, based on the three-dimensional structure of the target substance and the acquisition result regarding the number of substances considered to be involved in the binding of the target substance, from the search space including the surface of the target substance, the number of the target substance It is characterized by comprising an identifying means for searching for a place that is considered to be coincident with the acquisition result and identifying the binding site based on the searched place.
  • the binding site identification apparatus is characterized in that, in the binding site identification apparatus, the size of the search space is set based on the size of the binding substance.
  • the binding site identification apparatus is characterized in that, in the binding site identification apparatus, the size of the binding substance is a volume of the binding substance.
  • the size of the search space is set based on a radius of a minimum sphere including the binding substance.
  • the binding site identification device is characterized in that, in the binding site identification device, the shape of the search space is a sphere, an ellipsoid or a polyhedron.
  • the identification unit is set by a search point setting unit that sets a search point based on the three-dimensional structure and the search point setting unit.
  • a search space setting means for setting the search space based on the search point, and the number of the target substances existing in the search space set by the search space setting means based on the three-dimensional structure.
  • a substance number calculation means, a degree evaluation means for evaluating the degree of coincidence or mismatch between the calculation result obtained by the substance number calculation means and the acquisition result, and the degree evaluated by the degree evaluation means is a predetermined condition.
  • Search point determination means for determining the search point that satisfies the conditions, and search point reference identification means for identifying the binding site based on the search point determined by the search point determination means. Characterized in that it comprises.
  • the binding site identification device is the binding site identification device, wherein the degree evaluation means includes the number of the target substances included in the calculation result and the number of the target substances included in the acquisition result.
  • the absolute value of the difference is calculated as the degree
  • the predetermined condition is a condition that the absolute value is not more than a predetermined value.
  • the degree evaluation means when the acquisition result includes a plurality of patterns of the target substance, the degree evaluation means The absolute value is calculated, and the calculated sum of the absolute values is calculated as the degree, and the predetermined condition is a condition that the total is less than or equal to a predetermined value.
  • the binding site identification device is the binding site identification device, wherein the search point setting means may bind molecules from the viewpoint of interaction energy based on the three-dimensional structure. , And the detected point is set as the search point.
  • the binding site identification device is the above binding site identification device, wherein the identification unit is configured to generate noise based on the size of the binding substance from the search point determined by the search point determination unit.
  • the target substance in the binding site identification device, includes the number of substances considered to be involved in the binding of the target substance included in the acquisition result. It is calculated based on a threshold value defined in consideration of a ratio between the number of the target substances and the number of the target substances present in the binding site.
  • the binding site identification apparatus is the binding site identification apparatus, wherein the target substance is a protein, the binding substance binds to the protein, and the target substance is one type or a plurality of types. It is characterized by being an amino acid.
  • a binding site identification method is a binding site identification method that is executed by an information processing apparatus including a control unit that identifies a binding site of a binding substance in a target substance, and is executed by the control unit. From the search space including the surface of the target substance based on the three-dimensional structure of the target substance and the acquisition result regarding the number of substances considered to be involved in the binding of the target substance, the number of the target substance It includes an identification step of searching for a place that is considered to match the acquisition result and identifying the binding site based on the searched place.
  • the binding site identification program is a binding site identification program for identifying a binding site of a binding substance in a target substance and causing the information processing apparatus to include the control unit to execute the program.
  • the number of the target substances from the search space including the surface of the target substance based on the three-dimensional structure of the target substance and the acquisition result regarding the number of the target substances considered to be involved in the binding.
  • the method includes a step of searching for a place that is considered to be coincident with the acquisition result and identifying the binding site based on the searched place.
  • a recording medium is a non-transitory computer-readable recording medium, and a programmed instruction for causing an information processing apparatus to execute the binding site identification method (the binding site identification program described above). ).
  • the information processing apparatus is an information processing apparatus (search point setting apparatus) including a control unit, wherein the control unit sets a search point based on a three-dimensional structure of a target substance. And the search point set by the search point setting means is used when setting a search space including the surface of the target substance.
  • An information processing apparatus is an information processing apparatus (search space setting apparatus) including a control unit, and the control unit is configured to detect a surface of a target substance with reference to a preset search point.
  • a search space setting means for setting a search space including the search space set by the search space setting means when calculating the number of target substances existing in the search space based on the three-dimensional structure of the target substance The search point set in advance is set based on the three-dimensional structure.
  • the information processing apparatus is an information processing apparatus (substance number calculation apparatus) including a control unit, and the control unit is arranged in a preset search space based on the three-dimensional structure of the target substance.
  • a substance number calculating means for calculating the number of target substances present, and the calculation result obtained by the substance number calculating means is considered to be involved in the binding of the target substance and the binding substance for the target substance This is used when evaluating the degree of coincidence or mismatch between the number-related acquisition result and the calculation result, and the preset search space is set with reference to a search point set based on the three-dimensional structure And including a surface of the target substance.
  • the information processing apparatus is an information processing apparatus (degree evaluation apparatus) including a control unit, and the control unit includes a preset calculation result, a target substance and a binding substance for a target substance.
  • a degree evaluation means for evaluating the degree of coincidence or inconsistency with the obtained results relating to the number of things that are considered to be involved in the binding, and the degree evaluated by the degree evaluation means is based on the three-dimensional structure of the target substance It is used when determining a search point whose degree satisfies a predetermined condition from set search points, and the preset calculation result is based on a search point set based on the three-dimensional structure.
  • the number of the target substances existing in the search space including the surface of the target substance set as is calculated based on the three-dimensional structure.
  • the information processing apparatus is an information processing apparatus (search point determination apparatus) provided with a control unit, and the control unit has a predetermined degree set from a preset search point.
  • Search point determination means for determining a search point satisfying a condition is provided, and the search point determined by the search point determination means is used when identifying a binding site of a binding substance in a target substance based on the search point
  • the preset search point is set based on the three-dimensional structure of the target substance, and the preset degree is set based on the preset search point.
  • the calculation result calculated based on the three-dimensional structure with respect to the number of target substances existing in the search space including the surface of the target substance, and the target substance is involved in the binding of the target substance and the binding substance. And it shall have been evaluated for matching or the degree of mismatch between the results acquired for the number of those considered, characterized by.
  • the information processing apparatus is an information processing apparatus (search point reference identification apparatus) including a control unit, and the control unit is configured to bind a substance in a target substance based on a preset search point.
  • the degree of coincidence or mismatch is determined to satisfy a predetermined condition.
  • the information processing apparatus is an information processing apparatus (noise removal apparatus) including a control unit, and the control unit is configured to determine a size of a binding substance that binds to a target substance from a preset search point.
  • the search point remaining after the noise removal unit is executed is based on the search point, and the binding substance in the target substance is provided.
  • the search point set in advance includes the surface of the target substance set with reference to the search point set based on the three-dimensional structure of the target substance.
  • the calculation result calculated based on the three-dimensional structure with respect to the number of target substances existing in the search space and the target substance are considered to be involved in the binding between the target substance and the binding substance. It degree of match or mismatch between the acquisition results for the number of those that have been determined to satisfy a predetermined condition, characterized by.
  • the binding site identification system includes a search point setting unit that sets a search point based on a three-dimensional structure of a target substance, and the target that is based on the search point set by the search point setting unit.
  • Search space setting means for setting a search space including the surface of the substance, and a substance number calculation means for calculating the number of target substances existing in the search space set by the search space setting means based on the three-dimensional structure;
  • the number of target substances is acquired from the search space including the surface of the target substance, based on the three-dimensional structure of the target substance and the acquisition result regarding the number of substances considered to be involved in the binding of the target substance.
  • the part considered to be coincident with the result is searched, and the binding site is identified based on the searched part.
  • the size of the search space is set based on the size of the binding substance.
  • the size of the binding substance is the volume of the binding substance.
  • the size of the search space is set based on the radius of the smallest sphere that contains the binding substance.
  • the shape of the search space is a sphere, an ellipsoid, or a polyhedron.
  • the search point is set based on the three-dimensional structure
  • the search space based on the set search point is set, and the number of target substances existing in the set search space based on the three-dimensional structure
  • the degree of match or mismatch between the obtained calculation result and the obtained result is evaluated, a search point whose evaluated degree satisfies a predetermined condition is determined, and a binding site is determined based on the determined search point. Identify. Thereby, there exists an effect that a suitable binding site can be identified more reliably.
  • the absolute value of the difference between the number of target substances included in the calculation result and the number of target substances included in the acquisition result is calculated as a degree
  • the predetermined condition is that the absolute value is equal to or less than the predetermined value.
  • the absolute value is calculated for each pattern, the sum of the calculated absolute values is calculated as a degree, and the predetermined condition is The condition is that the sum is equal to or less than a predetermined value.
  • a point where a molecule may be bound is detected, and the detected point is set as a search point.
  • the search points there is an effect that it is possible to set a search point that may be considered to coincide with the acquisition result with respect to the number of target substances.
  • the number of grids considered in the calculation of the present invention can be reduced and the calculation time can be shortened.
  • the binding sites are determined based on the search points remaining after the noise deletion is performed. Identify. Thereby, there exists an effect that a suitable binding site can be identified more reliably.
  • the number of substances of interest that are considered to be involved in the binding included in the acquisition result is the ratio of the number of substances of interest contained in the target substance to the number of substances of interest present in the binding site. It is calculated based on a threshold value defined in consideration. Thereby, there exists an effect that an appropriate binding site can be identified based on a more accurate acquisition result.
  • the target substance is a protein
  • the binding substance binds to the protein
  • the target substance is one kind or a plurality of kinds of amino acids.
  • FIG. 1 is a diagram showing an outline of the present embodiment.
  • FIG. 2 is a block diagram illustrating an example of the configuration of the binding site identification apparatus 1.
  • FIG. 3 is a flowchart illustrating an example of main processing performed in the binding site identification apparatus 1.
  • FIG. 4 is a diagram showing an example of the three-dimensional structure of MAPK14 and an inhibitor.
  • FIG. 5 is a diagram showing an example of the results of an inhibitor titration experiment.
  • FIG. 6 is a diagram showing an example of a standardized chemical shift change amount for each signal derived from six types of amino acid residue species.
  • FIG. 7 is a diagram illustrating an example of a distribution of weighted average values.
  • FIG. 8 is a diagram illustrating an example of a processing result of the search point setting unit 10a1.
  • FIG. 8 is a diagram illustrating an example of a processing result of the search point setting unit 10a1.
  • FIG. 9 is a diagram illustrating an example of a processing result of the degree evaluation unit 10a4.
  • FIG. 10 is a diagram illustrating an example of a processing result of the search point determination unit 10a5.
  • FIG. 11 is a diagram illustrating an example of a processing result of the noise deletion unit 10a6.
  • FIG. 12 is a diagram illustrating an example of a processing result of the search point reference identification unit 10a7.
  • FIG. 13 is a diagram showing an example in which the identification result of the binding site shown in FIG. 12 is superimposed on the complex co-crystal structure of MAPK14 and an inhibitor.
  • FIG. 14 is a diagram showing an example of a complex structure of MAPK14 and lipid.
  • FIG. 15 is a diagram showing an example of the results of a decanoic acid titration experiment.
  • FIG. 16 is a diagram showing an example of a normalized chemical shift change amount for each signal derived from six types of amino acid residue species.
  • FIG. 17 is a diagram illustrating an example of a processing result of the search point determination unit 10a5.
  • FIG. 18 is a diagram illustrating an example of a processing result of the noise deletion unit 10a6 and a processing result of the search point reference identification unit 10a7.
  • FIG. 19 is a diagram showing an example in which the identification result of the binding site shown in FIG. 18 is superposed on the complex structure of MAPK14 and ⁇ -OG.
  • FIG. 19 is a diagram showing an example in which the identification result of the binding site shown in FIG. 18 is superposed on the complex structure of MAPK14 and ⁇ -OG.
  • FIG. 20 is a diagram showing an example of the results of titration experiment and mutation experiment of decanoic acid and ⁇ -OG.
  • FIG. 21 is a diagram showing an example of an apo structure used in the binding site identification apparatus 1 in which the identification result of the binding site is shown and a complex co-crystal structure of Factor Xa and an inhibitor are superimposed. It is.
  • FIG. 22 shows an example in which the apo structure used in the binding site identification apparatus 1 showing the identification result of the binding site and the complex co-crystal structure of Galectin-9 N-domain and lactose are superimposed.
  • FIG. 23 the apo structure used in the binding site identification apparatus 1 showing the identification result of the binding site and the angiotensin I converting enzyme N-domain / inhibitor complex co-crystal structure were superimposed. It is a figure which shows an example of a thing.
  • the target substance may be described as a protein, the binding substance as a ligand that specifically binds to the protein, and the target substance as an amino acid.
  • the target substance may be a nucleic acid (DNA or RNA) or a nucleic acid-protein complex.
  • the binding substance may be a protein, a nucleic acid, or a complex thereof.
  • the target substance is a nucleic acid (A, T, G or C if the target substance is DNA, A, U, G or C if the target substance is RNA), or the target substance is a complex If so, one or both of the nucleic acid and amino acid may be used.
  • the protein may be bound to a coenzyme.
  • FIG. 1 is a diagram showing an outline of the present embodiment. This embodiment does not search for a binding site of a ligand in the protein along the surface of the protein (in other words, only the surface of the protein as a search range) as in the conventional case, but from the space on the surface of the protein. It is characterized by searching for a binding site of a ligand in the protein.
  • the small molecule when used as a ligand, the small molecule often binds to an amino acid residue in the middle of a protein pocket. Therefore, the amino acid residues that bind to the small molecule are not necessarily located at close positions in the distance along the surface of the protein. Therefore, in this embodiment, it is possible to identify an appropriate binding site even for a binding substance such as a small molecule by searching for a binding site from the space on the surface of the protein.
  • the three-dimensional structure of the protein and the experimental results regarding the number of presumed to be involved in the binding of a plurality of preset amino acids are obtained in advance, and (ii) the three-dimensional structure In the space from the surface of the cell to the point considering the size of the ligand, search for a location that matches the experimental result for the number of amino acid residues of each type, and (iii) a binding site based on the searched location You may identify.
  • this implementation In the form, from the space from the surface of the three-dimensional structure to the point in consideration of the size of the ligand, search for a location where there are 5 Ala, 3 Met and 2 Trp, and combine the searched locations Identify as site.
  • the following steps (A) to (C) are performed in advance to obtain the experimental results, and (ii) it is limited to a spherical shape from an arbitrary point around the three-dimensional structure.
  • Calculated (counting) the number of a plurality of types of amino acid residues that are set in advance in the defined space (for example, a space whose size depends on the volume of the ligand), and (iii) included in the experimental results A predetermined score for the degree of coincidence or mismatch is calculated based on the number of amino acid residues of each type and the calculated number of amino acid residues of each type, and (iv) (ii) and (iii) above ) Is performed on each point around the three-dimensional structure (for example, each point existing at a specific interval), and (v) the binding site may be identified based on the calculated score at each point. .
  • the protein is labeled (labeled) with a predetermined isotope prepared for each type of amino acid so that the number of signals obtained by NMR analysis can be identified for each type of amino acid residue set in advance.
  • the type of amino acid residue is preferably set in consideration of, for example, the tendency to exist at the interface of the binding site.
  • B Change in chemical shift before and after ligand binding is detected by NMR analysis.
  • C Judging that a signal having a chemical shift change greater than a predetermined threshold is due to an amino acid residue involved in binding or existing at the binding site (assuming it), the number of the signals Is calculated for each type of amino acid residue.
  • the chemical shift change of the unobserved signal is not detected, and the number of signals in cases where the chemical shift change is larger or smaller than the predetermined threshold is calculated. May be.
  • the weighted average value using the weight representing the tendency to exist at the binding site is used as a threshold for selecting a signal having a large chemical shift change.
  • the weight w j defined by Equation 1 is calculated for each type of amino acid residue.
  • j is a character for identifying a preset type of amino acid residue.
  • N all and N j are the number of all amino acid residues and the number of amino acid residues j existing in a hypothetical protein having characteristics according to statistics with the total protein as a population.
  • n j is the number of amino acid residues j present at the binding site with the hypothetical protein ligand. That is, in Equation 1, the term “ ⁇ n j / N all ” represents the ratio of amino acid residues present in the binding site, and the term “n j / N j ” (the “included in the target substance” of the present invention).
  • the ratio of the number of target substances to be detected and the number of target substances present at the binding site corresponds to the ratio of amino acid residues j existing at the binding site, and the term“ N j / N all ” Represents the abundance ratio of amino acid residue j in the protein, and the term “n j / ⁇ n j ” represents the ratio of amino acid residue j to all amino acid residues present in the binding site. .
  • the weighted average value ⁇ threshold defined by Equation 2 (“ threshold value defined in consideration of the ratio between the number of target substances included in the target substance and the number of target substances present in the binding site” of the present invention) Is a signal derived from an amino acid residue involved in binding.
  • j is a character for identifying a preset type of amino acid residue.
  • w j is a weight for amino acid residue j obtained in Equation 1.
  • a j and k j are the sum of the chemical shift change amount of the signal derived from amino acid residue j and the number of amino acid residues j, respectively.
  • ⁇ t m W + C lig S w
  • the threshold defined in the above may be used as the threshold.
  • m W is a weighted average value defined by Equation 2
  • C lig is a predetermined constant (for example, 1/2, 3/2, etc.)
  • S w is a weighted standard deviation.
  • the identification of the binding site and the presence / absence of competition can be quickly determined even for a small molecule with weak binding.
  • a binding site of a hit compound in drug discovery can be rapidly validated.
  • binding such as a protein-protein interaction (PPI) inhibitor can be performed.
  • PPI protein-protein interaction
  • an allosteric site can be quickly identified.
  • the binding site of an enzyme inhibitor is often near the active center of the enzyme. The binding site can be identified without error.
  • the binding site of an agonist with an allosteric action can be identified.
  • identifying a binding site of the compound or the like it is possible to provide a guide for prediction of a binding mode by computational chemistry and development to synthesis by the binding mode.
  • a mutation location can be selected by identifying a substrate binding site.
  • protein modification based on structural information can be realized together with prediction of a binding mode by computational chemistry.
  • a binding site can be identified even if a model structure (for example, one constructed from a homologous structure) is used as a three-dimensional structure.
  • a model structure for example, one constructed from a homologous structure
  • the docking simulation is performed after the binding site is determined. It becomes possible to construct a complex model structure.
  • the experimental result is not limited to that obtained by NMR analysis, but may be obtained by other methods.
  • FIG. 2 is a diagram illustrating an example of the configuration of the binding site identification apparatus 1.
  • the binding site identification device 1 connects the device to the network 2 (for example, the Internet) via a control unit 10 such as a CPU that comprehensively controls the device, a communication device such as a router, and a wired or wireless communication line such as a dedicated line.
  • a communication interface unit 11 that is communicably connected to an intranet or a LAN (including both wired and wireless), a storage unit 12 that stores various databases, tables, files, and the like, an input device 14 and an output device 15 And an input / output interface unit 13 connected to each other, and these units are communicably connected via an arbitrary communication path.
  • the communication interface unit 11 mediates communication between the binding site identification device 1 and the network 2.
  • the communication interface unit 11 has a function of communicating data with other terminals via a communication line.
  • the input / output interface unit 13 is connected to the input device 14 and the output device 15.
  • the output device 15 in addition to a monitor (including a home television), a speaker or a printer can be used.
  • the input device 14 may be a keyboard, mouse, or microphone, or a monitor that realizes a pointing device function in cooperation with the mouse.
  • the storage unit 12 is a storage means.
  • a memory device such as a RAM / ROM, a fixed disk device such as a hard disk, a flexible disk, or an optical disk can be used.
  • the storage unit 12 stores a computer program for giving various instructions to the CPU in cooperation with an OS (Operating System).
  • the storage unit 12 includes a three-dimensional structure storage unit 12a, an experiment result storage unit 12b, a calculation result storage unit 12c, an evaluation result storage unit 12d, and an identification result storage unit 12e.
  • target substances for example, PDB (Protein Data Bank) etc.
  • PDB Protein Data Bank
  • the three-dimensional structure of protein is stored.
  • the three-dimensional structure includes the three-dimensional coordinates of each atom constituting the substance.
  • the three-dimensional structure may relate to a single target substance, or may relate to a complex of a certain ligand other than the target ligand (corresponding to the binding substance of the present invention) and the target substance.
  • the three-dimensional structure may be a model structure in which all or part of the three-dimensional structure is predicted by a computational chemical technique.
  • the three-dimensional structure may be a model structure constructed from a homologous structure of the target substance.
  • the experiment result storage unit 12b stores an experiment result (corresponding to the acquisition result of the present invention) regarding the number of substances considered to be involved in the binding of the target substance (for example, one type or a plurality of types of amino acids).
  • the experimental result includes, for example, the name (type) of the target substance and the number of the target substance obtained in the experiment. Moreover, what was obtained by NMR experiment as an experimental result is preferable.
  • the calculation result storage unit 12c stores a calculation result regarding the number of target substances present in the search space set by a search space setting unit 10a2 described later, which is obtained by a substance number calculation unit 10a3 described later.
  • the calculation result includes, for example, an identification number for uniquely identifying the search point set in the search space or the search point setting unit 10a1 described later, the name (type) of the substance of interest, and the calculated attention And the number of substances.
  • the search space is associated with the search points set by the search point setting unit 10a1.
  • the evaluation result storage unit 12d stores an evaluation result related to the degree of coincidence or mismatch between the calculation result and the experimental result obtained by the degree evaluation unit 10a4 described later.
  • the evaluation result includes, for example, an identification number for uniquely identifying the search point and a value related to the degree of matching or mismatching.
  • the identification result storage unit 12e stores the identification result regarding the binding site obtained by the identification unit 10a described later (specifically, the search point determination unit 10a5 described later).
  • the identification result includes, for example, a search point determined by the search point determination unit 10a5 (specifically, a search remaining after a search point regarded as noise is deleted by a noise deletion unit 10a6 described later)
  • the identification number for uniquely identifying the point) and the three-dimensional coordinates of the search point are included.
  • the control unit 10 has an internal memory for storing a control program such as an OS (Operating System), a program defining various processing procedures, and necessary data, and performs various information processing based on these programs. Execute.
  • the control unit 10 includes an identification unit 10a and a drawing unit 10b.
  • the identification unit 10a includes a search point setting unit 10a1, a search space setting unit 10a2, a substance number calculation unit 10a3, a degree evaluation unit 10a4, a search point determination unit 10a5, a noise deletion unit 10a6, and a search point reference identification unit 10a7.
  • the identification unit 10a determines the number of target substances from the search space including the surface of the target substance.
  • the part considered to be coincident with the experimental result is searched, and the binding site is identified based on the searched part.
  • the size of the search space is set based on the size of the target ligand (for example, the volume of the ligand) or the radius of the smallest sphere that includes the target ligand. May be good.
  • the shape of the search space may be a sphere, an ellipsoid, or a polyhedron.
  • the search point setting unit 10a1 searches for a search point (grid) in a three-dimensional space (for example, on the surface of the target substance or near the surface of the target substance (for example, For example, a search point in a three-dimensional region including a point whose distance from a certain point on the surface is 10 mm or less is set.
  • the search point setting unit 10a1 associates the set search point with the three-dimensional coordinates of the search point.
  • the search point setting unit 10a1 detects, based on the three-dimensional structure, a point where a molecule (for example, a hydrophobic or hydrophilic molecule) may be bound from the viewpoint of interaction energy.
  • the search point setting unit 10a1 may set the search points with a specific interval (for example, 0.5 mm).
  • the program published in the document “Bioinformatics., 2009, 25, 3185” is a probe for determining the interaction energy as CMET, and there is a possibility that molecules are bound from the viewpoint of interaction energy. You may perform by setting the threshold value at the time of detecting a point to -8.0.
  • the programs disclosed in the document are composed of two programs, EasyMIFs and SiteHound.
  • the search space setting unit 10a2 sets a search space based on the search point set by the search point setting unit 10a1.
  • the search space setting unit 10a2 associates the set search space with a search point.
  • the search space setting unit 10a2 may set the size of the search space based on the size of the target ligand (for example, the volume of the ligand), and the minimum including the target ligand. You may set based on the radius of a sphere. Further, the search space setting unit 10a2 may set the shape of the search space to a sphere, an ellipsoid, or a polyhedron.
  • the substance number calculation unit 10a3 calculates the number of target substances existing in the search space set by the search space setting unit 10a2 based on the three-dimensional structure stored in the three-dimensional structure storage unit 12a.
  • the substance number calculation unit 10a3 stores the calculated number of target substances in the calculation result storage unit 12c in association with the search space or the search point.
  • the degree evaluation unit 10a4 evaluates the degree of coincidence or mismatch between the calculation result stored in the calculation result storage unit 12c and the experiment result stored in the experiment result storage unit 12b.
  • the degree evaluation unit 10a4 stores a value related to the evaluated degree in the evaluation result storage unit 12d in association with the search point that is used as the reference of the search space.
  • the degree evaluation unit 10a4 may calculate the absolute value of the difference between the number of target substances included in the calculation result and the number of target substances included in the experiment result as the degree.
  • the degree evaluation unit 10a4 may calculate an absolute value for each pattern and calculate the sum of the calculated absolute values as a degree. .
  • the search point determination unit 10a5 determines a search point whose degree evaluated by the degree evaluation unit 10a4 satisfies a predetermined condition based on the evaluation result stored in the evaluation result storage unit 12d.
  • the search point determination unit 10a5 stores the determined search point in the identification result storage unit 12e in association with the three-dimensional coordinates of the search point.
  • the predetermined condition may be a condition that the absolute value is equal to or less than the predetermined value.
  • the predetermined condition may be a condition that the sum is equal to or less than a predetermined value.
  • the noise deletion unit 10a6 deletes what is regarded as noise from the search point stored in the identification result storage unit 12e based on the size of the target ligand (for example, the volume of the ligand).
  • the noise deletion unit 10a6 may delete a set of search points (a cluster of search points) that is smaller than the size of the target ligand from the identification result storage unit 12e as noise. Good.
  • the search point reference identifying unit 10a7 identifies the binding site based on the search point stored in the identification result storage unit 12e.
  • the search point reference identifying unit 10a7 may identify a three-dimensional region constituted by the search points stored in the identification result storage unit 12e as a region corresponding to the binding site.
  • the drawing unit 10b causes the monitor 15 to display the processing result obtained by each processing unit included in the identifying unit 10a.
  • the drawing unit 10b monitors, for example, the three-dimensional structure of the target substance or the complex structure of the target substance and the target binding substance in which the processing results obtained by the respective processing units included in the identification unit 10a are reflected. 15 is displayed.
  • FIG. 3 is a flowchart illustrating an example of main processing performed in the binding site identification apparatus 1.
  • the target substance is a protein and the binding substance specifically binds to the protein with a small molecule (for example, a molecule having a molecular weight of, for example, 1,000 or less (however, a molecule considered as a low molecule is However, the present invention is not limited to this example.)),
  • the substance of interest is a plurality of types of amino acids that are known to have a high tendency to exist at low molecular binding sites.
  • the experiment result stored in the experiment result storage unit 12b includes a plurality of patterns (cases) of the number of at least one type of amino acid.
  • the experimental results stored in the experimental result storage unit 12b are based on the method for amino acid-specific isotope labeling and the change in chemical shift for each type of amino acid obtained by NMR analysis is a predetermined threshold value.
  • the search point setting unit 10a1 sets a plurality of grids (lattice points) in a three-dimensional space at specific intervals (for example, 0.5 cm) based on the three-dimensional structure stored in the three-dimensional structure storage unit 12a. Set and associate each set grid with a three-dimensional coordinate (step SA1).
  • the search space setting unit 10a2 sets a search space for each grid set in step SA1, and associates each set search space with a grid (step SA2).
  • the search space set in step SA2 has a radius that is a value obtained by adding a predetermined value (for example, 4 ⁇ ) to the radius of the minimum sphere (minimum inclusion sphere) including the low molecule centered on the grid. The space inside the sphere.
  • the number-of-substances calculation unit 10a3 calculates the number of amino acids of each type existing in the search space for each search space set in step SA2 based on the three-dimensional structure stored in the three-dimensional structure storage unit 12a. And the number of each type of amino acid corresponding to each calculated search space is stored in the calculation result storage unit 12c in association with the grid (step SA3).
  • step SA3 amino acids having a solvent exposure of 2% or less are excluded from the calculation target.
  • the degree evaluation unit 10a4 includes the number of each type of amino acid included in the calculation result stored in the calculation result storage unit 12c and each type of amino acid included in the experiment result stored in the experiment result storage unit 12b.
  • the absolute value of the difference from the number of is calculated for each case included in the experiment result, the sum of the absolute values for each calculated case is calculated, and the evaluation result is obtained by associating the calculated sum with the grid.
  • the data is stored in the storage unit 12d, and these processes are executed for all the grids set in step SA1 (step SA4).
  • the degree evaluation unit 10a4 calculates the final penalty score S tot of the grid, which is defined by Formula 3, and represents the degree of inconsistency between the calculation result and the experiment result, and is calculated.
  • the penalty score S tot is stored in the evaluation result storage unit 12d in association with the grid.
  • i is an identification number for identifying a case, and a is a letter for identifying an amino acid residue type.
  • S i is a penalty score for the grid in case i
  • R a is the number of amino acid residue species a existing in the search space of the grid included in the calculation result
  • Da is the number of signals having a large chemical shift change of each type of amino acid residue species a included in the experimental results stored in the experimental result storage unit 12b.
  • the search point determination unit 10a5 based on the evaluation result stored in the evaluation result storage unit 12d, a grid that satisfies the condition that the sum calculated in step SA4 is less than or equal to a predetermined value (specifically, the step A grid that satisfies the condition that the penalty score S tot calculated in SA4 is equal to or less than a predetermined value is determined, and the determined grid is associated with the three-dimensional coordinates and stored in the identification result storage unit 12e (step SA5). ).
  • the noise deletion unit 10a6 considers a set of grids that are smaller than the size of low molecules (for example, the volume of low molecules) from the grid stored in the identification result storage unit 12e as noise. (Step SA6).
  • the search point reference identifying unit 10a7 identifies the three-dimensional region formed by the grid stored in the identification result storage unit 12e as a region corresponding to a low molecular binding site in the protein (step SA7).
  • a search point is set based on the three-dimensional structure, a search space based on the set search point is set, and each of the existing search spaces based on the three-dimensional structure is set.
  • Calculate the number of types of amino acids evaluate the degree of coincidence or discrepancy between the obtained calculation results and the experimental results, determine the search points where the evaluated degree satisfies a predetermined condition, and set the determined search points Based on this, the binding site of the ligand may be identified. Thereby, an appropriate binding site can be identified more reliably.
  • the size of the search space is set based on the size of the ligand (for example, the volume of the ligand), or based on the radius of the smallest sphere including the binding substance. It may be set, and the shape of the search space may be a sphere, an ellipsoid, or a polyhedron. This makes it possible to efficiently search for a location (search point) that is considered to match the experimental result with respect to the number of target substances.
  • the absolute value of the difference between the number of amino acids of each type included in the calculation result and the number of amino acids of each type included in the experimental result is calculated as a degree, and the calculated absolute value is Search points that satisfy the condition that the value is less than or equal to a predetermined value may be determined. This makes it possible to more reliably determine the search points that are considered to match the experimental results for the number of each type of amino acid.
  • the absolute value is calculated for each case, and the sum of the calculated absolute values is calculated as a degree.
  • a search point that satisfies the condition that the calculated sum is equal to or less than a predetermined value may be determined. This determines the search points that are considered to be consistent with the experimental results for the number of amino acids of each type, even if the experimental results take into account ambiguity or uncertainty about the number of amino acids of each type. can do.
  • a point where a molecule may bind may be detected, and the detected point may be set as a search point.
  • search points that may be considered to coincide with the experimental results for the number of amino acids of each type.
  • the number of grids considered in the calculation of the present embodiment can be reduced and the calculation time can be shortened.
  • calculation accuracy can be increased by taking the viewpoint of interaction energy into account.
  • the site may be identified. Thereby, an appropriate binding site can be identified more reliably.
  • the number of amino acids that are considered to be involved in binding for each type of amino acid included in the experimental results is the number of each type of amino acid included in the protein and each type that exists in the binding site.
  • all or a part of the processes described as being automatically performed can be manually performed, or all of the processes described as being manually performed can be performed.
  • a part can be automatically performed by a known method.
  • each illustrated component is functionally conceptual and does not necessarily need to be physically configured as illustrated.
  • the processing functions provided in the binding site identification apparatus 1, particularly the processing functions performed by the control unit 10, are interpreted and executed by a CPU (Central Processing Unit) and the CPU in its entirety or any part thereof. You may implement
  • the program is recorded on a non-transitory computer-readable recording medium including a programmed instruction for causing the information processing apparatus to execute the binding site identification method according to the present invention. It is mechanically read by the identification device 1. That is, in the storage unit 12 such as a ROM or an HDD, computer programs for performing various processes by giving instructions to the CPU in cooperation with an OS (Operating System) are recorded. This computer program is executed by being loaded into the RAM, and constitutes a control unit in cooperation with the CPU.
  • an OS Operating System
  • this computer program may be stored in an application program server connected to the binding site identification device 1 via an arbitrary network, and may be downloaded in whole or in part as necessary. It is.
  • binding site identification program may be stored in a non-temporary computer-readable recording medium, or may be configured as a program product.
  • the “recording medium” means a memory card, USB memory, SD card, flexible disk, magneto-optical disk, ROM, EPROM, EEPROM, CD-ROM, MO, DVD, and Blu-ray (registered trademark). It includes any “portable physical medium” such as Disc.
  • the “program” is a data processing method described in an arbitrary language or description method, and may be in the form of source code or binary code. Note that the “program” is not necessarily limited to a single configuration, but is distributed in the form of a plurality of modules and libraries, or in cooperation with a separate program typified by an OS (Operating System). Including those that achieve the function. In addition, a well-known structure and procedure can be used about the specific structure and reading procedure for reading a recording medium in each apparatus shown to embodiment, the installation procedure after reading, etc.
  • Various databases stored in the storage unit 12 are memory devices such as RAM and ROM, fixed disk devices such as hard disks, flexible disks, and storage means such as optical disks. It stores various programs, tables, databases, web page files, etc. used for various processes and website provision.
  • the binding site identification device 1 may be configured as an information processing device such as a known personal computer or workstation, or may be configured as the information processing device to which an arbitrary peripheral device is connected.
  • the binding site identification apparatus 1 may be realized by installing software (including a program or data) that causes the information processing apparatus to realize the binding site identification method of the present invention.
  • the specific form of distribution / integration of the devices is not limited to that shown in the figure, and all or a part of them may be functionally or physically in arbitrary units according to various additions or according to functional loads. It can be configured to be distributed and integrated. That is, the above-described embodiments may be arbitrarily combined and may be selectively implemented.
  • Example 1 that verifies the feasibility of the binding site identification apparatus 1 will be described with reference to FIGS.
  • FIG. 4 is a diagram showing an example of the three-dimensional structure of MAPK14 and an inhibitor.
  • DS Visualizer 2.5 (Accelrys Co., Ltd.) was used to create the three-dimensional structure image.
  • MAPK14 MAP kinase p38 ⁇
  • 2-amino-3-benzyl-oxypyridine is a binding substance
  • six kinds of amino acids Met, Ala, His, Tyr, Trp and Phe
  • Example 1 the three-dimensional structure of MAPK14 input to the binding site identification apparatus 1 was obtained from PDB (Protein Data Bank). The coordinates of the missing atoms were constructed using Swiss-PdbViewer 4.0.1.
  • FIG. 5 is a diagram showing an example of the results of an inhibitor titration experiment.
  • the number of amino acid residues considered to be involved in the binding is 3 for the Ala residue, 1 for the Met residue, 1 for the Tyr residue, 0 or 1 for the His residue, 0 or 1 for Phe residues and 0 for Trp residues.
  • FIG. 8 is a diagram illustrating an example of a processing result of the search point setting unit 10a1.
  • the search point setting unit 10a1 detects a point where a hydrophobic molecule may be bound from the viewpoint of interaction energy based on the three-dimensional structure, and sets the detected point as a grid, the setting is performed.
  • the grid (dots shown in purple in FIG. 8) was superimposed on the three-dimensional structure of MAPK14 and displayed on the monitor 15.
  • FIG. 9 is a diagram illustrating an example of a processing result of the degree evaluation unit 10a4.
  • the set grid is color-coded according to the value of the penalty score S tot (in other words, the high degree of coincidence), and is displayed on the monitor 15 so as to overlap with the three-dimensional structure of MAPK14.
  • a grid with a low penalty score S tot (in other words, a grid with a high degree of coincidence) is shown in red, and as the penalty score S tot increases, the color of the grid becomes orange, yellow, or green.
  • a grid having a higher penalty score S tot (in other words, a grid having a lower degree of coincidence) is shown in blue.
  • FIG. 10 is a diagram illustrating an example of a processing result of the search point determination unit 10a5.
  • the grid indicated by mainly red or orange having a low penalty score S tot was displayed on the monitor 15 so as to overlap with the three-dimensional structure of MAPK14.
  • FIG. 11 is a diagram illustrating an example of a processing result of the noise deletion unit 10a6.
  • a grid having a low penalty score S tot after a small block of the grid determined to be noise based on the size of the inhibitor was deleted was displayed on the monitor 15 in a superimposed manner with the three-dimensional structure of MAPK14. .
  • FIG. 12 is a diagram illustrating an example of a processing result of the search point reference identification unit 10a7.
  • the cluster of grids (the area shown in red in FIG. 12) remaining after the noise was deleted was displayed on the monitor 15 in a manner superimposed on the three-dimensional structure of MAPK14.
  • the mass of the grid was identified as a binding site (a space where an inhibitor is present at the time of binding).
  • each grid is a point indicating the center of the search space, and the space considered as a binding site is a wider range than the mass of the grid. Therefore, in FIG. Carbon atoms are displayed in CPK, and the surface is shown in red.
  • FIG. 13 is a diagram showing an example in which the identification result of the binding site shown in FIG. 12 and the known complex co-crystal structure of MAPK14 and inhibitor are superimposed.
  • the complex co-crystal structure is shown in green.
  • the binding site space identified by the binding site identification device 1 is almost the same as the binding site space confirmed in the complex co-crystal structure, and the binding site identification device 1 appropriately identifies the binding site of the inhibitor in MAPK14. We were able to.
  • the position of the binding site identified by the binding site identification apparatus 1 is slightly different from the position of the binding site confirmed in the complex co-crystal structure because the structure has changed due to the binding. It is considered a thing.
  • Example 2 that verifies the feasibility of the binding site identification apparatus 1 will be described with reference to FIGS. 14 to 20.
  • FIG. 14 is a diagram showing an example of a complex structure of MAPK14 and lipid ( ⁇ -OG).
  • the complex structure of MAPK14 and lipid is known, but experimental evidence that clearly shows the binding site of fatty acid in MAPK14 has not been obtained. Therefore, using the binding site identification apparatus 1, an attempt was made to identify the binding site of fatty acid (decanoic acid) in MAPK14.
  • MAPK14 was used as a target substance, and decanoic acid was used as a binding substance.
  • six types of amino acids (Met, Ala, His, Tyr, Trp, and Phe) were used as the substance of interest.
  • Example 2 the experimental results regarding the number of the six types of amino acids that are considered to be involved in the binding, which are input to the binding site identification apparatus 1, were obtained by the technique and NMR analysis regarding the amino acid-specific isotope label ( For example, refer to the document “J. Struct. Biol., 2011, 174, p. 434-442”).
  • the three-dimensional structure of MAPK14 input to the binding site identification apparatus 1 was obtained from PDB (Protein Data Bank). The coordinates of the missing atoms were constructed using Swiss-PdbViewer 4.0.1.
  • FIG. 15 is a diagram showing an example of the results of a decanoic acid titration experiment.
  • Decanoic acid titration experiments observed a significant shift in chemical shift at one Met residue, one Ala residue, one His residue, and one Trp residue, while the Tyr residue It was not observed that the chemical shift was significantly changed at the and Phe residues.
  • the number of amino acid residues considered to be involved in binding is 1 for the Ala residue, 1 for the Met residue, 0 for the Tyr residue, 1 or 2 for the His residue, 0 or 1 for Phe residues and 1 for Trp residues.
  • FIG. 17 is a diagram illustrating an example of a processing result of the search point determination unit 10a5.
  • the grid indicated by mainly red or orange having a low penalty score S tot was displayed on the monitor 15 so as to overlap with the three-dimensional structure of MAPK14.
  • FIG. 18 is a diagram illustrating an example of a processing result of the noise deletion unit 10a6 and a processing result of the search point reference identification unit 10a7.
  • a grid having a low penalty score S tot after a small lump of the grid determined to be noise based on the size of decanoic acid was deleted was displayed on the monitor 15 in a superimposed manner with the three-dimensional structure of MAPK14. .
  • the mass of the grid that remained after the noise was removed was identified as the binding site (the space in which decanoic acid was present at the time of binding).
  • FIG. 19 is a diagram showing an example in which the identification result of the binding site shown in FIG. 18 is overlapped with the known complex structure of MAPK14 and ⁇ -OG. It was shown that the space of the binding site of decanoic acid identified by the binding site identification apparatus 1 is almost the same as the space of the binding site of ⁇ -OG confirmed in the complex structure.
  • FIG. 20 is a diagram showing an example of the results of a titration experiment of decanoic acid and ⁇ -OG. It was confirmed by a titration experiment and a mutation experiment that the binding site identified by the binding site identification apparatus 1 was correct.
  • a change in chemical shift was examined for a signal derived from Trp residue by a titration experiment of decanoic acid and ⁇ -OG, a signal in which the chemical shift was greatly changed by the addition of decanoic acid and a chemical shift by ⁇ -OG were added. The signal for which the shift changed significantly was the same. Mutation experiments revealed that the signal whose chemical shift was greatly changed by the addition of decanoic acid or ⁇ -OG was the 197th Trp residue.
  • the 197th Trp residue is common to both the decanoic acid binding site and the ⁇ -OG binding site, that is, the decanoic acid binding site and the ⁇ -OG binding site are almost the same. There was found.
  • the complex structure of MAPK14 and ⁇ -OG also reveals that the 197th Trp residue is present at the binding site of ⁇ -OG. Therefore, it was confirmed that the binding site of decanoic acid is the binding site identified by the binding site identification apparatus 1.
  • Example 3 that verifies the feasibility of the binding site identification apparatus 1 will be described with reference to FIGS. PyMOL 1.5.0.3 (Schrödinger Co., Ltd.) was used for creating a three-dimensional image.
  • Example 3 in the crystal structure of a known protein-low molecular complex, the number of substances of interest existing within 5 cm of the periphery of the low molecule is acquired in advance, and acquisition regarding the number of substances considered to be involved in binding is acquired. As a result.
  • FIG. 21 shows an example in which the apo structure of Factor Xa showing the binding site identified by using the binding site identification apparatus 1 and the known complex co-crystal structure of Factor Xa and an inhibitor are superimposed.
  • FIG. 21 In identifying a binding site with the binding site identification apparatus 1, Factor Xa is a target substance, an inhibitor is a binding substance, and six types of amino acids (Met, Ala, His, Tyr, Trp, and Phe) are used as target substances. did.
  • the apo structure is shown in gray and the complex co-crystal structure is shown in green.
  • the space of the binding site of the inhibitor identified by the binding site identification apparatus 1 (the space shown in red in FIG. 21) is the inhibitor confirmed in the complex co-crystal structure (shown in blue in FIG. 21).
  • the binding site identification apparatus 1 was able to appropriately identify the binding site of the inhibitor from the apo structure.
  • FIG. 22 shows an apo structure of Galectin-9 N-domain showing a binding site identified using the binding site identification apparatus 1, and a known complex co-crystal structure of Galectin-9 N-domain and lactose. It is a figure which shows an example of what was superimposed.
  • Galectin-9, N-domain is a target substance
  • lactose is a binding substance
  • six types of amino acids (Met, Ala, His, Tyr, Trp, and Phe) are used. The substance of interest.
  • the apo structure is shown in gray, and the complex co-crystal structure is shown in green.
  • the lactose binding site space identified by the binding site identification apparatus 1 (the space shown in red in FIG. 22) is the lactose confirmed in the complex co-crystal structure (shown in blue in FIG. 22).
  • the binding site identification device 1 was able to appropriately identify the lactose binding site from the apo structure.
  • FIG. 23 shows the apo structure of Angiotensin I converting enzyme N-domain showing the binding site identified using the binding site identifying apparatus 1, and the known complex of Angiotensin I converting enzyme N-domain and an inhibitor. It is a figure which shows an example of what was superimposed with the crystal structure.
  • angiotensin I converting enzyme N-domain is a target substance
  • an inhibitor is a binding substance
  • six types of amino acids (Met, Ala, His, Tyr, Trp, and Phe are used. )
  • the apo structure is shown in gray and the complex co-crystal structure is shown in green.
  • the space of the binding site of the inhibitor identified by the binding site identification apparatus 1 (the space shown in red in FIG. 23) is the inhibitor confirmed in the complex co-crystal structure (shown in blue in FIG. 23).
  • the binding site identification apparatus 1 was able to appropriately identify the binding site of the inhibitor from the apo structure.
  • the present invention can be widely implemented in many industrial fields, in particular, molecular design including drug discovery and protein modification, and is extremely useful.

Landscapes

  • Physics & Mathematics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biotechnology (AREA)
  • Medicinal Chemistry (AREA)
  • Immunology (AREA)
  • Theoretical Computer Science (AREA)
  • Medical Informatics (AREA)
  • Evolutionary Biology (AREA)
  • High Energy & Nuclear Physics (AREA)
  • General Physics & Mathematics (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Biomedical Technology (AREA)
  • Urology & Nephrology (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Biochemistry (AREA)
  • Analytical Chemistry (AREA)
  • Hematology (AREA)
  • Pathology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Condensed Matter Physics & Semiconductors (AREA)
  • Microbiology (AREA)
  • Cell Biology (AREA)
  • Food Science & Technology (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • Investigating Or Analysing Biological Materials (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

低分子のような結合物質であっても適切な結合部位を同定することができる結合部位同定装置、結合部位同定方法および結合部位同定プログラムを提供することを課題とする。本実施形態では、標的物質の立体構造と、注目物質について結合に関与すると見做されるものの数に関する実験結果とに基づいて、標的物質の表面を含む探索空間から、注目物質の数について実験結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて結合部位を同定する。

Description

結合部位同定装置、結合部位同定方法および結合部位同定プログラム
 本発明は、結合部位同定装置、結合部位同定方法および結合部位同定プログラムに関するものである。
 非特許文献1には、タンパク質にリガンドを結合させた際の当該タンパク質における当該リガンドの結合部位を同定する技術が開示されている。具体的には、非特許文献1には、タンパク質の立体構造と、予め設定された複数種類のアミノ酸残基について結合に関与すると見做されるものの数に関する実験結果とに基づいて、当該タンパク質の表面から、当該複数種類のアミノ酸残基の数について当該実験結果と一致する箇所を探索し、探索された当該箇所を結合部位として同定する技術が開示されている。
 なお、その他の先行技術文献として特許文献1が挙げられる。特許文献1の段落0077には、結合ドメインの位置と配列を、既知のアミノ酸配列とNMR(Nuclear Magnetic Resonance)で同定されたアミノ酸を比較することで決定する、と記載されている。
 ここで、リガンドとして低分子を対象とする場合、当該低分子はタンパク質のポケットの途中にあるアミノ酸残基と結合することが多々ある。そのため、必ずしも、当該低分子と結合するアミノ酸残基同士が当該タンパク質の表面に沿った距離において近い位置にある、とは限らない。
 しかしながら、非特許文献1では、タンパク質の表面に沿って結合部位を探索する、つまりタンパク質の表面だけを探索範囲としているので、低分子を対象とする場合には適切な結合部位を同定することが困難である、という問題点があった。
 本発明は、上記問題点に鑑みてなされたもので、低分子のような結合物質であっても適切な結合部位を同定することができる結合部位同定装置、結合部位同定方法および結合部位同定プログラムを提供することを目的とする。
 上述した課題を解決し、目的を達成するために、本発明にかかる結合部位同定装置は、標的物質における結合物質の結合部位を同定する、制御部を備えた結合部位同定装置であって、前記制御部は、前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定手段を備えることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記探索空間の大きさは、前記結合物質の大きさに基づいて設定されたものであることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記結合物質の大きさは、前記結合物質の体積であることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記探索空間の大きさは、前記結合物質を包含する最小の球の半径に基づいて設定されたものであることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記探索空間の形状は、球、楕円体または多面体であることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記同定手段は、前記立体構造に基づいて探索点を設定する探索点設定手段と、前記探索点設定手段で設定された前記探索点を基準とする前記探索空間を設定する探索空間設定手段と、前記立体構造に基づいて、前記探索空間設定手段で設定された前記探索空間に存在する前記注目物質の数を計算する物質数計算手段と、前記物質数計算手段で得られた計算結果と前記取得結果との一致または不一致の度合いを評価する度合い評価手段と、前記度合い評価手段で評価された前記度合いが所定の条件を満たす前記探索点を決定する探索点決定手段と、前記探索点決定手段で決定された前記探索点に基づいて前記結合部位を同定する探索点基準同定手段と、をさらに備えることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記度合い評価手段は、前記計算結果に含まれる前記注目物質の数と前記取得結果に含まれる前記注目物質の数との差の絶対値を前記度合いとして計算し、前記所定の条件は、前記絶対値が所定値以下であるという条件であることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記取得結果に前記注目物質の数が複数パターン含まれている場合、前記度合い評価手段は、各々の前記パターンごとに前記絶対値を計算し、計算された前記絶対値の総和を前記度合いとして計算し、前記所定の条件は、前記総和が所定値以下であるという条件であることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記探索点設定手段は、前記立体構造に基づいて、相互作用エネルギーの観点から、分子が結合する可能性のある点を検出し、検出された当該点を前記探索点として設定することを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記同定手段は、前記探索点決定手段で決定された前記探索点から、前記結合物質の大きさに基づいて、ノイズと見做されるものを削除するノイズ削除手段をさらに備え、前記探索点基準同定手段は、前記ノイズ削除手段が実行された後に残った前記探索点に基づいて前記結合部位を同定することを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記取得結果に含まれる、前記注目物質について前記結合に関与すると見做されるものの数は、前記標的物質に含まれる前記注目物質の数と前記結合部位に存在する前記注目物質の数との比を考慮して定義された閾値に基づいて計算されたものであることを特徴とする。
 また、本発明にかかる結合部位同定装置は、前記の結合部位同定装置において、前記標的物質はタンパク質であり、前記結合物質は当該タンパク質と結合するものであり、前記注目物質は1種類または複数種類のアミノ酸であることを特徴とする。
 また、本発明にかかる結合部位同定方法は、標的物質における結合物質の結合部位を同定する、制御部を備えた情報処理装置で実行される結合部位同定方法であって、前記制御部で実行される、前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定ステップを含むことを特徴とする。
 また、本発明にかかる結合部位同定プログラムは、標的物質における結合物質の結合部位を同定する、制御部を備えた情報処理装置に実行させるための結合部位同定プログラムであって、前記制御部に実行させるための、前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定ステップを含むことを特徴とする。
 また、本発明にかかる記録媒体は、一時的でないコンピュータ読み取り可能な記録媒体であって、情報処理装置に前記の結合部位同定方法を実行させるためのプログラム化された命令(前記の結合部位同定プログラム)を含むこと、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(探索点設定装置)であって、前記制御部は、標的物質の立体構造に基づいて探索点を設定する探索点設定手段を備え、前記探索点設定手段で設定された探索点は、前記標的物質の表面を含む探索空間を設定するときに用いられるものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(探索空間設定装置)であって、前記制御部は、予め設定された探索点を基準とする、標的物質の表面を含む探索空間を設定する探索空間設定手段を備え、前記探索空間設定手段で設定された探索空間は、前記標的物質の立体構造に基づいて当該探索空間に存在する注目物質の数を計算するときに用いられるものであり、前記予め設定された探索点は、前記立体構造に基づいて設定されたものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(物質数計算装置)であって、前記制御部は、標的物質の立体構造に基づいて、予め設定された探索空間に存在する注目物質の数を計算する物質数計算手段を備え、前記物質数計算手段で得られた計算結果は、前記注目物質について前記標的物質と結合物質との結合に関与すると見做されるものの数に関する取得結果と当該計算結果との一致または不一致の度合いを評価するときに用いられるものであり、前記予め設定された探索空間は、前記立体構造に基づいて設定された探索点を基準として設定された、前記標的物質の表面を含むものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(度合い評価装置)であって、前記制御部は、予め設定された計算結果と、注目物質について標的物質と結合物質との結合に関与すると見做されるものの数に関する取得結果との一致または不一致の度合いを評価する度合い評価手段を備え、前記度合い評価手段で評価された度合いは、前記標的物質の立体構造に基づいて設定された探索点から、当該度合いが所定の条件を満たす探索点を決定するときに用いられるものであり、前記予め設定された計算結果は、前記立体構造に基づいて設定された探索点を基準として設定された前記標的物質の表面を含む探索空間に存在する前記注目物質の数に関して前記立体構造に基づき計算されたものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(探索点決定装置)であって、前記制御部は、予め設定された探索点から、予め設定された度合いが所定の条件を満たす探索点を決定する探索点決定手段を備え、前記探索点決定手段で決定された探索点は、当該探索点に基づいて標的物質における結合物質の結合部位を同定するときに用いられるものであり、前記予め設定された探索点は、前記標的物質の立体構造に基づいて設定されたものであり、前記予め設定された度合いは、前記予め設定された探索点を基準として設定された前記標的物質の表面を含む探索空間に存在する注目物質の数に関して前記立体構造に基づき計算された計算結果と、前記注目物質について前記標的物質と前記結合物質との結合に関与すると見做されるものの数に関する取得結果との一致または不一致の度合いに関して評価されたものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(探索点基準同定装置)であって、前記制御部は、予め設定された探索点に基づいて、標的物質における結合物質の結合部位を同定する探索点基準同定手段を備え、前記予め設定された探索点は、前記標的物質の立体構造に基づいて設定された探索点を基準として設定された前記標的物質の表面を含む探索空間に存在する注目物質の数に関して前記立体構造に基づき計算された計算結果と、前記注目物質について前記標的物質と前記結合物質との結合に関与すると見做されるものの数に関する取得結果との一致または不一致の度合いが所定の条件を満たすと決定されたものであること、を特徴とする。
 また、本発明にかかる情報処理装置は、制御部を備えた情報処理装置(ノイズ削除装置)であって、前記制御部は、予め設定された探索点から、標的物質に結合する結合物質の大きさに基づいて、ノイズと見做されるものを削除するノイズ削除手段を備え、前記ノイズ削除手段が実行された後に残った前記探索点は、当該探索点に基づいて前記標的物質における前記結合物質の結合部位を同定するときに用いられるものであり、前記予め設定された探索点は、前記標的物質の立体構造に基づいて設定された探索点を基準として設定された前記標的物質の表面を含む探索空間に存在する注目物質の数に関して前記立体構造に基づき計算された計算結果と、前記注目物質について前記標的物質と前記結合物質との結合に関与すると見做されるものの数に関する取得結果との一致または不一致の度合いが所定の条件を満たすと決定されたものであること、を特徴とする。
 また、本発明にかかる結合部位同定システムは、標的物質の立体構造に基づいて探索点を設定する探索点設定手段と、前記探索点設定手段で設定された前記探索点を基準とする、前記標的物質の表面を含む探索空間を設定する探索空間設定手段と、前記立体構造に基づいて、前記探索空間設定手段で設定された前記探索空間に存在する注目物質の数を計算する物質数計算手段と、前記物質数計算手段で得られた計算結果と、前記注目物質について前記標的物質と結合物質との結合に関与すると見做されるものの数に関する取得結果との一致または不一致の度合いを評価する度合い評価手段と、前記度合い評価手段で評価された前記度合いが所定の条件を満たす前記探索点を決定する探索点決定手段と、前記探索点決定手段で決定された前記探索点に基づいて、前記標的物質における前記結合物質の結合部位を同定する探索点基準同定手段と、を備えたことを特徴とする。
 本発明によれば、標的物質の立体構造と、注目物質について結合に関与すると見做されるものの数に関する取得結果とに基づいて、標的物質の表面を含む探索空間から、注目物質の数について取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて結合部位を同定する。これにより、低分子のような結合物質であっても適切な結合部位を同定することができるという効果を奏する。
 本発明によれば、探索空間の大きさは、結合物質の大きさに基づいて設定されたものである。これにより、注目物質の数について取得結果と一致すると見做される箇所を効率よく探索することができるという効果を奏する。
 本発明によれば、結合物質の大きさは、結合物質の体積である。これにより、注目物質の数について取得結果と一致すると見做される箇所を効率よく探索することができるという効果を奏する。
 本発明によれば、探索空間の大きさは、結合物質を包含する最小の球の半径に基づいて設定されたものである。これにより、注目物質の数について取得結果と一致すると見做される箇所を効率よく探索することができるという効果を奏する。
 本発明によれば、探索空間の形状は、球、楕円体または多面体である。これにより、注目物質の数について取得結果と一致すると見做される箇所を効率よく探索することができるという効果を奏する。
 本発明によれば、立体構造に基づいて探索点を設定し、設定された探索点を基準とする探索空間を設定し、立体構造に基づいて、設定された探索空間に存在する注目物質の数を計算し、得られた計算結果と取得結果との一致または不一致の度合いを評価し、評価された度合いが所定の条件を満たす探索点を決定し、決定された探索点に基づいて結合部位を同定する。これにより、適切な結合部位を、より確実に同定することができるという効果を奏する。
 本発明によれば、計算結果に含まれる注目物質の数と取得結果に含まれる注目物質の数との差の絶対値を度合いとして計算し、所定の条件は、絶対値が所定値以下であるという条件である。これにより、注目物質の数について取得結果と一致すると見做される探索点を、より確実に決定することができるという効果を奏する。
 本発明によれば、取得結果に注目物質の数が複数パターン含まれている場合、各々のパターンごとに絶対値を計算し、計算された絶対値の総和を度合いとして計算し、所定の条件は、総和が所定値以下であるという条件である。これにより、取得結果が、注目物質の数について曖昧さまたは不確実さが考慮されたものであっても、注目物質の数について取得結果と一致すると見做される探索点を決定することができるという効果を奏する。
 本発明によれば、立体構造に基づいて、相互作用エネルギーの観点から、分子が結合する可能性のある点を検出し、検出された当該点を探索点として設定する。これにより、注目物質の数について取得結果と一致すると見做される可能性のある探索点を設定することができるという効果を奏する。また、相互作用エネルギーの観点から探索点をある程度絞ることで、本発明の計算で考慮するグリッドの数を減らし、計算時間を短くすることができるという効果を奏する。また、相互作用エネルギーの観点を取り入れることで、計算精度を上げることができるという効果を奏する。
 本発明によれば、決定された探索点から、結合物質の大きさに基づいて、ノイズと見做されるものを削除し、ノイズ削除が実行された後に残った探索点に基づいて結合部位を同定する。これにより、適切な結合部位を、より確実に同定することができるという効果を奏する。
 本発明によれば、取得結果に含まれる、注目物質について結合に関与すると見做されるものの数は、標的物質に含まれる注目物質の数と結合部位に存在する注目物質の数との比を考慮して定義された閾値に基づいて計算されたものである。これにより、より正確な取得結果に基づいて適切な結合部位を同定することができるという効果を奏する。
 本発明によれば、標的物質はタンパク質であり、結合物質は当該タンパク質と結合するものであり、注目物質は1種類または複数種類のアミノ酸である。これにより、アミノ酸に注目して、タンパク質における結合物質の適切な結合部位を同定することができるという効果を奏する。
図1は、本実施形態の概要を示す図である。 図2は、結合部位同定装置1の構成の一例を示すブロック図である。 図3は、結合部位同定装置1で行われるメイン処理の一例を示すフローチャートである。 図4は、MAPK14および阻害剤の立体構造の一例を示す図である。 図5は、阻害剤滴定実験の結果の一例を示す図である。 図6は、6種類のアミノ酸残基種に由来する各シグナルについての、規格化された化学シフト変化量の一例を示す図である。 図7は、加重平均値の分布の一例を示す図である。 図8は、探索点設定部10a1の処理結果の一例を示す図である。 図9は、度合い評価部10a4の処理結果の一例を示す図である。 図10は、探索点決定部10a5の処理結果の一例を示す図である。 図11は、ノイズ削除部10a6の処理結果の一例を示す図である。 図12は、探索点基準同定部10a7の処理結果の一例を示す図である。 図13は、図12に示されている結合部位の同定結果と、MAPK14と阻害剤の複合体共結晶構造と、を重ね合わせたものの一例を示す図である。 図14は、MAPK14と脂質の複合体構造の一例を示す図である。 図15は、デカン酸滴定実験の結果の一例を示す図である。 図16は、6種類のアミノ酸残基種に由来する各シグナルについての、規格化された化学シフト変化量の一例を示す図である。 図17は、探索点決定部10a5の処理結果の一例を示す図である。 図18は、ノイズ削除部10a6の処理結果および探索点基準同定部10a7の処理結果の一例を示す図である。 図19は、図18に示す結合部位の同定結果と、MAPK14とβ−OGの複合体構造と、を重ね合わせたものの一例を示す図である。 図20は、デカン酸およびβ−OGの滴定実験および変異実験の結果の一例を示す図である。 図21は、結合部位同定装置1で用いられた、結合部位の同定結果が示されているapo構造と、Factor Xaと阻害剤の複合体共結晶構造と、が重ね合わされたものの一例を示す図である。 図22は、結合部位同定装置1で用いられた、結合部位の同定結果が示されているapo構造と、Galectin−9 N−domainとラクトースの複合体共結晶構造と、が重ね合わされたものの一例を示す図である。 図23は、結合部位同定装置1で用いられた、結合部位の同定結果が示されているapo構造と、Angiotensin I 変換酵素 N−domainと阻害剤の複合体共結晶構造と、が重ね合わされたものの一例を示す図である。
 以下に、本発明にかかる結合部位同定装置、結合部位同定方法および結合部位同定プログラムの実施形態を、図面に基づいて説明する。なお、本実施形態により本発明が限定されるものではない。特に、本実施形態では、標的物質をタンパク質とし、結合物質を当該タンパク質に特異的に結合するリガンドとし、注目物質をアミノ酸として説明する場合があるが、本発明はこの場合に限定されるものではない。例えば、標的物質を、核酸(DNAまたはRNA)または核酸−タンパク質複合体としてもよい。また、例えば、結合物質を、タンパク質、核酸、またはこれらの複合体としてもよい。また、例えば、注目物質を、核酸(標的物質がDNAであればA、T、GまたはC、標的物質がRNAであればA、U、GまたはC)、または、標的物質が複合体の場合であれば上記核酸とアミノ酸の片方または両方としてもよい。なお、標的物質および結合物質のいずれにおいても、タンパク質は、補酵素と結合したものでもよい。
[1.本実施形態の概要]
 図1は、本実施形態の概要を示す図である。本実施形態は、従来のように、タンパク質の表面に沿って(換言すると、タンパク質の表面だけを探索範囲として)当該タンパク質におけるリガンドの結合部位を探索するのではなく、タンパク質の表面上の空間から当該タンパク質におけるリガンドの結合部位を探索する、ことが特徴である。
 ここで、低分子をリガンドとする場合、当該低分子はタンパク質のポケットの途中にあるアミノ酸残基と結合することが多々ある。そのため、必ずしも、当該低分子と結合するアミノ酸残基同士が当該タンパク質の表面に沿った距離において近い位置にある、とは限らない。そこで、本実施形態では、タンパク質の表面上の空間から結合部位を探索することにより、低分子のような結合物質であっても適切な結合部位を同定することが可能となる。
 なお、本実施形態では、(i)タンパク質の立体構造と、予め設定した複数種類のアミノ酸について結合に関与すると見做されるものの数に関する実験結果とを予め取得しておき、(ii)立体構造の表面上からリガンドのサイズを考慮した点までの空間内に、各種類のアミノ酸残基の数について実験結果と一致する箇所を探索し、(iii)探索された当該箇所に基づいて結合部位を同定してもよい。例えば、予め設定された複数種類のアミノ酸がAla、MetおよびTrpで、実験結果が、結合部位にAlaが5個、Metが3個およびTrpが2個存在するというデータを含むものである場合、本実施形態では、立体構造の表面上からリガンドのサイズを考慮した点までの空間内から、Alaが5個、Metが3個およびTrpが2個存在する箇所を探索し、探索された当該箇所を結合部位として同定する。
 また、本実施形態では、(i)事前に、以下の(A)から(C)の工程を実施して実験結果を取得しておき、(ii)立体構造周辺の任意の一点から球状に限定された空間(例えばリガンドの体積に依存する大きさの空間など)内に在る、予め設定した複数種類のアミノ酸残基の数を計算し(カウントし)、(iii)実験結果に含まれている各種類のアミノ酸残基の数と計算された各種類のアミノ酸残基の数とに基づいて、一致度または不一致度に関する所定のスコアを計算し、(iv)前記の(ii)および(iii)の処理を、立体構造周辺の各点(例えば、特定の間隔で存在する各点)に対し実行し、(v)計算された各点でのスコアに基づいて結合部位を同定してもよい。
(A)NMR解析で得られるシグナルの数が、予め設定した各種類のアミノ酸残基ごとに識別できるように、各種類のアミノ酸ごとに用意された所定の同位体でタンパク質を標識(ラベル)する。なお、アミノ酸残基の種類は、例えば、結合部位の界面に存在する傾向を考慮して設定することが好ましい。
(B)リガンドの結合前後での化学シフト変化をNMR解析で検出する。
(C)化学シフト変化が所定の閾値よりも大きいシグナルを、結合に関与するまたは結合部位に存在するアミノ酸残基に因るものであると判断して(見做して)、当該シグナルの個数を各種類のアミノ酸残基ごとに計算する。なお、NMRスペクトル上で観測されないシグナルがあった場合には、当該観測されないシグナルの化学シフト変化は検出せず、化学シフト変化が所定の閾値よりも大きいケースと小さいケースでのシグナルの個数を計算してもよい。このようにして実験結果にシグナルの個数について複数のケースを含めることで、NMR解析で得られる情報に曖昧さまたは不確実さが含まれる場合でも、適切な結合部位を探索することができる。
 ここで、化学シフト変化が大きいシグナルの選び方について簡単に説明する。従来では、得られた全てのシグナルの化学シフト変化量の平均値を、化学シフト変化が大きいシグナルを選ぶ際の閾値とすることが多かった。また、相互作用部位(結合部位)の界面に存在する傾向が、アミノ酸の種類により異なる(例えば、文献「J.Chem.Inf.Model.,2007,47,400」参照)。そこで、本実施形態では、一部のアミノ酸に由来するシグナルのみをNMR解析で観測する。
 しかし、観測対象とするアミノ酸を変更すると、これに応じて閾値が変わるので、最終的に得られる結果(化学シフト変化が大きいシグナルの数)も変わる可能性がある。そこで、本実施形態では、結合部位への存在傾向を表す重みを用いた加重平均値を、化学シフト変化が大きいシグナルを選ぶ際の閾値とする。
 具体的には、まず、アミノ酸残基の種類ごとに、数式1で定義される重みwを計算する。数式1において、jは、予め設定したアミノ酸残基の種類を識別するための文字である。数式1において、NallおよびNは、全タンパク質を母集団とする統計に従う特徴を有する仮想的なタンパク質に存在する、全アミノ酸残基の数およびアミノ酸残基jの数である。nは、前記仮想的なタンパク質のリガンドとの結合部位に存在するアミノ酸残基jの数である。つまり、数式1において、「Σn/Nall」の項は、結合部位に存在するアミノ酸残基の割合を表しており、「n/N」の項(本発明の「標的物質に含まれる注目物質の数と結合部位に存在する注目物質の数との比」に相当)は、アミノ酸残基jについて結合部位に存在する割合を表しており、「N/Nall」の項は、タンパク質中におけるアミノ酸残基jの存在比率を表しており、「n/Σn」の項は、結合部位に存在する全てのアミノ酸残基のうちアミノ酸残基jが占める割合を表している。
Figure JPOXMLDOC01-appb-M000001
 つぎに、数式2で定義される加重平均値Δωthreshold(本発明の「標的物質に含まれる注目物質の数と結合部位に存在する注目物質の数との比を考慮して定義された閾値」に相当)を閾値とし、この閾値より化学シフト変化量が大きいシグナルを、結合に関与しているアミノ酸残基に由来するシグナルであるとする。数式2において、jは、予め設定したアミノ酸残基の種類を識別するための文字である。数式2において、wは、数式1で得られる、アミノ酸残基jについての重みである。数式2において、aおよびkは、それぞれ、アミノ酸残基jに由来するシグナルの化学シフト変化量の和、およびアミノ酸残基jの個数である。
Figure JPOXMLDOC01-appb-M000002
 なお、式「ω=m+Clig」で定義されるω(本発明の「標的物質に含まれる注目物質の数と結合部位に存在する注目物質の数との比を考慮して定義された閾値」に相当)を閾値としてもよい。当該式において、mは数式2で定義される加重平均値であり、Cligは所定の定数(例えば、1/2、または3/2など)であり、Sは加重標準偏差である。
 以上、本実施形態の概要について説明したが、本実施形態によれば、結合が弱い小さな分子であっても、結合部位の同定、および競合の有無などを迅速に判断することができる。また、本実施形態によれば、例えば創薬などにおけるヒット化合物の結合部位についてのバリデーションを迅速に行うことができ、例えば創薬においては、タンパク質間相互作用(PPI)阻害剤のような、結合部位が容易に推測できない薬剤について、結合部位の同定が可能である。また、本実施形態によれば、アロステリック部位を迅速に同定することができ、例えば酵素阻害剤の結合部位は酵素の活性中心付近であることが多いが、アロステリックな作用による酵素阻害剤について、その結合部位を誤ることなく同定できる。また、アロステリックな作用による作動薬について、その結合部位を同定可能である。特に、本実施形態によれば、前記化合物などの結合部位を同定することで、計算化学による結合様式の予測と当該結合様式による合成への展開の指針を提供することができる。また、本実施形態によれば、基質結合部位を同定することで、変異箇所を選定することができる。特に、本実施形態によれば、構造情報に基づくタンパク質改変を、計算化学による結合様式の予測と併せて実現することが可能となる。
 また、本実施形態によれば、立体構造としてモデル構造(例えばホモログの構造から構築されたものなど)を用いても、結合部位を同定することができる。また、本実施形態によれば、結合に関与しているアミノ酸残基の一部が予測(ほぼ決定)することができるので、結合部位が決まった後、ドッキングシミュレーションを行うことにより、精度の高い複合体モデル構造を構築することが可能となる。また、本実施形態によれば、実験結果はNMR解析で得られるものに限らず、他の手法で得られるものでもよい。
[2.本実施形態の構成]
 図2は、結合部位同定装置1の構成の一例を示す図である。結合部位同定装置1は、当該装置を統括的に制御するCPU等の制御部10と、ルータ等の通信装置および専用線等の有線または無線の通信回線を介して当該装置をネットワーク2(例えばインターネット、イントラネットまたはLAN(有線/無線の双方を含む)等)に通信可能に接続する通信インターフェース部11と、各種のデータベース、テーブルまたはファイルなどを格納する記憶部12と、入力装置14および出力装置15に接続する入出力インターフェース部13と、を備え、これら各部は任意の通信路を介して通信可能に接続される。
 通信インターフェース部11は、結合部位同定装置1とネットワーク2との間における通信を媒介する。通信インターフェース部11は、他の端末と通信回線を介してデータを通信する機能を有する。
 入出力インターフェース部13は、入力装置14および出力装置15と接続される。出力装置15には、モニタ(家庭用テレビを含む)の他、スピーカまたはプリンタなどを用いることができる。入力装置14には、キーボード、マウスまたはマイクの他、マウスと協働してポインティングデバイス機能を実現するモニタを用いることができる。
 記憶部12は、ストレージ手段である。記憶部12として、例えば、RAM・ROM等のメモリ装置、ハードディスクのような固定ディスク装置、フレキシブルディスク、または光ディスク等を用いることができる。記憶部12には、OS(Operating System)と協働してCPUに命令を与えて各種処理を行うためのコンピュータプログラムが記録されている。記憶部12には、立体構造記憶部12a、実験結果記憶部12b、計算結果記憶部12c、評価結果記憶部12dおよび同定結果記憶部12eが含まれる。
 立体構造記憶部12aには、各種物質(例えばタンパク質など)の立体構造に関するデータを公開している外部データベース(例えばPDB(Protein Data Bank)など)からネットワーク2を介して取得された、標的物質(例えばタンパク質など)の立体構造が格納される。なお、立体構造には、物質を構成する各原子の3次元座標などが含まれる。また、立体構造は、標的物質単体に関するものでもよく、対象とするリガンド(本発明の結合物質に相当)以外の別の或るリガンドと標的物質の複合体に関するものでもよい。また、立体構造は、その全部または一部が計算化学的な手法によって予測されたモデル構造でもよい。また、立体構造は、標的物質のホモログの構造から構築されたモデル構造でもよい。
 実験結果記憶部12bには、注目物質(例えば1種類または複数種類のアミノ酸など)について結合に関与すると見做されるものの数に関する実験結果(本発明の取得結果に相当)が格納される。ここで、実験結果には、例えば、注目物質の名称(種類)と、実験で得られた注目物質の数と、が含まれる。また、実験結果は、NMR実験で得られたものが好ましい。また、実験結果に含まれる、注目物質について結合に関与すると見做されるものの数は、例えば、標的物質に含まれる注目物質の数と結合部位に存在する注目物質の数との比(例えば前記の数式1で定義される重みwなど)を考慮して定義された閾値(例えば、前記の数式2で定義される加重平均値Δωthresholdまたは前記の式「ω=m+Clig」で定義されるωなど)に基づいて計算されたものでもよい。
 計算結果記憶部12cには、後述する物質数計算部10a3で得られた、後述する探索空間設定部10a2で設定された探索空間に存在する注目物質の数に関する計算結果が格納される。ここで、計算結果には、例えば、探索空間または後述する探索点設定部10a1で設定された探索点を一意に識別するための識別番号と、注目物質の名称(種類)と、計算された注目物質の数と、が含まれる。なお、探索空間は、探索点設定部10a1で設定された探索点と関連付けられている。
 評価結果記憶部12dには、後述する度合い評価部10a4で得られた、計算結果と実験結果との一致または不一致の度合いに関する評価結果が格納される。ここで、評価結果には、例えば、探索点を一意に識別するための識別番号と、一致または不一致の度合いに関する値と、が含まれる。
 同定結果記憶部12eには、後述する同定部10a(具体的には、後述する探索点決定部10a5)で得られた、結合部位に関する同定結果が格納される。ここで、同定結果には、例えば、探索点決定部10a5で決定された探索点(具体的には、後述するノイズ削除部10a6でノイズと見做される探索点が削除された後に残った探索点)を一意に識別するための識別番号と、探索点の3次元座標と、が含まれる。
 制御部10は、OS(Operating System)等の制御プログラム・各種の処理手順等を規定したプログラム・所要データなどを格納するための内部メモリを有し、これらのプログラムに基づいて種々の情報処理を実行する。制御部10は、同定部10aおよび描画部10bを備える。同定部10aは、探索点設定部10a1、探索空間設定部10a2、物質数計算部10a3、度合い評価部10a4、探索点決定部10a5、ノイズ削除部10a6および探索点基準同定部10a7を備える。
 同定部10aは、立体構造記憶部12aに格納されている立体構造と実験結果記憶部12bに格納されている実験結果とに基づいて、標的物質の表面を含む探索空間から、注目物質の数について実験結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて結合部位を同定する。ここで、探索空間の大きさは、対象とするリガンドの大きさ(例えばリガンドの体積など)に基づいて設定されたもの、または対象とするリガンドを包含する最小の球の半径に基づいて設定されたものでもよい。また、探索空間の形状は、球、楕円体または多面体でもよい。
 探索点設定部10a1は、立体構造記憶部12aに格納されている立体構造に基づいて、3次元空間上の探索点(グリッド)(例えば、標的物質の表面上または標的物質の表面近傍(例えば、表面上の或る点との距離が例えば10Å以下である点を含む3次元領域など)に在る探索点など)を設定する。探索点設定部10a1は、設定された探索点を、当該探索点の3次元座標と関連付ける。ここで、探索点設定部10a1は、立体構造に基づいて、相互作用エネルギーの観点から、分子(例えば疎水的または親水的な分子など)が結合する可能性のある点を検出し、検出された当該点を探索点として設定してもよい。具体的には、例えば、文献「Bioinformatics.,2009,25,3185」に公開されているプログラムを用いて、分子が結合する可能性のある点を検出してもよい。また、探索点設定部10a1は、探索点を、特定の間隔(例えば0.5Åなど)を空けて設定してもよい。
 具体的には、文献「Bioinformatics.,2009,25,3185」に公開されているプログラムを、相互作用エネルギーを求める際のプローブをCMETとし、相互作用エネルギーの観点から分子が結合する可能性のある点を検出する際の閾値を−8.0として、実行してもよい。ここで、当該文献で公開されているプログラムは、EasyMIFsとSiteHoundの2つのプログラムからなる。
 EasyMIFsでは、タンパク質周囲にグリッドを設け、指定したプローブが存在した場合にどの程度のエネルギー的寄与が得られるかを、各グリッドについて計算する。ここで、後述する実施例では、プローブにCMET、グリッドの間隔に0.5Åを設定した。
 SiteHoundでは、プローブの存在により一定以上(域値は自分で指定)のエネルギー的寄与が見込まれるグリッドのみを残す。なお、グリッドのクラスタリングを行ってもよい。ここで、後述する実施例では、閾値に−8.0、クラスタリングの際のリンクにaverage、カットオフに7.8を設定した。なお、リンクとカットオフをどのように設定しても、得られる「分子が結合する可能性のある点」は変わらない。
 探索空間設定部10a2は、探索点設定部10a1で設定された探索点を基準とする探索空間を設定する。探索空間設定部10a2は、設定された探索空間を、探索点と関連付ける。ここで、探索空間設定部10a2は、探索空間の大きさを、対象とするリガンドの大きさ(例えばリガンドの体積など)に基づいて設定してもよく、また対象とするリガンドを包含する最小の球の半径に基づいて設定してもよい。また、探索空間設定部10a2は、探索空間の形状を、球、楕円体または多面体としてもよい。
 物質数計算部10a3は、立体構造記憶部12aに格納されている立体構造に基づいて、探索空間設定部10a2で設定された探索空間に存在する注目物質の数を計算する。物質数計算部10a3は、計算された注目物質の数を、探索空間または探索点と関連付けて計算結果記憶部12cに格納する。
 度合い評価部10a4は、計算結果記憶部12cに格納されている計算結果と実験結果記憶部12bに格納されている実験結果との一致または不一致の度合いを評価する。度合い評価部10a4は、評価された度合いに関する値を、探索空間の基準とされた探索点と関連付けて評価結果記憶部12dに格納する。ここで、度合い評価部10a4は、計算結果に含まれる注目物質の数と実験結果に含まれる注目物質の数との差の絶対値を度合いとして計算してもよい。また、実験結果に注目物質の数が複数パターン含まれている場合、度合い評価部10a4は、各々のパターンごとに絶対値を計算し、計算された絶対値の総和を度合いとして計算してもよい。
 探索点決定部10a5は、評価結果記憶部12dに格納されている評価結果に基づいて、度合い評価部10a4で評価された度合いが所定の条件を満たす探索点を決定する。探索点決定部10a5は、決定された探索点を、当該探索点の3次元座標と関連付けて同定結果記憶部12eに格納する。ここで、度合い評価部10a4が前記の絶対値を度合いとして計算する場合には、所定の条件は当該絶対値が所定値以下であるという条件でもよい。また、度合い評価部10a4が前記の総和を度合いとして計算する場合には、所定の条件は当該総和が所定値以下であるという条件でもよい。
 ノイズ削除部10a6は、同定結果記憶部12eに格納されている探索点から、対象とするリガンドの大きさ(例えばリガンドの体積など)に基づいて、ノイズと見做されるものを削除する。ここで、ノイズ削除部10a6は、例えば、対象とするリガンドの大きさに対して小さい探索点の集合(探索点の塊)を、ノイズと見做して同定結果記憶部12eから削除してもよい。
 探索点基準同定部10a7は、同定結果記憶部12eに格納されている探索点に基づいて結合部位を同定する。ここで、探索点基準同定部10a7は、同定結果記憶部12eに格納されている探索点で構成される3次元領域を、結合部位に相当する領域として同定してもよい。
 描画部10bは、同定部10aに含まれている各処理部で得られた処理結果を、モニタ15に表示させる。描画部10bは、例えば、同定部10aに含まれている各処理部で得られた処理結果が反映された、標的物質の立体構造または標的物質と対象とする結合物質の複合体構造を、モニタ15に表示させる。
[3.本実施形態の処理]
 図3は、結合部位同定装置1で行われるメイン処理の一例を示すフローチャートである。なお、本説明では、標的物質をタンパク質とし、結合物質を当該タンパク質と特異的に結合する低分子(例えば、分子量が例えば1,000以下の分子など(ただし、低分子と見做される分子は、この例示に限定されるものではない。))とし、注目物質を、低分子の結合部位に存在している傾向が高いことが知られている複数種類のアミノ酸とする。また、本説明では、実験結果記憶部12bに格納されている実験結果は、少なくとも1種類のアミノ酸についてその数を複数パターン(ケース)含むものであるとする。また、本説明では、実験結果記憶部12bに格納されている実験結果は、アミノ酸特異的な同位体ラベルに関する手法およびNMR解析によって得られた、各種類のアミノ酸について化学シフトの変化が所定の閾値(例えば、前記の数式2で定義される加重平均値Δωthresholdまたは前記の式「ω=m+Clig」で定義されるωなど)に対して大きいシグナルの個数を含むものであるとする。
 まず、探索点設定部10a1は、立体構造記憶部12aに格納されている立体構造に基づいて、3次元空間上のグリッド(格子点)を特定の間隔(例えば0.5Åなど)を空けて複数設定し、設定された各グリッドを3次元座標と関連付ける(ステップSA1)。
 つぎに、探索空間設定部10a2は、ステップSA1で設定された各グリッド対し探索空間を設定し、設定された各探索空間をグリッドと関連付ける(ステップSA2)。ここで、ステップSA2で設定される探索空間は、グリッドを中心とし、低分子を包含する最小の球(最小包含球)の半径に所定値(例えば4Åなど)を足し合わせた値を半径とする球の内側の空間とする。
 つぎに、物質数計算部10a3は、立体構造記憶部12aに格納されている立体構造に基づいて、ステップSA2で設定された各探索空間に対して、探索空間に存在する各種類のアミノ酸の数を計算し、計算された各探索空間に対応する各種類のアミノ酸の数を、グリッドと関連付けて計算結果記憶部12cに格納する(ステップSA3)。ここで、ステップSA3では、溶媒露出度が2%以下であるアミノ酸は、計算対象から除外される。
 つぎに、度合い評価部10a4は、計算結果記憶部12cに格納されている計算結果に含まれる各種類のアミノ酸の数と実験結果記憶部12bに格納されている実験結果に含まれる各種類のアミノ酸の数との差の絶対値を、当該実験結果に含まれている各ケースに対し計算し、計算された各ケースに対する絶対値の総和を計算し、計算された総和をグリッドと関連付けて評価結果記憶部12dに格納し、そしてこれらの処理を、ステップSA1で設定された全てのグリッドに対し実行する(ステップSA4)。具体的には、度合い評価部10a4は、計算結果と実験結果の不一致度を表す、数式3で定義される、グリッドが有する最終的なペナルティスコアStotを、各グリッドに対し計算し、計算されたペナルティスコアStotをグリッドと関連付けて評価結果記憶部12dに格納する。なお、数式3において、iはケースを識別するための識別番号であり、aはアミノ酸残基種を識別するための文字である。また、数式3において、Sは、ケースiにおける、グリッドに対するペナルティスコアであり、Rは、計算結果に含まれている、グリッドの探索空間に存在するアミノ酸残基種aの数であり、Dは、実験結果記憶部12bに格納されている実験結果に含まれている、各種類のアミノ酸残基種aの、化学シフト変化が大きいシグナルの数である。
Figure JPOXMLDOC01-appb-M000003
 ここで、数式3に基づくペナルティスコアの計算例について説明する。例えば、実験結果に「Metの数が1、Alaの数が2、Pheの数が0または1」というデータが含まれているとし、また、或るグリッドに対する計算結果に「Metの数が0、Alaの数が2、Pheの数が0」というデータが含まれているとする。この場合、実験結果から、ケース1として「Metの数が1、Alaの数が2、Pheの数が0」が設定され、ケース2として「Metの数が1、Alaの数が2、Pheの数が1」が設定される。そして、ケース1におけるペナルティスコアSの値は1と計算され、ケース2におけるペナルティスコアSの値は2と計算される。そして、ペナルティスコアStotの値は、Sの値である1とSの値である2を足し合わせて、3と計算される。
 つぎに、探索点決定部10a5は、評価結果記憶部12dに格納されている評価結果に基づいて、ステップSA4で計算された総和が所定値以下である条件を満たすグリッド(具体的には、ステップSA4で計算されたペナルティスコアStotの値が所定値以下であるという条件を満たすグリッド)を決定し、決定されたグリッドをその3次元座標と関連付けて同定結果記憶部12eに格納する(ステップSA5)。
 つぎに、ノイズ削除部10a6は、同定結果記憶部12eに格納されているグリッドから、低分子の大きさ(例えば低分子の体積など)に対して小さい、グリッドの集合を、ノイズと見做して削除する(ステップSA6)。
 つぎに、探索点基準同定部10a7は、同定結果記憶部12eに格納されているグリッドで構成される3次元領域を、タンパク質における低分子の結合部位に相当する領域として同定する(ステップSA7)。
[4.本実施形態のまとめ、および他の実施形態]
 以上、本実施形態によれば、タンパク質の立体構造と、複数種類のアミノ酸について結合に関与すると見做されるものの数に関する実験結果とに基づいて、タンパク質の表面を含む探索空間から、各種類のアミノ酸の数について実験結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて、タンパク質におけるリガンド(例えば低分子など)の結合部位を同定する。これにより、低分子のような結合物質であっても適切な結合部位を同定することができる。
 なお、本実施形態によれば、立体構造に基づいて探索点を設定し、設定された探索点を基準とする探索空間を設定し、立体構造に基づいて、設定された探索空間に存在する各種類のアミノ酸の数を計算し、得られた計算結果と実験結果との一致または不一致の度合いを評価し、評価された度合いが所定の条件を満たす探索点を決定し、決定された探索点に基づいてリガンドの結合部位を同定してもよい。これにより、適切な結合部位を、より確実に同定することができる。
 ここで、本実施形態によれば、探索空間の大きさは、リガンドの大きさ(例えばリガンドの体積など)に基づいて設定されたもの、または結合物質を包含する最小の球の半径に基づいて設定されたものでもよく、さらに探索空間の形状は、球、楕円体または多面体でもよい。これにより、注目物質の数について実験結果と一致すると見做される箇所(探索点)を効率よく探索することができる。
 また、本実施形態によれば、計算結果に含まれる各種類のアミノ酸の数と実験結果に含まれる各種類のアミノ酸の数との差の絶対値を度合いとして計算し、計算された絶対値が所定値以下であるという条件を満たす探索点を決定してもよい。これにより、各種類のアミノ酸の数について実験結果と一致すると見做される探索点を、より確実に決定することができる。また、本実施形態によれば、実験結果に各種類のアミノ酸の数が複数ケース含まれている場合、各々のケースごとに絶対値を計算し、計算された絶対値の総和を度合いとして計算し、計算された総和が所定値以下であるという条件を満たす探索点を決定してもよい。これにより、実験結果が、各種類のアミノ酸の数について曖昧さまたは不確実さが考慮されたものであっても、各種類のアミノ酸の数について実験結果と一致すると見做される探索点を決定することができる。
 また、本実施形態によれば、立体構造に基づいて、相互作用エネルギーの観点から、分子が結合する可能性のある点を検出し、検出された当該点を探索点として設定してもよい。これにより、各種類のアミノ酸の数について実験結果と一致すると見做される可能性のある探索点を設定することができる。また、相互作用エネルギーの観点から探索点をある程度絞ることで、本実施形態の計算で考慮するグリッドの数を減らし、計算時間を短くすることができる。また、相互作用エネルギーの観点を取り入れることで、計算精度を上げることができる。
 また、本実施形態によれば、決定された探索点から、リガンドの大きさに基づいて、ノイズと見做されるものを削除し、ノイズ削除が実行された後に残った探索点に基づいて結合部位を同定してもよい。これにより、適切な結合部位を、より確実に同定することができる。
 また、本実施形態によれば、実験結果に含まれる、各種類のアミノ酸について結合に関与すると見做されるものの数は、タンパク質に含まれる各種類のアミノ酸の数と結合部位に存在する各種類のアミノ酸の数との比(例えば前記の数式1で定義される重みwなど)を考慮して定義された閾値(例えば、前記の数式2で定義される加重平均値Δωthresholdまたは前記の式「ω=m+Clig」で定義されるωなど)に基づいて計算されたものでもよい。これにより、より正確な実験結果に基づいて適切な結合部位を同定することができる。
 さて、これまで本発明の実施形態について説明したが、本発明は、上述した実施形態以外にも、特許請求の範囲に記載した技術的思想の範囲内において種々の異なる実施形態にて実施されてよいものである。
 また、実施形態において説明した各処理のうち、自動的に行われるものとして説明した処理の全部または一部を手動的に行うこともでき、あるいは、手動的に行われるものとして説明した処理の全部または一部を公知の方法で自動的に行うこともできる。
 このほか、上記文献中や図面中で示した処理手順、制御手順、具体的名称、各処理の登録データや検索条件等のパラメータを含む情報、画面例、データベース構成については、特記する場合を除いて任意に変更することができる。
 また、結合部位同定装置1に関して、図示の各構成要素は機能概念的なものであり、必ずしも物理的に図示の如く構成されていることを要しない。
 例えば、結合部位同定装置1が備える処理機能、特に制御部10にて行われる各処理機能については、その全部または任意の一部を、CPU(Central Processing Unit)および当該CPUにて解釈実行されるプログラムにて実現してもよく、また、ワイヤードロジックによるハードウェアとして実現してもよい。尚、プログラムは、情報処理装置に本発明にかかる結合部位同定方法を実行させるためのプログラム化された命令を含む一時的でないコンピュータ読み取り可能な記録媒体に記録されており、必要に応じて結合部位同定装置1に機械的に読み取られる。すなわち、ROMまたはHDDなどの記憶部12などには、OS(Operating System)と協働してCPUに命令を与え、各種処理を行うためのコンピュータプログラムが記録されている。このコンピュータプログラムは、RAMにロードされることによって実行され、CPUと協働して制御部を構成する。
 また、このコンピュータプログラムは、結合部位同定装置1に対して任意のネットワークを介して接続されたアプリケーションプログラムサーバに記憶されていてもよく、必要に応じてその全部または一部をダウンロードすることも可能である。
 また、本発明にかかる結合部位同定プログラムを、一時的でないコンピュータ読み取り可能な記録媒体に格納してもよく、また、プログラム製品として構成することもできる。ここで、この「記録媒体」とは、メモリーカード、USBメモリ、SDカード、フレキシブルディスク、光磁気ディスク、ROM、EPROM、EEPROM、CD−ROM、MO、DVD、および、Blu−ray(登録商標) Disc等の任意の「可搬用の物理媒体」を含むものとする。
 また、「プログラム」とは、任意の言語または記述方法にて記述されたデータ処理方法であり、ソースコードまたはバイナリコード等の形式を問わない。なお、「プログラム」は必ずしも単一的に構成されるものに限られず、複数のモジュールやライブラリとして分散構成されるものや、OS(Operating System)に代表される別個のプログラムと協働してその機能を達成するものをも含む。なお、実施形態に示した各装置において記録媒体を読み取るための具体的な構成および読み取り手順ならびに読み取り後のインストール手順等については、周知の構成や手順を用いることができる。
 記憶部12に格納される各種のデータベース等(立体構造記憶部12aに格納される立体構造、実験結果記憶部12bに格納される実験結果、計算結果記憶部12cに格納される計算結果、評価結果記憶部12dに格納される評価結果および同定結果記憶部12eに格納される同定結果など)は、RAM、ROM等のメモリ装置、ハードディスク等の固定ディスク装置、フレキシブルディスク、および、光ディスク等のストレージ手段であり、各種処理やウェブサイト提供に用いる各種のプログラム、テーブル、データベース、および、ウェブページ用ファイル等を格納する。
 また、結合部位同定装置1は、既知のパーソナルコンピュータまたはワークステーション等の情報処理装置として構成してもよく、また、任意の周辺装置が接続された当該情報処理装置として構成してもよい。また、結合部位同定装置1は、当該情報処理装置に本発明の結合部位同定方法を実現させるソフトウェア(プログラムまたはデータ等を含む)を実装することにより実現してもよい。
 更に、装置の分散・統合の具体的形態は図示するものに限られず、その全部または一部を、各種の付加等に応じてまたは機能負荷に応じて、任意の単位で機能的または物理的に分散・統合して構成することができる。すなわち、上述した実施形態を任意に組み合わせて実施してもよく、実施形態を選択的に実施してもよい。
 結合部位同定装置1の実施可能性を検証した実施例1を、図4から図13を参照して説明する。
 図4は、MAPK14および阻害剤の立体構造の一例を示す図である。立体構造の画像作成には、DS Visualizer 2.5(アクセルリス株式会社)を用いた。実施例1では、MAPK14(MAP kinase p38α)を標的物質とし、その阻害剤である2−amino−3−benzyl−oxypyridineを結合物質とし、6種類のアミノ酸(Met、Ala、His、Tyr、TrpおよびPhe)を注目物質とした。実施例1では、結合部位同定装置1に入力する、6種類のアミノ酸について結合に関与すると見做されるものの数に関する実験結果を、アミノ酸特異的な同位体ラベルに関する手法およびNMR解析で得た(例えば、文献「J.Struct.Biol.,2011,174,p.434−442」を参照)。実施例1では、結合部位同定装置1に入力するMAPK14の立体構造を、PDB(Protein Data Bank)から得た。なお、欠損している原子の座標は、Swiss−PdbViewer 4.0.1を用いて構築した。
 図5は、阻害剤滴定実験の結果の一例を示す図である。図6は、6種類のアミノ酸残基種に由来する各シグナルについての、文献「J.Struct.Biol.,2011,174,434」に記載されているように規格化された化学シフト変化量(ΔωRMS)の一例を示す図であり、図中横線は、前記の式「ω=m+Clig」で表される閾値である。図7は、規格化された化学シフト変化量(ΔωRMS)の分布の一例を示す図であり、ωは前記の式「ω=m+Clig」で表される閾値である。阻害剤滴定実験により、3つのAla残基、1つのMet残基および1つのTyr残基で、化学シフトが有意に変化していることが観測され、一方、His残基、Trp残基およびPhe残基で、化学シフトが有意に変化していることは観測されなかった。ただし、His残基およびPhe残基のそれぞれで、阻害剤の添加前後のいずれにおいても、シグナルが1つ観測されなかった。そこで、観測されなかった1つのシグナルが結合に関与する可能性を考慮して、結合に関与すると見做されるHis残基の数およびPhe残基の数として、それぞれ0または1という2つのケースを設定した。従って、全ての組合せを考慮し、4つのケース((His,Phe)={(0,0),(0,1),(1,0),(1,1)})を設定した。
 以上、阻害剤滴定実験の結果、結合に関与すると見做されるアミノ酸残基の数は、Ala残基で3、Met残基で1、Tyr残基で1、His残基で0または1、Phe残基で0または1、そしてTrp残基で0であった。
 図8は、探索点設定部10a1の処理結果の一例を示す図である。探索点設定部10a1が、立体構造に基づいて、相互作用エネルギーの観点から、疎水的な分子が結合する可能性のある点を検出し、検出された当該点をグリッドとして設定した場合において、設定された当該グリッド(図8において紫色で示されている点)をMAPK14の立体構造と重ね合わせてモニタ15に表示した。
 図9は、度合い評価部10a4の処理結果の一例を示す図である。設定されたグリッドを、ペナルティスコアStotの値(換言すると一致度の高さ)に応じて色分けして、MAPK14の立体構造と重ね合わせてモニタ15に表示した。図9において、ペナルティスコアStotの値が低いグリッド(換言すると一致度が高いグリッド)は赤色で示され、ペナルティスコアStotの値が高くなるにつれて、グリッドの色が、オレンジ色、黄色、緑色の順に変化し、そしてペナルティスコアStotの値が高いグリッド(換言すると一致度が低いグリッド)は青色で示されている。
 図10は、探索点決定部10a5の処理結果の一例を示す図である。設定されたグリッドのうち、ペナルティスコアStotの値が低かった、主に赤色またはオレンジ色などで示されているグリッドを、MAPK14の立体構造と重ね合わせてモニタ15に表示した。
 図11は、ノイズ削除部10a6の処理結果の一例を示す図である。阻害剤の大きさに基づいてノイズであると判断されたグリッドの小さな塊が削除された後のペナルティスコアStotの値が低かったグリッドを、MAPK14の立体構造と重ね合わせてモニタ15に表示した。
 図12は、探索点基準同定部10a7の処理結果の一例を示す図である。ノイズが削除された後に残ったグリッドの塊(図12において赤色で示されている領域)を、MAPK14の立体構造と重ね合わせてモニタ15に表示した。当該グリッドの塊は、結合部位(結合時に阻害剤が存在している空間)として同定された。ここで、各グリッドは探索空間の中心を示す点であり、結合部位として考えられる空間は当該グリッドの塊よりも広い範囲となるため、図12では、例として炭素原子を置いた塊に置き換えて、炭素原子をCPK表示し、その表面を赤色で示している。
 図13は、図12に示されている結合部位の同定結果と、既知の、MAPK14と阻害剤の複合体共結晶構造と、を重ね合わせたものの一例を示す図である。図13において、複合体共結晶構造は緑色で示されている。結合部位同定装置1で同定された結合部位の空間は複合体共結晶構造で確認される結合部位の空間とほぼ同じであり、結合部位同定装置1はMAPK14における阻害剤の結合部位を適切に同定することができた。なお、結合部位同定装置1で同定された結合部位の位置が、複合体共結晶構造で確認される結合部位の位置から若干ずれているのは、結合に伴う構造変化があったことに因るものと考えられる。
 結合部位同定装置1の実施可能性を検証した実施例2を、図14から図20を参照して説明する。
 図14は、MAPK14と脂質(β−OG)の複合体構造の一例を示す図である。MAPK14と脂質の複合体構造は既知であるが、MAPK14における脂肪酸の結合部位を明確に示した実験的証拠は得られていない。そこで、結合部位同定装置1を用いて、MAPK14における脂肪酸(デカン酸)の結合部位の同定を試みた。なお、実施例2では、MAPK14を標的物質とし、デカン酸を結合物質として用いた。また、実施例2では、6種類のアミノ酸(Met、Ala、His、Tyr、TrpおよびPhe)を注目物質として用いた。実施例2では、結合部位同定装置1に入力する、6種類のアミノ酸について結合に関与すると見做されるものの数に関する実験結果を、アミノ酸特異的な同位体ラベルに関する手法およびNMR解析で得た(例えば、文献「J.Struct.Biol.,2011,174,p.434−442」を参照)。実施例2では、結合部位同定装置1に入力するMAPK14の立体構造を、PDB(Protein Data Bank)から得た。なお、欠損している原子の座標は、Swiss−PdbViewer 4.0.1を用いて構築した。
 図15は、デカン酸滴定実験の結果の一例を示す図である。図16は、6種類のアミノ酸残基種に由来する各シグナルついての、文献「J.Struct.Biol.,2011,174,434」に記載されているように規格化された化学シフト変化量(ΔωRMS)の一例を示す図であり、図中横線は、前記の式「ω=m+Clig」で表される閾値である。デカン酸滴定実験により、1つのMet残基、1つのAla残基、1つのHis残基および1つのTrp残基で、化学シフトが有意に変化していることが観測され、一方、Tyr残基およびPhe残基で、化学シフトが有意に変化していることは観測されなかった。ただし、His残基およびPhe残基のそれぞれで、デカン酸の添加前後のいずれにおいても、シグナルが1つ観測されなかった。そこで、観測されなかった1つのシグナルが結合に関与する可能性を考慮して、結合に関与すると見做されるHis残基の数として1または2という2つのケースを設定し、結合に関与すると見做されるPhe残基の数として0または1という2つのケースを設定した。従って、全ての組合せを考慮し、4つのケース((His,Phe)={(1,0),(1,1),(2,0),(2,1)})を設定した。
 以上、デカン酸滴定実験の結果、結合に関与すると見做されるアミノ酸残基の数は、Ala残基で1、Met残基で1、Tyr残基で0、His残基で1または2、Phe残基で0または1、そしてTrp残基で1であった。
 図17は、探索点決定部10a5の処理結果の一例を示す図である。設定されたグリッドのうち、ペナルティスコアStotの値が低かった、主に赤色またはオレンジ色などで示されているグリッドを、MAPK14の立体構造と重ね合わせてモニタ15に表示した。
 図18は、ノイズ削除部10a6の処理結果および探索点基準同定部10a7の処理結果の一例を示す図である。デカン酸の大きさに基づいてノイズであると判断されたグリッドの小さな塊が削除された後のペナルティスコアStotの値が低かったグリッドを、MAPK14の立体構造と重ね合わせてモニタ15に表示した。ノイズが削除された後に残ったグリッドの塊(図18においてオレンジ色で示されている領域)は、結合部位(結合時にデカン酸が存在している空間)として同定された。
 図19は、図18に示す結合部位の同定結果と、既知の、MAPK14とβ−OGの複合体構造と、を重ね合わせたものの一例を示す図である。結合部位同定装置1で同定されたデカン酸の結合部位の空間は、複合体構造で確認されるβ−OGの結合部位の空間とほぼ同じであることが示された。
 図20は、デカン酸およびβ−OGの滴定実験の結果の一例を示す図である。結合部位同定装置1で同定された結合部位が正しいことを、滴定実験および変異実験で確認した。まず、デカン酸およびβ−OGの滴定実験により、Trp残基由来のシグナルについて化学シフトの変化を調べたところ、デカン酸の添加により化学シフトが大きく変化したシグナルと、β−OGの添加により化学シフトが大きく変化したシグナルは同じであった。変異実験により、デカン酸またはβ−OGの添加により化学シフトが大きく変化したシグナルは、197番目のTrp残基であることが判明した。従って、デカン酸の結合部位およびβ−OGの結合部位のどちらにも共通して197番目のTrp残基が存在する、つまりデカン酸の結合部位とβ−OGの結合部位はほぼ同じであることが判明した。MAPK14とβ−OGの複合体構造からも、197番目のTrp残基はβ−OGの結合部位に存在することが明らかとなっている。よって、デカン酸の結合部位は、結合部位同定装置1で同定された結合部位であることが確認された。
 結合部位同定装置1の実施可能性を検証した実施例3を、図21から図23を参照して説明する。立体構造の画像作成には、PyMOL 1.5.0.3(シュレーディンガー株式会社)を用いた。なお、実施例3では、既知のタンパク質−低分子複合体の結晶構造において、低分子の周囲5Å以内に存在する注目物質の数を予め取得し、結合に関与すると見做されるものの数に関する取得結果とした。
 図21は、結合部位同定装置1を用いて同定された結合部位が示されているFactor Xaのapo構造と、Factor Xaと阻害剤の既知の複合体共結晶構造と、が重ね合わされたものの一例を示す図である。なお、結合部位同定装置1で結合部位を同定するにあたり、Factor Xaを標的物質とし、阻害剤を結合物質とし、6種類のアミノ酸(Met、Ala、His、Tyr、TrpおよびPhe)を注目物質とした。図21において、apo構造は灰色で示されており、複合体共結晶構造は緑色で示されている。結合部位同定装置1で同定された阻害剤の結合部位の空間(図21において赤色で示されている空間)は、複合体共結晶構造で確認される阻害剤(図21において青色で示されている)が存在する空間とほぼ同じであり、結合部位同定装置1はapo構造から阻害剤の結合部位を適切に同定することができた。
 図22は、結合部位同定装置1を用いて同定された結合部位が示されているGalectin−9 N−domainのapo構造と、Galectin−9 N−domainとラクトースの既知の複合体共結晶構造と、が重ね合わされたものの一例を示す図である。なお、結合部位同定装置1で結合部位を同定するにあたり、Galectin−9 N−domainを標的物質とし、ラクトースを結合物質とし、6種類のアミノ酸(Met、Ala、His、Tyr、TrpおよびPhe)を注目物質とした。図22において、apo構造は灰色で示されており、複合体共結晶構造は緑色で示されている。結合部位同定装置1で同定されたラクトースの結合部位の空間(図22において赤色で示されている空間)は、複合体共結晶構造で確認されるラクトース(図22において青色で示されている)が存在する空間とほぼ同じであり、結合部位同定装置1はapo構造からラクトースの結合部位を適切に同定することができた。
 図23は、結合部位同定装置1を用いて同定された結合部位が示されているAngiotensin I変換酵素 N−domainのapo構造と、Angiotensin I変換酵素 N−domainと阻害剤の既知の複合体共結晶構造と、が重ね合わされたものの一例を示す図である。なお、結合部位同定装置1で結合部位を同定するにあたり、Angiotensin I変換酵素 N−domainを標的物質とし、阻害剤を結合物質とし、6種類のアミノ酸(Met、Ala、His、Tyr、TrpおよびPhe)を注目物質とした。図23において、apo構造は灰色で示されており、複合体共結晶構造は緑色で示されている。結合部位同定装置1で同定された阻害剤の結合部位の空間(図23において赤色で示されている空間)は、複合体共結晶構造で確認される阻害剤(図23において青色で示されている)が存在する空間とほぼ同じであり、結合部位同定装置1はapo構造から阻害剤の結合部位を適切に同定することができた。
 本発明は、産業上の多くの分野、特に創薬を含む分子設計やタンパク質改変などの分野で広く実施することができ、極めて有用である。
1 結合部位同定装置
 10 制御部
   10a 同定部
     10a1 探索点設定部
     10a2 探索空間設定部
     10a3 物質数計算部
     10a4 度合い評価部
     10a5 探索点決定部
     10a6 ノイズ削除部
     10a7 探索点基準同定部
   10b 描画部
 11 通信インターフェース部
 12 記憶部
   12a 立体構造記憶部
   12b 実験結果記憶部
   12c 計算結果記憶部
   12d 評価結果記憶部
   12e 同定結果記憶部
 13 入出力インターフェース部
 14 入力装置
 15 出力装置
2 ネットワーク

Claims (14)

  1.  標的物質における結合物質の結合部位を同定する、制御部を備えた結合部位同定装置であって、
     前記制御部は、
     前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定手段を備えること
     を特徴とする結合部位同定装置。
  2.  前記探索空間の大きさは、前記結合物質の大きさに基づいて設定されたものであること
     を特徴とする請求項1に記載の結合部位同定装置。
  3.  前記結合物質の大きさは、前記結合物質の体積であること
     を特徴とする請求項2に記載の結合部位同定装置。
  4.  前記探索空間の大きさは、前記結合物質を包含する最小の球の半径に基づいて設定されたものであること
     を特徴とする請求項2に記載の結合部位同定装置。
  5.  前記探索空間の形状は、球、楕円体または多面体であること
     を特徴とする請求項1から4のいずれか一つに記載の結合部位同定装置。
  6.  前記同定手段は、
     前記立体構造に基づいて探索点を設定する探索点設定手段と、
     前記探索点設定手段で設定された前記探索点を基準とする前記探索空間を設定する探索空間設定手段と、
     前記立体構造に基づいて、前記探索空間設定手段で設定された前記探索空間に存在する前記注目物質の数を計算する物質数計算手段と、
     前記物質数計算手段で得られた計算結果と前記取得結果との一致または不一致の度合いを評価する度合い評価手段と、
     前記度合い評価手段で評価された前記度合いが所定の条件を満たす前記探索点を決定する探索点決定手段と、
     前記探索点決定手段で決定された前記探索点に基づいて前記結合部位を同定する探索点基準同定手段と、
     をさらに備えること
     を特徴とする請求項1から5のいずれか一つに記載の結合部位同定装置。
  7.  前記度合い評価手段は、前記計算結果に含まれる前記注目物質の数と前記取得結果に含まれる前記注目物質の数との差の絶対値を前記度合いとして計算し、
     前記所定の条件は、前記絶対値が所定値以下であるという条件であること
     を特徴とする請求項6に記載の結合部位同定装置。
  8.  前記取得結果に前記注目物質の数が複数パターン含まれている場合、前記度合い評価手段は、各々の前記パターンごとに前記絶対値を計算し、計算された前記絶対値の総和を前記度合いとして計算し、
     前記所定の条件は、前記総和が所定値以下であるという条件であること
     を特徴とする請求項7に記載の結合部位同定装置。
  9.  前記探索点設定手段は、前記立体構造に基づいて、相互作用エネルギーの観点から、分子が結合する可能性のある点を検出し、検出された当該点を前記探索点として設定すること
     を特徴とする請求項6から8のいずれか一つに記載の結合部位同定装置。
  10.  前記同定手段は、前記探索点決定手段で決定された前記探索点から、前記結合物質の大きさに基づいて、ノイズと見做されるものを削除するノイズ削除手段をさらに備え、
     前記探索点基準同定手段は、前記ノイズ削除手段が実行された後に残った前記探索点に基づいて前記結合部位を同定すること
     を特徴とする請求項6から9のいずれか一つに記載の結合部位同定装置。
  11.  前記取得結果に含まれる、前記注目物質について前記結合に関与すると見做されるものの数は、前記標的物質に含まれる前記注目物質の数と前記結合部位に存在する前記注目物質の数との比を考慮して定義された閾値に基づいて計算されたものであること
     を特徴とする請求項1から10のいずれか一つに記載の結合部位同定装置。
  12.  前記標的物質はタンパク質であり、前記結合物質は当該タンパク質と結合するものであり、前記注目物質は1種類または複数種類のアミノ酸であること
     を特徴とする請求項1から11のいずれか一つに記載の結合部位同定装置。
  13.  標的物質における結合物質の結合部位を同定する、制御部を備えた情報処理装置で実行される結合部位同定方法であって、
     前記制御部で実行される、
     前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定ステップを含むこと
     を特徴とする結合部位同定方法。
  14.  標的物質における結合物質の結合部位を同定する、制御部を備えた情報処理装置に実行させるための結合部位同定プログラムであって、
     前記制御部に実行させるための、
     前記標的物質の立体構造と、注目物質について前記結合に関与すると見做されるものの数に関する取得結果とに基づいて、前記標的物質の表面を含む探索空間から、前記注目物質の数について前記取得結果と一致すると見做される箇所を探索し、探索された当該箇所に基づいて前記結合部位を同定する同定ステップを含むこと
     を特徴とする結合部位同定プログラム。
PCT/JP2013/085330 2012-12-25 2013-12-25 結合部位同定装置、結合部位同定方法および結合部位同定プログラム Ceased WO2014104401A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2014554633A JP6344768B2 (ja) 2012-12-25 2013-12-25 結合部位同定装置、結合部位同定方法および結合部位同定プログラム

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2012-281823 2012-12-25
JP2012281823 2012-12-25

Publications (1)

Publication Number Publication Date
WO2014104401A1 true WO2014104401A1 (ja) 2014-07-03

Family

ID=51021455

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2013/085330 Ceased WO2014104401A1 (ja) 2012-12-25 2013-12-25 結合部位同定装置、結合部位同定方法および結合部位同定プログラム

Country Status (2)

Country Link
JP (1) JP6344768B2 (ja)
WO (1) WO2014104401A1 (ja)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2009086331A1 (en) * 2007-12-20 2009-07-09 Georgia Tech Research Corporation Elucidating ligand-binding information based on protein templates
US20100138205A1 (en) * 2008-10-10 2010-06-03 Los Alamos National Security, Llc Stochastic molecular binding simulation

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
KEIJI KAKUMOTO: "A Statistical Analysis of an Effective Method to Conduct In Silico Screening for Active Compounds", 12 February 2005 (2005-02-12), Retrieved from the Internet <URL:http://www.rs.kagu.tus.ac.jp/yoshilab/iyaku/2004ME/kakumoto.pdf> *
SHINJI SOGA: "Tanpakushitsu Hyomenjo no Aminosan Sosei ni yoru Ligand Ketsugo Bui Yosoku Hoho no Kaihatsu", NAGOYA UNIVERSITY HAKASE GAKUI RONBUN, 25 March 2011 (2011-03-25), pages 1 - 116 *

Also Published As

Publication number Publication date
JPWO2014104401A1 (ja) 2017-01-19
JP6344768B2 (ja) 2018-06-20

Similar Documents

Publication Publication Date Title
Pfab et al. DeepTracer for fast de novo cryo-EM protein structure modeling and special studies on CoV-related complexes
Schmidtke et al. Large-scale comparison of four binding site detection algorithms
Ortiz et al. MAMMOTH (matching molecular models obtained from theory): an automated method for model comparison
Zhang et al. Identification of cavities on protein surface using multiple computational approaches for drug binding site prediction
Deutsch et al. Trans‐Proteomic Pipeline, a standardized data processing pipeline for large‐scale reproducible proteomics informatics
Huang et al. SGPPI: structure-aware prediction of protein–protein interactions in rigorous conditions with graph convolutional network
Gill et al. Emerging role of bioinformatics tools and software in evolution of clinical research
Wierbowski et al. Cross‐docking benchmark for automated pose and ranking prediction of ligand binding
Sinha et al. Docking by structural similarity at protein‐protein interfaces
JP5990862B2 (ja) 承認予測装置、承認予測方法、および、プログラム
Yang et al. SPOT-Seq-RNA: predicting protein–RNA complex structure and RNA-binding function by fold recognition and binding affinity prediction
JP5905781B2 (ja) 相互作用予測装置、相互作用予測方法、および、プログラム
Carbery et al. Learnt representations of proteins can be used for accurate prediction of small molecule binding sites on experimentally determined and predicted protein structures
Marin et al. FROST: a filter‐based fold recognition method
Gao et al. EAGLE: an algorithm that utilizes a small number of genomic features to predict tissue/cell type-specific enhancer-gene interactions
Ghersi et al. Beyond structural genomics: computational approaches for the identification of ligand binding sites in protein structures
Gorelik et al. High quality binding modes in docking ligands to proteins
Moshawih et al. Consensus holistic virtual screening for drug discovery: a novel machine learning model approach
Waldron et al. Meta-analysis in gene expression studies
Beltrao et al. Comparative genomics and disorder prediction identify biologically relevant SH3 protein interactions
Wu et al. Cosbin: cosine score-based iterative normalization of biologically diverse samples
Konc et al. ProBiS tools (algorithm, database, and web servers) for predicting and modeling of biologically interesting proteins
Morita et al. Highly accurate method for ligand‐binding site prediction in unbound state (apo) protein structures
Fabris et al. Elucidating the higher‐order structure of biopolymers by structural probing and mass spectrometry: MS3D
Tang et al. Virtual screening for lead discovery

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13868104

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2014554633

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13868104

Country of ref document: EP

Kind code of ref document: A1