WO2014046284A1 - 二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置 - Google Patents

二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置 Download PDF

Info

Publication number
WO2014046284A1
WO2014046284A1 PCT/JP2013/075702 JP2013075702W WO2014046284A1 WO 2014046284 A1 WO2014046284 A1 WO 2014046284A1 JP 2013075702 W JP2013075702 W JP 2013075702W WO 2014046284 A1 WO2014046284 A1 WO 2014046284A1
Authority
WO
WIPO (PCT)
Prior art keywords
gene cluster
gene
genes
region
reference value
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2013/075702
Other languages
English (en)
French (fr)
Inventor
町田 雅之
舞子 梅村
英明 小池
至 竹田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
National Institute of Advanced Industrial Science and Technology AIST
Original Assignee
National Institute of Advanced Industrial Science and Technology AIST
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by National Institute of Advanced Industrial Science and Technology AIST filed Critical National Institute of Advanced Industrial Science and Technology AIST
Priority to JP2014536957A priority Critical patent/JP5946149B2/ja
Priority to US14/427,349 priority patent/US20150310168A1/en
Publication of WO2014046284A1 publication Critical patent/WO2014046284A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • G16B30/10Sequence alignment; Homology search
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B10/00ICT specially adapted for evolutionary bioinformatics, e.g. phylogenetic tree construction or analysis
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B40/00ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding

Definitions

  • the present invention relates to a method, a prediction program, and a prediction apparatus for predicting a gene cluster including a secondary metabolic gene among gene clusters composed of a plurality of genes.
  • Secondary metabolites are highly likely to have physiological activity and are extremely useful as lead compounds for pharmaceuticals. Secondary metabolites are diverse and have been discovered from various species such as actinomycetes, fungi, plants, etc., but the conditions for expression are often special and unknown, and many secondary properties with useful properties Metabolites are thought to be sleeping without being discovered. Even if it is discovered, it is a problem in use that it is difficult to produce a stable and sufficient amount.
  • secondary metabolic genes genes involved in the discovery of useful unknown secondary metabolites and the biosynthesis of secondary metabolites (secondary metabolic genes) can be identified.
  • secondary metabolic genes have been difficult to identify with high accuracy using current comparative genomic analysis methods. This is because secondary metabolic genes often contradict evolutionary phylogenetic trees of the genus / species and there are many unknown genes whose functions have not been elucidated at all.
  • core genes such as polyketide synthase gene (PKS gene) and nonribosomal peptide synthetase gene (NRPS gene) It was a method of predicting as a cluster including genes associated therewith.
  • core genes such as polyketide synthase gene (PKS gene) and nonribosomal peptide synthetase gene (NRPS gene)
  • PKS gene polyketide synthase gene
  • NRPS gene nonribosomal peptide synthetase gene
  • SMURF described in Non-Patent Document 1
  • antiSMASH described in Non-Patent Document 2
  • CLUSEAN described in Non-Patent Document 3
  • ClustScan described in Non-Patent Document 4.
  • clusters detected by these methods are limited to secondary metabolic gene clusters having a core gene, and are only a part of the entire cluster including secondary metabolic genes. In other words, with these methods, it was impossible to predict a secondary metabolic gene cluster that does not include the core gene that is expected to account for more than half of the total.
  • CLUSEAN A computer-based framework for the automated analysis of bacterial secondary metabolite biosynthetic gene clusters.JOURNAL OF BIOTECHNOLOGY. 140, Starcevic Antonio; Zucko Jurica; Simunkovic Jurica; et al.
  • ClustScan an integrated program package for the semi-automatic annotation of modular biosynthetic gene clusters and in sil 6821 2008)
  • the present invention provides a method, a prediction program, and a prediction apparatus capable of predicting a gene cluster including a secondary metabolic gene with high accuracy without depending on information on the core gene in view of the above-described actual situation. With the goal.
  • the present invention that has achieved the above-described object includes the following.
  • the gene cluster is a gene cluster including a secondary metabolic gene (1)
  • the synteny-like region includes at least two ortholog genes, and the distance between adjacent ortholog genes is a region within a predetermined distance in each of the genomic base sequence information and the other genomic base sequence information.
  • the synteny-like region includes at least one pair of genomic base sequence information used for comparison and third genomic base sequence information different from the pair of genomic base sequence information.
  • the prediction method according to (1) wherein a synteny region / non-synteny region is determined in advance, and the synteny region is defined as a synteny-like region.
  • the number of homologous genes included in the identified gene cluster and / or the number of all genes included in the identified gene cluster are compared with a predetermined reference value, and the homologous gene Performing the above-described process for determining whether the gene cluster is a gene cluster including a secondary metabolic system gene for a gene cluster in which the number of genes is greater than or equal to a reference value and / or a gene cluster in which the number of all genes is less than the reference value (1)
  • the prediction method according to (1) The prediction method according to (1).
  • the number of all genes included in the specified gene cluster is compared with a predetermined reference value, or the length of the specified gene cluster is compared with a predetermined reference value. Then, for a gene cluster whose total number or length is less than the reference value, a step is performed to determine whether the gene cluster is a gene cluster including a secondary metabolic gene, and the gene cluster includes a secondary metabolic gene. In the step of determining whether it is a cluster, the gene cluster is corrected by adding a gene adjacent to the gene cluster so that the number of genes included in the determination target gene cluster is equal to the number of the reference values.
  • a prediction method according to (1) characterized in that a synteny-like region is identified for a modified gene cluster composed of a reference number of genes. .
  • the step of specifying the gene cluster After the step of specifying the gene cluster, the number of all genes included in the specified gene cluster is compared with a predetermined reference value, or the length of the specified gene cluster is compared with a predetermined reference value. Then, for a gene cluster whose total number or length is less than the reference value, a step is performed to determine whether the gene cluster is a gene cluster including a secondary metabolic gene, and the gene cluster includes a secondary metabolic gene. In the step of determining whether it is a cluster, a predetermined number of genes or a region of a predetermined length is added to the determination target gene cluster to correct the gene cluster, and the corrected gene cluster is a synteny-like region. (1) The prediction method according to (1).
  • the gene cluster is specified by tracing back from the cell showing the maximum score (1 ) The prediction method described.
  • the sequence of other genes is stored by substituting 0 for the score of the cell included in the specified gene cluster and tracing back the Smith-Waterman matrix.
  • the region according to claim 12, wherein the region in which the gene sequence is stored is identified again by the Smith-Waterman algorithm for the identified region, and the region is identified as a gene cluster.
  • the number of all genes included in the specified gene cluster is compared with a predetermined reference value, or the length of the specified gene cluster is compared with a predetermined reference value. And adding a predetermined number of genes or a region of a predetermined length to the gene cluster to expand the gene cluster to be the reference value, For each gene constituting the expanded gene cluster, a positive score is obtained if there is a homology with the genes constituting the gene cluster to be compared in other genomic nucleotide sequence information, and a minus is obtained if there is no homology.
  • FIG. 5 is a flowchart relating to a method for predicting gene clusters including secondary metabolic genes according to the present invention.
  • the prediction method which concerns on this invention it is a conceptual diagram of the matrix produced in Smith-Waterman algorithm at the time of specifying a gene cluster.
  • the prediction method which concerns on this invention it is a flowchart which shows the process until it identifies a gene cluster, performs an ortholog check to the identified gene cluster, and finally identifies the gene cluster containing a secondary metabolic system gene.
  • It is a schematic diagram for demonstrating the step which performs an ortholog check by applying the prediction method which concerns on this invention.
  • the method for predicting a gene cluster including a secondary metabolic gene specifies a gene cluster based on a sequence of genes in the compared genome by using the result of homology search for genes contained in at least a pair of genomes. And a step of determining whether the identified gene cluster is a gene cluster including a secondary metabolic gene (FIG. 1).
  • secondary metabolic genes mean genes involved in the biosynthesis of secondary metabolites.
  • Secondary metabolites are metabolites that are not directly involved in the life activity of an organism.
  • a metabolite will be comprised from a primary metabolite and a secondary metabolite.
  • the secondary metabolite can be said to be a metabolite excluding the primary metabolite.
  • the primary metabolite is a substance that directly participates in the life activity of an organism, and means, for example, sugar, amino acid, lipid, and nucleic acid.
  • secondary metabolites may be defined as substances other than sugars, amino acids, lipids and nucleic acids among metabolites. Examples of secondary metabolites include antibiotics, alkaloids, terpenoids, flavonoids, polyketides, phenols, glycosides, and special amino acids that do not constitute proteins.
  • genes involved in biosynthesis of secondary metabolites include genes encoding enzymes involved in secondary metabolite assimilation or catabolism, and proteins involved in translocation / accumulation of secondary metabolites. And a gene encoding a protein involved in regulation of the expression of these genes.
  • secondary metabolic genes include genes involved in biosynthesis of polyketides, non-ribosomal peptides, alkaloids, terpenoids, flavonoids, and other compounds not belonging to primary metabolism. It should be noted that the gene cluster predicted by the prediction method according to the present invention does not necessarily include these specifically exemplified secondary metabolic genes, but may also include other secondary metabolic genes.
  • a gene cluster is a group of a plurality of genes included in a predetermined continuous region, and a group of a plurality of genes whose arrangement is preserved between a plurality of genomes (for example, between a pair of genomes).
  • the continuous region means a region included in the whole genome or a part of the genome composed of nucleic acids such as chromosomes and mitochondria. That is, the gene cluster means a group of a plurality of genes whose arrangement is preserved in a continuous region included in the whole genome or a part of the genome.
  • Genome base sequence information is text data in which four types of nucleotides consisting of adenine, guanine, cytosine, and thymine are expressed as A, G, C, and G, respectively. Genome base sequence information is expressed in the direction from the 5 ′ end to the 3 ′ end.
  • one or both of the pair of genome base sequence information may use data acquired from a database storing various genome base sequence information or the like, or unknown or publicly known by applying DNA sequencing technology Data obtained from other organisms may be used.
  • DNA sequencing technology for example, all methods described in Chapter 11 of Molecular-Cloning, A-Laboratory-Manual, Fourth-Edition (Cold-Spring-Harbor-Laboratoty-Press) can be applied.
  • the prediction method according to the present invention is not limited to biological species at all, and gene clusters including secondary metabolic genes can be predicted.
  • the biological species include plants, bacteria, actinomycetes, fungi, filamentous fungi, mushrooms and the like.
  • the genome base sequence information may be data in which the species of origin is not clarified.
  • a base sequence can be determined for DNA directly extracted from an environment such as soil, sludge, lake water, and seawater without culturing, so-called environmental DNA, and this can be used as genome base sequence information. That is, in the prediction method according to the present invention, a gene cluster including secondary metabolic genes existing in environmental DNA can be predicted.
  • a gene cluster from at least a pair of genomic base sequence information, first, based on the genomic base sequence information, the arrangement of a plurality of genes existing in the pair of genomic base sequence information is compared, and the gene sequence is stored. Identify the region.
  • the amino acid sequence is used as a query sequence (query sequence), and the homology is searched using the amino acid sequence related to the gene included in the other genome base sequence information as a database sequence.
  • query sequence amino acid sequence related to the gene included in the other genome base sequence information
  • homology search conventionally known homology analysis software such as Blastp, FASTA and Clustal can be used.
  • homology search is similarly performed by replacing the genome base sequence information as the query sequence and the genome base sequence information as the database sequence.
  • genes having high sequence similarity can be mutually identified between a pair of genome base sequence information.
  • a threshold can be set for a value indicating sequence similarity, and a combination of genes exceeding the threshold can be specified as a homologous gene.
  • combinations that satisfy a predetermined standard can be specified as orthologous genes.
  • an ortholog gene is defined as a gene having homology formed by differentiation of a single gene having a common origin.
  • examples of the value indicating sequence similarity include values such as E value (e-value), bits value (bits), and amino acid identity (Identities) in Blast search. Therefore, a combination of genes can be specified as a homologous gene by setting a threshold value for one or more of these values. For example, specifically, a homology search between a query sequence and a database sequence, and a homology search performed by exchanging the query sequence and the database sequence (these are collectively referred to as a “set of homology searches”)
  • the threshold value can be set to, for example, 1.0e-20, preferably 1.0e-15, particularly preferably 1.0e-10. Then, in a set of homology searches, a combination of genes each having an E value equal to or less than a threshold value can be specified as a homologous gene in both genome base sequence information.
  • a standard is set so that a combination that matches the above-mentioned definition of the ortholog gene can be selected.
  • a standard is set so that a combination that matches the above-mentioned definition of the ortholog gene can be selected.
  • the method for specifying a combination of ortholog genes from a combination of homologous genes specified by a set of homology searches is not limited to the above-described method.
  • the arrangement of the genes in the pair of genome base sequence information is compared, and the region where the gene sequence is stored is specified.
  • a plurality of genes existing in the genomic base sequence information are regarded as character strings by considering the genes as characters, and character string search and character Algorithms that compare the similarity of columns can be applied.
  • Examples of algorithms that can be used in this process include a Smith-Waterman algorithm, a Needleman-Wunsch algorithm, and a k-tuple method that are character string search algorithms.
  • a matrix (two-dimensional) of (J + 1) ⁇ (I + 1) is created (FIG. 2).
  • a score calculated according to the following procedure is entered in each cell of the matrix. That is, for each cell, when x i and y j are homologous
  • a set of coordinates of cells having high homology with x i and y j is R 0 .
  • gap and mismatch are penalty scores, which are set in the range of about -0.4 to -0.1, but preferably both are -0.2.
  • R 0 is in a region where gene sequences are stored and is a set of coordinates of highly homologous gene pairs.
  • this R 0 is a gene cluster, that is, a group of a plurality of genes whose arrangement is preserved in a pair of genome base sequence information.
  • the prediction method according to the present invention it is possible to determine whether the gene cluster R 0 specified as described above is a gene cluster including a secondary metabolic gene, as will be described in detail later.
  • the prediction method according to the present invention is not limited to the gene cluster R 0 identified as described above, and further different gene clusters R ′ 0 are identified by the following procedure, and two of these gene clusters R ′ 0 are identified. It can also be determined whether the gene cluster includes a secondary metabolic gene (FIG. 3).
  • 0 is assigned to all the cells indicated by the coordinates included in the set obtained by the previous operation (initially R 0 ). That is,
  • the synteny-like region in the identified gene cluster can be evaluated using the number and distance of orthologous genes included in the gene cluster.
  • the gene cluster or R 0, R '0, R ''0 ... gene clusters shown in shown in R 0, a combination of a homologous gene, for example, two or more, and preferably contains three or more
  • the total number of genes is, for example, within 50, preferably within 40, and more preferably within 35.
  • an ortholog check is performed on a gene cluster that satisfies the above-described conditions (for example, condition * 2).
  • the gene cluster is corrected so that the number of genes included in the gene cluster becomes the total number of genes under the above conditions (for example, 35 genes in the case of condition * 2) (FIG. 4).
  • genes existing in the vicinity of the gene cluster are added to the gene cluster specified in the above step so that, for example, a total of 35 genes are obtained. For example, by adding the same number of genes from both ends of the gene cluster specified in the above step, it can be corrected to a gene cluster consisting of a total of 35 genes, for example.
  • the number of genes to be expanded is an odd number, there is no particular limitation, but one more or less genes may be added to the 3 ′ end of the gene cluster.
  • the total number of genes is set to 35 genes, for example, errors at the boundaries of the gene clusters identified in the above process can be taken into account, and the distribution of orthologous gene pairs in the nearby region is also averaged. Can be evaluated.
  • the number of genes is partially omitted for simplicity.
  • the synteny-like region means a region containing a plurality of orthologous genes, and the distance between adjacent orthologous genes (which may contain other genes in between) is not more than a reference value.
  • the reference value can be, for example, 10 to 30 kb, 10 to 20 bk, or 10 kb.
  • the regions 1 to A and 2 to a are synteny-like regions. It can be. Similarly, all the combinations of orthologous genes contained in X and Y are confirmed to be synteny-like regions. There may be a plurality of synteny-like regions.
  • the synteny-like regions identified in X and Y are represented by X and Y subsets xSB and ySB, respectively.
  • the gene cluster consisting of x i and the gene cluster consisting of y j It is determined that the gene cluster is included.
  • the predetermined ratio is not particularly limited, but may be 30%, 25%, or 20%. For example, when the predetermined ratio is 25% (condition * 3)
  • the method for predicting a gene cluster including a secondary metabolic gene is not limited to a method using a synteny-like region specified according to the above procedure, and a synteny-like region specified by another method may be used.
  • a method for specifying a synteny-like region for example, a method of predetermining a synteny region and a non-synteny region using genomic base sequence information and annotation information of different species can be mentioned.
  • a gene cluster containing secondary metabolic genes can be predicted in the same manner as described above. That is, even if the method for specifying the synteny-like region from the synteny region is used as described above, the method similar to the determination of the synteny-like region can be applied as described with reference to FIG. That is, orthologous genes for genes predicted on two types of genomes are determined in advance, the synteny region defined as described above is specified, and regions other than the synteny region in the genomic nucleotide sequence information are defined as non-synteny regions. Define.
  • R 0 gene cluster or R 0 shown by a, R '0, R'' 0 ... gene clusters shown in extend the cluster length by the same method as described above (e.g., 35 genes)
  • the synteny region is less than 25% of the total, it can be predicted as a gene cluster including secondary metabolic genes.
  • This method may be able to predict a gene cluster containing a secondary metabolic gene with higher accuracy than the method of identifying a synteny-like region after detecting a gene cluster as described above.
  • a gene cluster containing a secondary metabolic gene with higher accuracy than the method of identifying a synteny-like region after detecting a gene cluster as described above.
  • closely related species such as A. flavus and A. oryzae
  • there is no second gene cluster that is highly homologous to the same gene cluster in other strains of A.usflavus or A.oryzae so it is considered that the aflatoxin biosynthesis cluster present in A. flavus is not detected. .
  • a synteny region is defined as a region of a gene that exists in common among relatively related species such as the genus Aspergillus.
  • the gene cluster to be evaluated is limited based on the number of genes included in the gene cluster as described above.
  • the gene cluster to be evaluated may be limited based on the length of the gene cluster. That is, the length of a gene cluster can be compared with a predetermined reference value, and an ortholog check can be performed for a gene cluster having a length shorter than the reference value.
  • the reference value is not particularly limited, but for example, 125 kb (equivalent to about 50 genes), preferably 100 kb (equivalent to about 40 genes), more preferably 87.5 kb (about 35 genes). Equivalent).
  • the number of genes included in the gene cluster is set to a predetermined number (for example, 35) before the ortholog check is performed.
  • the gene cluster was corrected.
  • the gene cluster is corrected by adding a predetermined number of genes or a predetermined length region to the gene cluster, and the corrected gene cluster An ortholog check may be performed for.
  • Examples of the method for correcting gene clusters include a method for correcting the boundaries of gene clusters, as will be described below. That is, it is a method of correcting the boundaries of gene clusters indicated by the identified R 0 , R ′ 0 , R ′′ 0 .
  • To correct the boundaries of a gene cluster is to determine whether the gene cluster specified by the method described in the section ⁇ Specify gene cluster> described above includes a gene that exists outside the gene cluster. It is synonymous.
  • the number of genes included in the gene cluster is 15 as described above.
  • the gene cluster is expanded to ⁇ 65 genes, more preferably 35 genes (not limited to 35 genes).
  • a positive score is given when a gene with high homology exists in the gene cluster to be compared, and a negative score is obtained when there is no gene with high homology.
  • the scores given to the genes are summed in order from the genes located at the center of the expanded gene cluster, and the total score value is given to each gene.
  • the gene having the maximum score given to each gene included in the expanded gene cluster is identified, and the identified gene is set as the boundary of the gene cluster. This process may leave the genes that are the boundaries of the gene clusters unmodified and remain in the original gene clusters.
  • a one-dimensional array SC composed of n (X) elements was prepared.
  • Each element of this array can contain, for example, a score calculated according to the following formula: When x i is homologous to at least one of y c , y c-1 ,... y d-1 , y d
  • the element numbers of the elements having the maximum scores in the respective ranges (1) and (2) are set as i start and i stop .
  • the same operation is performed on the set Y.
  • the i start and i stop specified in this way are used as the boundaries of the gene cluster. That is, the gene clusters whose boundaries are corrected are
  • negative values are, for example, ⁇ 0.1, ⁇ 0.2, ⁇ 0.3 , -0.4, -0.5, or -1.
  • the method for predicting a gene cluster including a secondary metabolic gene according to the present invention described above includes an input means such as a mouse and a keyboard, a central processing means (CPU), a storage means including a volatile and / or nonvolatile memory. And a computer having output means such as a display. At this time, the computer is preferably connected to a storage device such as an external database or an external computer system via a communication network such as the Internet or an intranet. That is, the prediction method according to the present invention can be provided as a prediction program capable of predicting a gene cluster including a secondary metabolic system gene using the computer device having the above configuration. In other words, the computer on which this prediction program is installed serves as a prediction device for gene clusters including secondary metabolic genes.
  • a pair of genome base sequence information may be input from an external storage means or computer system to the computer via the communication network, or via a DNA sequencer and an interface. May be connected and input to the computer.
  • a pair of genome base sequence information may be read into a computer using a storage medium such as a DVD or CD.
  • the central processing unit can execute the homology search for the pair of genome base sequence information, and the result of the homology search can be stored in the storage device.
  • the above-described ⁇ identification of gene cluster> procedure and ⁇ determination of gene cluster including secondary metabolic gene> are realized by software implementing a character string search algorithm such as the Smith-Waterman algorithm. be able to.
  • Example 1 In this example, eight types of genome data sets were used. Among them, Aspergillus oryzae data used was equivalent to that registered in GenBank (AP007150-AP007177). Aspergillus flavus data was downloaded from GenBank in genbank format. (genbank accession EQ963472 to EQ963493). Data on Aspergillus fumigatas, Aspergillus nidulans, Aspergillus terreus, Magnaporthe grisea, Fusarium graminearum and Chaetomium globosum were downloaded from BROAD INSTITUTE and used.
  • homologous genes are those whose E value (E-value) is 1.0e-10 or less in the homology search.
  • E-value E value
  • the conservation of the gene sequence was verified by the Smith-Waterman algorithm, and gene clusters R 0 , R ′ 0 , R ′′ 0 .
  • the number of combinations of homologous genes contained in the identified gene cluster was set to 3 or more, and the total number of genes was set to less than 35.
  • the synteny-like region is a region containing a plurality of orthologous genes, and the distance between adjacent orthologous genes (which may contain other genes between them) is 10 bk or less, 20 kb or less, or 30 kb. The following areas were defined.
  • the number of genes (number of elements in the subset) included in the synteny-like region (X and Y subsets xSB and ySB) is less than 25% (less than 8) of 35 genes.
  • the original gene cluster was predicted as a gene cluster containing secondary metabolic genes.
  • the number of gene clusters containing secondary metabolic genes predicted using 10 genome base sequences of filamentous fungi that have been subjected to genome analysis such as A. flavus and A. oryzae is shown. It was shown in 1.
  • Table 1-1 shows the results when the distance between adjacent orthologous genes in the synteny-like region is defined as 10 bk or less, and Table 1-2 shows the results when the distance is 20 kb or less.
  • -3 shows the result when the distance is 30 kb or less. From this result, it was clarified that if the distance between the orthologue genes adjacent to the synteny-like region is defined as 10 to 30 bk or less, the result does not vary greatly.
  • Table 2 shows the result of calculating the ratio of those predicted to be gene clusters containing secondary metabolic genes in this example and containing the Q gene.
  • the Q gene is a gene classification that is considered as a secondary metabolic system in the functional classification of the COG (Cluster of Orthologous Group).
  • Example 2 In this example, in the same manner as in Example 1, the gene sequence conservation was verified by the Smith-Waterman algorithm, and gene clusters R 0 , R ′ 0 , R ′′ 0 .
  • +1 is given to each gene included in the gene cluster expanded to 35 genes when there is a homologous gene, If not, -0.3 is given, and the total score is calculated from the center of the expanded gene cluster, and the gene taking the maximum of the total value is used as the boundary of the gene cluster.
  • gene clusters including secondary metabolic genes were predicted.
  • Table 3 shows a part of the gene cluster including the secondary metabolic genes predicted in this example. Further, similarly to Example 1, Table 4 shows gene clusters including secondary metabolic genes predicted without correcting the boundaries of the gene clusters.
  • the error column indicates how much the predicted gene cluster is in the upstream direction (5 ′ side) and downstream direction (3 ′ side) with respect to the actual gene cluster of the secondary metabolic system gene. The number of genes is shown whether it is wrong.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Medical Informatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biotechnology (AREA)
  • Evolutionary Biology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Theoretical Computer Science (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Analytical Chemistry (AREA)
  • Chemical & Material Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Molecular Biology (AREA)
  • Physiology (AREA)
  • Animal Behavior & Ethology (AREA)
  • Artificial Intelligence (AREA)
  • Bioethics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Epidemiology (AREA)
  • Evolutionary Computation (AREA)
  • Public Health (AREA)
  • Software Systems (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Description

二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置
 本発明は、複数の遺伝子から構成される遺伝子クラスタのうち二次代謝系遺伝子を含む遺伝子クラスタを予測する方法、予測プログラム及び予測装置に関する。
 二次代謝物質は、生理活性を有する可能性が高く、医薬のリード化合物として極めて有用である。二次代謝物質は多様で、放線菌、真菌、植物などの様々な生物種から発見されているが、発現する条件が特殊で知られていないことが多く、有用な性質を持つ多数の二次代謝物質が発見されないままに眠っていると考えられている。また発見されたとしても、安定で十分な量の生産が困難であることが利用の際の問題である。
 一方、近年、DNAシークエンス技術の革新的な発展により、様々な生物種、特に微生物のゲノム情報の蓄積は加速度的に増加しており、数年後には数千種以上の微生物のゲノム塩基配列が明らかになることは確実である。また、DNAシークエンス技術によれば、ゲノム情報が未知の生物についても迅速、且つ低コストにゲノム情報を取得することができる。このようなゲノム情報の蓄積、ゲノム情報の簡便な解析により、ホールゲノム解析、シンテニー解析といった比較ゲノム解析を広範な生物種に対して行うことができる。
 このような状況であれば、詳細かつ膨大なゲノム情報を収集して構築されたデータベースを利用し、更に二次代謝物質の構造、多様性、生物界での分布などに関する情報を利用することによって、有用な未知の二次代謝物質の発見及び二次代謝物質の生合成に関与する遺伝子(二次代謝系遺伝子)を同定できると期待される。しかし、二次代謝系遺伝子は、現状の比較ゲノム解析方法を利用して高精度に同定することが困難であった。これは、二次代謝系遺伝子は、属・種の進化系統樹と矛盾することが多い上、機能が全く解明されていない未知の遺伝子が多数存在するためである。
 従来、二次代謝系遺伝子を解析する手法として、ポリケチドシンターゼ遺伝子(PKS遺伝子)やノンリボゾーマルペプチドシンテターゼ遺伝子(NRPS遺伝子)等、既知の配列相同性が高い遺伝子(コア遺伝子)の検出を基盤として、これに付随する遺伝子を含めたクラスタとして予測する方法であった。具体的には、非特許文献1に記載されたSMURF、非特許文献2に記載されたantiSMASH、非特許文献3に記載されたCLUSEAN及び非特許文献4に記載されたClustScanを挙げることができる。
 しかし、これら方法で検出されるクラスタは、コア遺伝子を有する二次代謝系遺伝子クラスタに限定され、二次代謝系遺伝子を含むクラスタ全体の一部でしかない。換言すると、これらの方法では、全体の半数以上を占めると予想されるコア遺伝子を含まない二次代謝系遺伝子クラスタを予測することは不可能であった。
Khaldi Nora; Seifuddin Fayaz T.; Turner Geoff; et al. SMURF: Genomic mapping of fungal secondary metabolite clusters. FUNGAL GENETICS AND BIOLOGY. 47, 9, 736-741 (2010) Medema Marnix H.; Blin Kai; Cimermancic Peter; et al. antiSMASH: rapid identification, annotation and analysis of secondary metabolite biosynthesis gene clusters in bacterial and fungal genome sequences. NUCLEIC ACIDS RESEARCH. 39, 339-346 (2011) Weber T.; Rausch C.; Lopez P.; et al. CLUSEAN: A computer-based framework for the automated analysis of bacterial secondary metabolite biosynthetic gene clusters. JOURNAL OF BIOTECHNOLOGY. 140, 1-2, 13-17 (2009) Starcevic Antonio; Zucko Jurica; Simunkovic Jurica; et al. ClustScan: an integrated program package for the semi-automatic annotation of modular biosynthetic gene clusters and in silico prediction of novel chemical structures. NUCLEIC ACIDS RESEARCH. 36, 21, 6882-6892 (2008)
 本発明は、上述したような実情に鑑み、コア遺伝子に関する情報に依存せず、二次代謝系遺伝子を含む遺伝子クラスタを高精度に予測することができる方法、予測プログラム及び予測装置を提供することを目的とする。
 上述した目的を達成した本発明は以下を包含する。
 (1)少なくとも一対のゲノム塩基配列情報に含まれる遺伝子に関して相互に相同性検索し、これらゲノム塩基配列情報間で相同遺伝子の組み合わせ、相同遺伝子の組み合わせのなかでオーソログ遺伝子の組み合わせを同定する工程と、上記相同性検索の結果に基づいて他のゲノム塩基配列情報との間で遺伝子の並びが保存されている領域を遺伝子クラスタとして特定する工程と、上記工程で特定された遺伝子クラスタのうち、上記相同性検索の結果のうちオーソログ遺伝子の存在に基づいてシンテニー様領域を特定し、当該シンテニー様領域が占める割合に基づいて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程とを含む二次代謝系遺伝子を含む遺伝子クラスタの予測方法。
 (2)上記シンテニー様領域に含まれる遺伝子数が上記遺伝子クラスタ全体に含まれる遺伝子数に対して、所定の割合以下である場合に当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであると判定することを特徴とする(1)記載の予測方法。
 (3)上記所定の割合を25%とすることを特徴とする(2)記載の予測方法。
 (4)上記シンテニー様領域は、少なくとも2以上のオーソログ遺伝子を含み、隣接するオーソログ遺伝子の距離が上記ゲノム塩基配列情報及び上記他のゲノム塩基配列情報のそれぞれにおいて所定の距離以内にある領域とすることを特徴とする(1)記載の予測方法。
 (5)上記所定の距離は、10~30kbとすることを特徴とする(4)記載の予測方法。
 (6)上記シンテニー様領域は、比較に用いる少なくとも一対のゲノム塩基配列情報のうちの1種のゲノム塩基配列情報と、これら一対のゲノム塩基配列情報とは異なる第3のゲノム塩基配列情報とを用いてシンテニー領域/非シンテニー領域をあらかじめ決定し、このシンテニー領域をシンテニー様領域とすることを特徴とする(1)記載の予測方法。
 (7)上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる相同遺伝子の数及び/又は特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較し、相同遺伝子の数が基準値以上である遺伝子クラスタ、及び/又は全遺伝子の数が基準値未満の遺伝子クラスタについて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記工程を行うことを特徴とする(1)記載の予測方法。
 (8)相同遺伝子の数に関する上記基準値を3個とし、全遺伝子の数に関する上記基準値を35個とすることを特徴とする(7)記載の予測方法。
 (9)上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに含まれる遺伝子の数が上記基準値の個数となるように、当該遺伝子クラスタに隣接する遺伝子を追加して当該遺伝子クラスタを修正し、当該基準値の個数の遺伝子からなる修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする(1)記載の予測方法。
 (10)全遺伝子の数に関する上記基準値を35個とすることを特徴とする(9)記載の予測方法。
 (11)上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに対して、所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタを修正し、修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする(1)記載の予測方法。
 (12)上記遺伝子クラスタを特定する工程では、Smith-Watermanアルゴリズムで作成されるSmith-Watermanマトリックスにおいて、最大のスコアを示すセルからトレースバックすることで遺伝子クラスタを特定することを特徴とする(1)記載の予測方法。
 (13)上記遺伝子クラスタを特定する工程では、上記特定された遺伝子クラスタに含まれるセルのスコアに0を代入し、このSmith-Watermanマトリックスについてトレースバックすることで他の遺伝子の並びが保存されている領域を特定し、特定された領域について再度Smith-Watermanアルゴリズムにて遺伝子の並びが保存された領域を特定し、これを遺伝子クラスタとして特定することを特徴とする請求項(12)記載の予測方法。
 (14)上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、遺伝子クラスタに対して所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタが当該基準値となるように拡張し、
 拡張した遺伝子クラスタを構成する各遺伝子について、他のゲノム塩基配列情報における比較対象の遺伝子クラスタを構成する遺伝子との間で相同性がある場合にはプラスのスコア、相同性がない場合にはマイナスのスコアを与え、
 遺伝子クラスタの中心に位置する遺伝子から端に向かって順にスコアの合計値を算出し、スコアの合計値が極大値となる遺伝子を遺伝子クラスタの境界として特定し、
 境界として特定された遺伝子により挟み込まれる領域を遺伝子クラスタとすることを特徴とする(1)記載の予測方法。
 (15)全遺伝子の数を予め定めた上記基準値を15~65個とすることを特徴とする(14)記載の予測方法。
 本明細書は本願の優先権の基礎である日本国特許出願2012-210044号の明細書及び/又は図面に記載される内容を包含する。
 本発明では、比較ゲノム科学的手法を用い、遺伝子の並びを配列と見立てて塩基配列を比較する技術を応用すること、および、単純なシンテニーと区別することにより、コア遺伝子を含むか否かに拘わらず、全く新規な二次代謝系遺伝子クラスタを予測することを可能とする。
本発明に係る二次代謝系遺伝子を含む遺伝子クラスタの予測方法に関するフローチャートである。 本発明に係る予測方法において、遺伝子クラスタを特定する際のSmith-Watermanアルゴリズムにおいて作成されるマトリックスの概念図である。 本発明に係る予測方法において、遺伝子クラスタを特定し、特定した遺伝子クラスタにオーソログチェックを行い、最終的に二次代謝系遺伝子を含む遺伝子クラスタを特定するまでの工程を示すフローチャートである。 本発明に係る予測方法を適用してオーソログチェックを行うステップを説明するための模式図である。 本発明に係る予測方法を適用してオーソログチェックを行うステップを説明するための模式図である。 本発明に係る予測方法を適用してオーソログチェックを行うステップを説明するための模式図である。 本発明に係る予測方法において、遺伝子クラスタの境界を修正するステップを説明するための模式図である。
 以下、本発明を図面を参照しながら詳細に説明する。
 本発明に係る二次代謝系遺伝子を含む遺伝子クラスタの予測方法は、少なくとも一対のゲノムに含まれる遺伝子に関して相同性検索した結果を利用し、比較したゲノムの遺伝子の並びに基づいて遺伝子クラスタを特定する工程と、特定した遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるかを判定する工程と含んでいる(図1)。
 ここで、二次代謝系遺伝子とは、二次代謝産物の生合成に関与する遺伝子を意味する。二次代謝産物とは、生物の生命活動に直接関与しない代謝産物である。なお、生物が合成する物質を代謝産物と総称する場合、代謝産物は一次代謝産物と二次代謝産物とから構成されることとなる。この場合、二次代謝産物とは、一次代謝産物を除く代謝産物ということができる。なお、一次代謝産物とは、生物の生命活動に直接関与する物質であり、例えば糖、アミノ酸、脂質及び核酸を意味する。よって、二次代謝産物とは、代謝産物のうち糖、アミノ酸、脂質及び核酸以外の物質と定義しても良い。二次代謝産物としては、例えば、抗生物質、アルカロイド、テルペノイド、フラボノイド、ポリケチド、フェノール類、配糖体、タンパク質を構成しない特殊なアミノ酸等を挙げることができる。
 また、二次代謝産物の生合成に関与する遺伝子とは、二次代謝物質の同化反応又は異化反応に関与する酵素をコードする遺伝子、二次代謝物質の転流・蓄積に関与するタンパク質をコードする遺伝子、及びこれら遺伝子の発現制御に関与するタンパク質をコードする遺伝子を含む意味である。
 より具体的に、二次代謝系遺伝子としては、ポリケチド、非リボソームペプチド、アルカロイド、テルペノイド、フラボノイド、その他の一次代謝に属さない化合物の生合成に関わる遺伝子を挙げることができる。なお、本発明に係る予測方法によって予測される遺伝子クラスタは、これら具体的に例示した二次代謝系遺伝子を含むとは限らず、他の二次代謝系遺伝子を含むこともある。
<遺伝子クラスタの特定>
 本方法においては、先ず、遺伝子クラスタを特定する。ここで、遺伝子クラスタとは、所定の連続した領域に含まれる複数の遺伝子からなる群であって、複数のゲノム間(例えば、一対のゲノム間)で並びが保存された複数の遺伝子の群を意味する。連続した領域とは、染色体、ミトコンドリアといった核酸から構成されるゲノム全体或いはゲノムの一部に含まれる領域を意味する。すなわち、遺伝子クラスタとは、ゲノム全体或いはゲノムの一部に含まれる連続した領域において、並びが保存された複数の遺伝子の群を意味する。
 遺伝子クラスタを特定するには、少なくとも一対のゲノム塩基配列情報を準備する。ゲノム塩基配列情報とは、アデニン、グアニン、シトシン及びチミンからなる4種類のヌクレオチドをそれぞれA、G、C及びGとして表記したテキストデータである。なお、ゲノム塩基配列情報は、5'末端から3'末端の方向で表記される。ここで、一対のゲノム塩基配列情報のうち一方又は両方は、各種のゲノム塩基配列情報等が格納されたデータベースから取得したデータを使用しても良いし、DNAシーケンス技術を適用して未知或いは公知の生物から取得したデータを使用しても良い。DNAシーケンス技術としては、例えばMolecular Cloning, A Laboratory Manual, Fourth Edition (Cold Spring Harbor Laboratoty Press)のChapter 11に記載される全ての方法を適用することができる。
 また、ゲノム塩基配列情報としては、如何なる生物種由来のデータを使用しても良い。言い換えると、本発明に係る予測方法では、生物種には何ら限定されず、二次代謝系遺伝子を含む遺伝子クラスタを予測することができる。具体的に、生物種としては、例えば、植物、細菌、放線菌、真菌、糸状菌、キノコ等を挙げることができる。さらに、ゲノム塩基配列情報としては、由来する生物種が明らかになっていないデータであっても良い。例えば、土壌や汚泥、湖水、海水等の環境から培養を経ずに直接抽出されたDNA、いわゆる環境DNAについて塩基配列を決定し、これをゲノム塩基配列情報として利用することもできる。すなわち、本発明に係る予測方法では、環境DNAに存在する、二次代謝系遺伝子を含む遺伝子クラスタを予測することができる。
 少なくとも一対のゲノム塩基配列情報から遺伝子クラスタを特定するには、先ず、ゲノム塩基配列情報に基づいて、一対のゲノム塩基配列情報に存在する複数の遺伝子の配置を比較し、遺伝子の並びが保存された領域を特定する。
 遺伝子の配置を比較するには、対象となる一対のゲノム塩基配列情報に含まれる遺伝子に関して相互に相同性検索し、これらゲノム塩基配列情報間で相同遺伝子の組み合わせ、相同遺伝子の組み合わせのなかでオーソログ遺伝子の組み合わせを同定する。これには先ず、対象となる一対のゲノム塩基配列情報に含まれる複数の遺伝子について、これら遺伝子によりコードされるアミノ酸配列を推定する。アミノ酸配列の推定には、いわゆるOpen Reading Frame解析ソフトウェアを使用することができる。この解析ソフトウェアによれば、5'末端から3'末端の方向で表記されたゲノム塩基配列情報及びその相補鎖について、それぞれ3つのフレームのORFを特定することができる。このとき一方のゲノム塩基配列情報にある遺伝子をxi(i=1, 2,…, I)とし、他方のゲノム塩基配列情報上にある遺伝子をyj(j=1, 2,…, J)とする。
 次に、一方のゲノム塩基配列情報に含まれる全ての遺伝子について、それぞれアミノ酸配列を問い合わせ配列(クエリー配列)とし、他方のゲノム塩基配列情報に含まれる遺伝子に関するアミノ酸配列をデータベース配列として相同性を検索する。相同性の検索には、Blastp、FASTA及びClustalといった従来公知の相同性解析ソフトウェアを使用することができる。また、問い合わせ配列としたゲノム塩基配列情報と、データベース配列としたゲノム塩基配列情報とを入れ替えて同様に相同性検索を実施する。
 この相同性検索によれば、一対のゲノム塩基配列情報の間で、配列類似性の高い遺伝子を相互に特定することができる。例えば、配列類似性を示す値に対して閾値を設定し、当該閾値を超える遺伝子の組み合わせを相同遺伝子として特定することができる。さらに、相同遺伝子として特定された遺伝子の組み合わせのなかで、所定の基準をクリアした組み合わせをオーソログ遺伝子として特定することができる。なお、オーソログ遺伝子とは、共通の起源をもつ単一の遺伝子から種の分化によってできた相同性を持つ遺伝子と定義される。
 ここで、配列類似性を示す値としては、Blast検索におけるE値(e-value)、bits値(bits)、アミノ酸一致度(Identities)といった値を挙げることができる。よって、これらのうち1又は複数の値について閾値を設定することで、相同遺伝子として遺伝子の組み合わせを特定することができる。例えば具体的には、問い合わせ配列とデータベース配列との間の相同性検索及びこれら問い合わせ配列とデータベース配列とを入れ替えて実施する相同性検索(これらをまとめて「一組の相同性検索」と称す)において、閾値としてE値(E-value)を、例えば1.0e-20、好ましくは1.0e-15、特に好ましくは1.0e-10を設定することができる。そして、一組の相同性検索において、それぞれE値の値が閾値以下となっている遺伝子の組み合わせを、両方のゲノム塩基配列情報における相同遺伝子として特定することができる。
 また、このように特定された相同遺伝子のなかからオーソログ遺伝子を特定するには、上述したオーソログ遺伝子の定義に合致する組み合わせを選択できるように基準を設定する。具体的には、上述した一組の相同性検索の結果として、配列類似性の高い順(E値の低い順)で作成される一組の遺伝子リストにおいて、互いに上位5番目以内、好ましくは上位3位以内、特に好ましくは上記1位にリストされるときにオーソログ遺伝子の組み合わせと定義することができる。また、一組の相同性検索にて特定された相同性遺伝子の組み合わせのなかからオーソログ遺伝子の組み合わせを特定するには、上述した方法に限定されるものではない。
 次に、上述した相同性検索の結果を用いて、一対のゲノム塩基配列情報における遺伝子の配置を比較し、遺伝子の並びが保存されている領域を特定する。このとき、「一対のゲノム塩基配列情報における遺伝子の配置を比較」するには、遺伝子を文字に見立てることでゲノム塩基配列情報に存在する複数の遺伝子を文字列とみなし、文字列の検索や文字列の類似性を比較するアルゴリズムを適用することができる。
 本工程で使用できるアルゴリズムとしては、文字列検索アルゴリズムであるSmith-Watermanアルゴリズム、Needleman-Wunsch アルゴリズム、k-タプル法等を挙げることができる。特に、Smith-Watermanアルゴリズムを適用することが好ましい。これは高感度に局所アライメントが得られるためである。
 Smith-Watermanアルゴリズムを適用することで、具体的には以下のようにして一対のゲノム塩基配列情報における遺伝子の配置を比較することができる。一方のゲノム塩基配列情報にある遺伝子をxi(i=1, 2,…, I)とし、他方のゲノム塩基配列情報上にある遺伝子をyj(j=1, 2,…, J)とする。Smith-Watermanアルゴリズムでは、一方のゲノム塩基配列情報の遺伝子xi(i=1, 2,…, I)と他方のゲノム塩基配列情報上にある遺伝子をyj(j=1, 2,…, J)とについて、(J+1)×(I+1)のマトリックス(二次元)を作成する(図2)。
 そして、マトリックスの各セルには、以下の手順に従って計算されるスコアが記入される。すなわち、各セルについて、xiとyjに相同性があるとき
Figure JPOXMLDOC01-appb-M000001
とし、相同性がないとき
Figure JPOXMLDOC01-appb-M000002
とする。マトリックスの全てのセルに対して以上のスコア計算を行った後、最大のスコアを持つセルからスコアが0になっているセルまでトレースバックを行う。この時に通過したセルのうち、xiとyjに相同性が高いセルの座標の集合をR0とする。ここで、gapとmismatchはペナルティスコアであり、-0.4~-0.1程度の範囲で設定されるが、好ましくは、いずれも-0.2である。
Figure JPOXMLDOC01-appb-M000003
 ここで、R0は遺伝子の並びが保存されている領域内にあり、かつ相同性の高い遺伝子ペアの座標の集合になる。すなわち、このR0が遺伝子クラスタ、すなわち、一対のゲノム塩基配列情報において並びが保存された複数の遺伝子の群となる。なお、(J+1)×(I+1)のマトリックスにおいて、最大のスコアを持つセルが複数ある場合には、上述した工程において複数の遺伝子クラスタが特定されることとなる。
 本発明に係る予測方法では、上述のように特定された遺伝子クラスタR0について、詳細を後述するように、二次代謝系遺伝子を含む遺伝子クラスタであるかを判定することができる。しかしながら、本発明に係る予測方法では、上述のように特定された遺伝子クラスタR0に限定されず、以下の手順により、更に異なる遺伝子クラスタR'0を特定し、これら遺伝子クラスタR'0について二次代謝系遺伝子を含む遺伝子クラスタであるかを判定することもできる(図3)。
 遺伝子クラスタR'0とは、上述したR0以外であって、xi(i=1, 2,…, I)とyj(j=1, 2,…, J)との対応関係で遺伝子の並びが保存されている領域として特定された遺伝子クラスタRm(m=1, 2, 3, …)を再度アライメント解析(図3中、アライメント_2と表記)した結果として導き出される遺伝子クラスタである。
 二次代謝系遺伝子を含む遺伝子クラスタは、当該クラスタを構成する各遺伝子についても多様性が高く、遺伝子クラスタ同士を比較する場合には遺伝子単位での挿入や欠失により大きなギャップが予想される。そこで、ギャップを多く含む領域も遺伝子クラスタとして検出できるように、上述した計算の際に得た(J+1)×(I+1)のマトリックス及びスコアを利用して遺伝子クラスタ(Rm、m=1, 2, 3, …)を特定する。この遺伝子クラスタ(Rm、m=1, 2, 3, …)を特定する手法としては、特に限定されないが、以下の手順を適用することができる。
 まず、ひとつ前の操作で得た集合(最初はR0)に含まれる座標が示すセルすべてに0を代入する。すなわち、
Figure JPOXMLDOC01-appb-M000004
 次に、R0の各セルに0が代入された(J+1)×(I+1)のマトリックスの中でスコアが1よりも大きく、かつ最大であるセルからスコアが0になっているセルまで再度トレースバックを行う。なお、スコアが1よりも大きく、かつ最大であるセルとは
Figure JPOXMLDOC01-appb-M000005
となる。これを条件※1と称する。
 条件※1を満たすセルからスコアが0になっているセルまでトレースバックすることで、xiとyjに相同性が高いセルの座標の集合Rmを特定することができる。なお、R0の各セルに0が代入された(J+1)×(I+1)のマトリックスにおいて、条件※1を満たすセルが複数ある場合には、上述した工程において複数の遺伝子クラスタRm(m=1, 2, 3, …)が特定されることとなる。
 また、以上のように特定された複数の遺伝子クラスタRm(m=1, 2, 3, …)については、既に特定されている遺伝子クラスタR0を十分に近接していると、R0含まれるセルのスコアに影響を受けたものとなる。よって、以上のように複数の遺伝子クラスタRm(m=1, 2, 3, …)を特定した後、R0含まれるセルのスコアに影響を排除するため、特定した遺伝子クラスタRmに対して、Smith-Watermanアルゴリズム等の文字列検索アルゴリズムによって再度、保存されている遺伝子の並びを特定し直すことが好ましい。
 具体的には、集合Rm(m=1, 2, 3, …)のうちn(Rm)≧3を満たすものについて、
Figure JPOXMLDOC01-appb-M000006
を満たす領域を取り出し、上述と同様にマトリックス(二次元)を作成するとともに再度スコアを計算する。これにより、上述のように特定したRm(m=1, 2, 3, …)から、新たに構築した遺伝子クラスタR'0を導き出すことができる。
 以上の処理を上記条件※1を満たすセルからスコアが0になっているセルまでトレースバックすることができなくなるまで繰り返すことで、二次代謝系遺伝子を含む遺伝子クラスタであるかを判定する対象となる、遺伝子クラスタ(R0、R'0、R''0…)を特定することができる。
<二次代謝系遺伝子を含む遺伝子クラスタの判定>
 以上のようにして特定された、R0で示す遺伝子クラスタ又はR0、R'0、R''0…で示す遺伝子クラスタ群について、二次代謝系遺伝子を含む遺伝子クラスタであるかを判定する(図3における「オーソログチェック」)。
 本予測方法では、二次代謝系遺伝子は多様性が高く、異なる種間でオーソロガスな遺伝子をほとんど含まないという特徴を考慮することで、二次代謝系遺伝子を含む遺伝子クラスタであるかの判定を行う。この特徴は、二次代謝系遺伝子を含む遺伝子クラスタであればシンテニー様領域の割合が低いことと同義である。したがって、特定された遺伝子クラスタについて、シンテニー様領域を特定し、当該シンテニー様領域が遺伝子クラスタに占める割合に基づいて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定することができる。
 より具体的には、特定された遺伝子クラスタにおけるシンテニー様領域は、当該遺伝子クラスタに含まれるオーソログ遺伝子の数と距離を用いて評価することができる。このとき、遺伝子クラスタの大きさや当該遺伝子クラスタに含まれる相同性遺伝子の数によって、評価対象の遺伝子クラスタを限定することが好ましい。具体的には、R0で示す遺伝子クラスタ又はR0、R'0、R''0…で示す遺伝子クラスタ群について、相同遺伝子の組み合わせが例えば2個以上、好ましくは3個以上含まれており、且つ、全体の遺伝子の数が例えば50以内、好ましくは40以内、より好ましくは35以内であるか確認する。これら両条件を満たす遺伝子クラスタについて、シンテニー様領域を特定するためのオーソログチェックを実施することが好ましい。これら条件の少なくとも一方を満たさない遺伝子クラスタについては以後の処理を行わず、二次代謝系遺伝子を含まない遺伝子クラスタとして棄却する。例えば、この段階で、相同遺伝子の組み合わせの数を3個とし、且つ、全体の遺伝子の数を35個という基準を設けた場合、この段階では下記条件(※2)で遺伝子クラスタを絞り込むこととなる。
Figure JPOXMLDOC01-appb-M000007
n:遺伝子クラスタ内での遺伝子の位置
in:ゲノム中での遺伝子の位置
 次に、上述の条件(例えば条件※2)を満たす遺伝子クラスタについてオーソログチェックを行う。オーソログチェックを実施する前に、遺伝子クラスタに含まれる遺伝子の数が上記条件における全体の遺伝子数(例えば条件※2の場合には35遺伝子)となるように遺伝子クラスタを修正する(図4)。具体的には、上記工程で特定された遺伝子クラスタに対して、例えば合計35遺伝子になるように、遺伝子クラスタの近傍に存在する遺伝子を追加する。例えば、上記工程で特定された遺伝子クラスタの両端から、同じ数の遺伝子を加えることで例えば合計35遺伝子からなる遺伝子クラスタに修正することができる。なお、拡充する遺伝子の数が奇数である場合、特に限定されないが、遺伝子クラスタの3'末端に1つ多い数又は少ない数の遺伝子を加えればよい。このように、合計の遺伝子の数を例えば35遺伝子とすることで、上記工程で特定された遺伝子クラスタの境界における誤差を考慮することができ、近傍の領域におけるオーソロガス遺伝子ペアの分布も平均化して評価することができる。なお、図4は、簡単のために遺伝子数を一部省略している。
 xi(i=1, 2,…, I)とyj(j=1, 2,…, J)について、遺伝子クラスタに含まれる遺伝子の数を35遺伝子としたときの、全遺伝子の集合をそれぞれX及びYとする。
Figure JPOXMLDOC01-appb-M000008
 そして、X及びYに含まれる遺伝子の間でオーソログ遺伝子の組み合わせが存在しているか、上述した相同性検索の結果に基づいて判断する(図4の破線矢印)。X及びYに含まれる遺伝子の間でオーソログ遺伝子の組み合わせが二つ以上ある場合には、シンテニー様領域を特定する。ここでシンテニー様領域とは、複数のオーソログ遺伝子を含む領域であって隣接するオーソログ遺伝子(間に他の遺伝子を含んでいても良い)の距離が基準値以下の領域を意味する。ここで、基準値としては、例えば10~30kbとすることができ、10~20bkとすることができ、10kbとすることができる。例えば、図5の吹き出しにおいて、Aとa、1と2の二つのペアは1~A <10kbかつ2~a<10kbを満たしているため、1~A、2~aの領域をシンテニー様領域とすることができる。X及びYに含まれるオーソログ遺伝子の組み合わせ全てについて同様にシンテニー様領域であるか確認する。なお、複数のシンテニー様領域が存在することもある。
 X及びYにおいて特定されたシンテニー様領域を、それぞれX及びYの部分集合xSB, ySBで表す。これらの部分集合の要素の数がともにX及びY全体の要素の数に対して所定の割合以下である場合に、xiからなる遺伝子クラスタ及びyjからなる遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであると判定する。ここで、所定の割合としては、特に限定されないが、30%、25%或いは20%とすることができる。例えば、所定の割合を25%としたとき(条件※3)に、
Figure JPOXMLDOC01-appb-M000009
を二次代謝系遺伝子を含む遺伝子クラスタとして予測することができる。
Figure JPOXMLDOC01-appb-M000010
 具体的に、図6においてシンテニー様領域に該当しなかった領域を破線枠で囲った。そして、シンテニー様領域(図6中、実線枠)内の遺伝子が双方ともに8遺伝子以下(全体35遺伝子に対する25%未満)であれば、最初に検出されたA~G、a~hを、それぞれ二次代謝系遺伝子を含む遺伝子クラスタとして予測する。
 二次代謝系遺伝子を含む遺伝子クラスタの予測方法としては、上記手順に従って特定したシンテニー様領域を利用する方法に限定されず、他の方法により特定されたシンテニー様領域を利用しても良い。シンテニー様領域を特定する方法としては、例えば、異なる種のゲノム塩基配列情報とアノテーション情報などを用いて、シンテニー領域と非シンテニー領域を予め決定しておく方法が挙げられる。
 予め決定したシンテニー領域を、本方法におけるシンテニー様領域として利用することで、上述した方法と同様にして二次代謝系遺伝子を含む遺伝子クラスタを予測することができる。すなわち、このようにシンテニー領域からシンテニー様領域を特定する方法であっても、図5にて説明したように、シンテニー様領域の決定と同様の方法を適用することができる。即ち、あらかじめ2種のゲノム上に予測された遺伝子についてオーソログ遺伝子を決定しておき、上述したように定義されたシンテニー領域を特定し、ゲノム塩基配列情報におけるシンテニー領域以外の領域を非シンテニー領域として定義する。そして、上記により特定された、R0で示す遺伝子クラスタ又はR0、R'0、R''0…で示す遺伝子クラスタ群について、上記と同様の方法によってクラスタ長を延長し(例えば35遺伝子)、シンテニー領域(上述した方法におけるシンテニー様領域)が全体の25%未満となる場合に、二次代謝系遺伝子を含む遺伝子クラスタとして予測することができる。
 この方法では、上述したように遺伝子クラスタを検出してからシンテニー様領域を特定する方法と比較して、より高精度に二次代謝系遺伝子を含む遺伝子クラスタを予測できる場合がある。例えば、A. flavusとA. oryzaeの様に近縁度が高い種間の比較においては、A. oryzaeの一部の株にアフラトキシン生合成クラスタと相同性の高い遺伝子クラスタが存在する場合があり、かつA. flavusあるいはA.oryzaeの他の株には同遺伝子クラスタと相同性が高い第2の遺伝子クラスタが存在しないため、A. flavusに存在するアフラトキシン生合成クラスタが検出されないことが考えられる。この場合に、実際に比較する生物種2種のゲノムの一方について、第3の種のゲノムを用いてシンテニー領域を予め決定しておくことにより、予測の可能性を向上させることができる。この方法では、シンテニー領域はAspergillus属などの比較的近縁種間で共通に存在する遺伝子の領域として定義される。
 一方、二次代謝系遺伝子を含む遺伝子クラスタの予測方法としては、上述したように、遺伝子クラスタに含まれる遺伝子の数に基づいて、評価対象の遺伝子クラスタを限定していた。しかしながら、本発明に係る予測方法では、遺伝子クラスタの長さに基づいて評価対象の遺伝子クラスタを限定してもよい。すなわち、遺伝子クラスタな長さを所定の基準値と比較し、当該基準値未満の長さの遺伝子クラスタについてオーソログチェックを実施することもできる。ここで、基準値としては、特に限定されないが、例えば125kb(50個の遺伝子程度に相当)、好ましくは100kb(40個の遺伝子程度に相当)、より好ましくは87.5kb(35個の遺伝子程度に相当)とすることができる。
 また、二次代謝系遺伝子を含む遺伝子クラスタの予測方法としては、上述したように、オーソログチェックを実施する前に、遺伝子クラスタに含まれる遺伝子の数が所定の数(例えば35個)となるように遺伝子クラスタを修正していた。しかしながら、本発明に係る予測方法では、オーソログチェックを行う前に、遺伝子クラスタに対して所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタを修正し、修正された遺伝子クラスタについてオーソログチェックを行うようにしても良い。
 遺伝子クラスタを修正する方法としては、例えば、以下に説明するように、遺伝子クラスタの境界を修正する方法が挙げられる。すなわち、特定されたR0、R'0、R''0…で示す遺伝子クラスタについて、その遺伝子クラスタの境界を修正する方法である。遺伝子クラスタの境界を修正するとは、上述した<遺伝子クラスタの特定>の欄に記載の方法で特定された遺伝子クラスタに対して、当該遺伝子クラスタの外側に存在する遺伝子を含ませるか判断することと同義である。
 具体的には、図7(a)に示すように、先ず、R0、R'0、R''0…で示す遺伝子クラスタについて、上述のように、遺伝子クラスタに含まれる遺伝子の数が15~65遺伝子、より具体的に好ましくは35遺伝子(35遺伝子には限定されない)となるように遺伝子クラスタを拡張する。次に、拡張した遺伝子クラスタを構成する各遺伝子について、比較する遺伝子クラスタ内に相同性が高い遺伝子が存在する場合にはプラスのスコアを与え、相同性が高い遺伝子が存在しない場合にはマイナスのスコアを与える。次に、図7(b)に示すように、拡張した遺伝子クラスタの中心に位置する遺伝子から順に両端に向かって遺伝子に与えられたスコアを合計し、各遺伝子にスコアの合計値を与える。そして、拡張した遺伝子クラスタに含まれる各遺伝子に与えられたスコアの合計値が極大値となる遺伝子を特定し、特定した遺伝子を遺伝子クラスタの境界とする。なお、この処理により、遺伝子クラスタの境界となる遺伝子が修正されずオリジナルの遺伝子クラスタのままとなる場合もある。
 より具体的に、例えばxi(i=1, 2,…, I)とyj(j=1, 2,…, J)について、遺伝子クラスタに含まれる遺伝子の数を35遺伝子に拡張したときの、全遺伝子の集合をそれぞれX及びYとする。
Figure JPOXMLDOC01-appb-M000011
 そして、遺伝子クラスタの境界を修正するため、n(X)個の要素からなる一次元配列SCを作製した。この配列の各要素には、例えば以下の式に従って計算されたスコアを入れることができる。xiがyc, yc-1, … yd-1, ydの少なくとも一つに相同性があるとき
Figure JPOXMLDOC01-appb-M000012
xiがyc, yc-1, … yd-1, ydのいずれとも相同性がないとき
Figure JPOXMLDOC01-appb-M000013
とした。そして、配列の全ての要素に対して以上のスコア計算を行った後、(1)及び(2)それぞれの範囲で最大のスコアを持つ要素の要素番号をistart及びistopとする。同様の操作を集合Yに対しても行う。
 このようにして特定されたistart及びistopを遺伝子クラスタの境界とする。すなわち、境界が修正された遺伝子クラスタは、それぞれ、
Figure JPOXMLDOC01-appb-M000014
となる。なお、xiがyc, yc-1, … yd-1, ydのいずれとも相同性がないときのスコアSC(j)において、negativeの値は例えば-0.1、-0.2、-0.3、-0.4、-0.5或いは-1とすることができる。
 以上のようにして、R0、R'0、R''0…で示す遺伝子クラスタの境界を修正することで、上述したオーソログチェックを経て予測される二次代謝系遺伝子の遺伝子クラスタの正確性を向上させることができる。なお、上述のように、遺伝子クラスタの境界を修正する処理は、上述したオーソログチェックの前でも良いし、オーソログチェックの後でも良い。
<予測装置及び予測プログラム>
 以上で説明した本発明に係る二次代謝系遺伝子を含む遺伝子クラスタの予測方法は、マウスやキーボード等の入力手段、中央演算処理手段(CPU)、揮発性及び/又は不揮発性メモリを含む記憶手段及びディスプレイ等の出力手段を備えるコンピュータにて実行することができる。このとき、コンピュータは、外部のデータベース等の記憶装置や外部のコンピュータシステム等に対してインターネットやイントラネット等の通信回線網を介して接続されていることが好ましい。すなわち、本発明に係る予測方法は、上記構成のコンピュータ装置を用いて二次代謝系遺伝子を含む遺伝子クラスタを予測することができる予測プログラムとして提供することができる。また、この予測プログラムをインストールしたコンピュータは、言い換えると、二次代謝系遺伝子を含む遺伝子クラスタの予測装置となる。
 上述した予測方法をコンピュータにて実現するには、一対のゲノム塩基配列情報を外部の記憶手段やコンピュータシステムから上記通信回線網を介してコンピュータに入力しても良いし、DNAシーケンサーとインターフェイスを介して接続されてコンピュータに入力されても良い。或いは、DVDやCD等の記憶媒体を利用して一対のゲノム塩基配列情報をコンピュータに読み込んでも良い。
 また、コンピュータによれば、中央演算装置により一対のゲノム塩基配列情報について相同性検索を実行することができ、相同性検索の結果を記憶装置に格納することができる。また、コンピュータによれば、上述した<遺伝子クラスタの特定>手順及び<二次代謝系遺伝子を含む遺伝子クラスタの判定>については、Smith-Watermanアルゴリズム等の文字列検索アルゴリズムを実装したソフトウェアにより実現することができる。
 以下、実施例を用いて本発明をより詳細に説明するが、本発明の技術的範囲は以下の実施例に限定されるものではない。
〔実施例1〕
 本実施例では、8種類のゲノムのデータセットを使用した。このうち、Aspergillus oryzaeのデータはGenBankに登録したもの(AP007150-AP007177)と同等のものを使用した。Aspergillus flavusのデータはGenBankよりgenbank形式のファイルをダウンロードして使用した。(genbank accession EQ963472~EQ963493)。Aspergillus fumigatas、Aspergillus nidulans、Aspergillus terreus、Magnaporthe grisea、Fusarium graminearum及びChaetomium globosumのデータはBROAD INSTITUTE よりダウンロードして使用した。
 本実施例では、相同性検索においてE値(E-value)が1.0e-10以下のものを相同遺伝子とした。また、本実施例では、相同性検索の結果として配列類似性の高い順(E値の低い順)で作成される一組の遺伝子リストにおいて互いに1位にリストされるときにオーソログ遺伝子の組み合わせとした。
 また、本実施例では、Smith-Watermanアルゴリズムにより遺伝子の並びの保存性を検証し、遺伝子クラスタR0、R'0、R''0…を特定した。また、本実施例では、シンテニー様領域を特定するため、特定した遺伝子クラスタに含まれる相同遺伝子の組み合わせの数を3個以上とし、且つ、全体の遺伝子の数を35個未満という基準を設けた。また、本実施例において、シンテニー様領域とは、複数のオーソログ遺伝子を含む領域であって隣接するオーソログ遺伝子(間に他の遺伝子を含んでいても良い)の距離が10bk以下、20kb以下或いは30kb以下の領域として定義した。
 さらに、本実施例では、シンテニー様領域(X及びYの部分集合xSB,及びySB)に含まれる遺伝子の数(部分集合の要素の数)が35遺伝子中の25%未満(8個以下)のときに、もとの遺伝子クラスタを二次代謝系遺伝子を含む遺伝子クラスタとして予測した。
 上記に記載の方法を用いて、A. flavus、A. oryzaeなどのゲノム解析が完了した糸状菌の10個のゲノム塩基配列を用いて予測された二次代謝遺伝子を含む遺伝子クラスタの数を表1に示した。なお、表1-1は、シンテニー様領域を隣接するオーソログ遺伝子の距離が10bk以下と定義した場合の結果であり、表1-2は同距離を20kb以下とした場合の結果であり、表1-3は同距離を30kb以下とした場合の結果を示している。この結果より、シンテニー様領域を隣接するオーソログ遺伝子の距離が10~30bk以下と定義すれば結果について特に大きな変動がないことが明らかとなった。
Figure JPOXMLDOC01-appb-T000015
Figure JPOXMLDOC01-appb-T000016
Figure JPOXMLDOC01-appb-T000017
 また、本実施例で二次代謝系遺伝子を含む遺伝子クラスタとして予測されたもののなかでQ遺伝子が含まれているものの割合を計算した結果を表2に示す。なお、Q遺伝子とは、COG(Cluster of Orthologous Group)の機能分類において、二次代謝系と考えられる遺伝子分類である。
Figure JPOXMLDOC01-appb-T000018
Figure JPOXMLDOC01-appb-T000019
Figure JPOXMLDOC01-appb-T000020
 表2に示す結果から、本実施例によって、二次代謝系遺伝子を含む遺伝子クラスタと予測されたものは、非常に高い確率でQ遺伝子を含んでいることが明らかとなった。この結果から、本実施例に示した方法によれば、二次代謝系遺伝子を含む遺伝子クラスタを高精度に予測することができ、従来の方法論では見つけられなかった二次代謝系遺伝子を含む遺伝子クラスタを見つけられる可能性が非常に高いことが判った。
〔実施例2〕
 本実施例では、実施例1と同様にして、Smith-Watermanアルゴリズムにより遺伝子の並びの保存性を検証し、遺伝子クラスタR0、R'0、R''0…を特定した。また、本実施例では、特定した遺伝子クラスタの境界を修正する工程として、35遺伝子となるように拡張した遺伝子クラスタに含まれる各遺伝子について、相同遺伝子がある場合には+1を与え、相同遺伝子がない場合には-0.3を与え、拡張した遺伝子クラスタの中心からスコアの合計値を算出し、合計値の極大値を取る遺伝子を遺伝子クラスタの境界とした以外は、実施例1と同様にして、二次代謝系遺伝子を含む遺伝子クラスタを予測した。
 本実施例で予測した二次代謝系遺伝子を含む遺伝子クラスタの一部を表3に示す。また、実施例1と同様に、遺伝子クラスタの境界を修正することなく予測した二次代謝系遺伝子を含む遺伝子クラスタを表4に示す。
Figure JPOXMLDOC01-appb-T000021
Figure JPOXMLDOC01-appb-T000022
 なお、表3及び4において、エラーの欄は、実際の二次代謝系遺伝子の遺伝子クラスタに対して、予測した遺伝子クラスタが上流方向(5’側)及び下流方向(3’側)にどの程度誤っているかを遺伝子の個数として示している。
 表4から判るように、遺伝子クラスタの境界を修正しない場合には、エラーとして94個の遺伝子がカウントされた。これは、表4に示した21個の遺伝子クラスタに平均4.5個のエラーが含まれることになる。これに対して、遺伝子クラスタの境界を修正した場合には、エラーとして82個の遺伝子がカウントされ、21個の遺伝子クラスタに平均3.9個のエラーとなった。このように遺伝子クラスタの境界を修正することによって、二次代謝系遺伝子を含む遺伝子クラスタをより正確に検出することができる。
 本明細書で引用した全ての刊行物、特許および特許出願をそのまま参考として本明細書にとり入れるものとする。

Claims (45)

  1.  少なくとも一対のゲノム塩基配列情報に含まれる遺伝子に関して相互に相同性検索し、これらゲノム塩基配列情報間で相同遺伝子の組み合わせ、相同遺伝子の組み合わせのなかでオーソログ遺伝子の組み合わせを同定する工程と、
     上記相同性検索の結果に基づいて他のゲノム塩基配列情報との間で遺伝子の並びが保存されている領域を遺伝子クラスタとして特定する工程と、
     上記工程で特定された遺伝子クラスタのうち、上記相同性検索の結果のうちオーソログ遺伝子の存在に基づいてシンテニー様領域を特定し、当該シンテニー様領域が占める割合に基づいて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程と
     を含む二次代謝系遺伝子を含む遺伝子クラスタの予測方法。
  2.  上記シンテニー様領域に含まれる遺伝子数が上記遺伝子クラスタ全体に含まれる遺伝子数に対して、所定の割合以下である場合に当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであると判定することを特徴とする請求項1記載の予測方法。
  3.  上記所定の割合を25%とすることを特徴とする請求項2記載の予測方法。
  4.  上記シンテニー様領域は、少なくとも2以上のオーソログ遺伝子を含み、隣接するオーソログ遺伝子の距離が上記ゲノム塩基配列情報及び上記他のゲノム塩基配列情報のそれぞれにおいて所定の距離以内にある領域とすることを特徴とする請求項1記載の予測方法。
  5.  上記所定の距離は、10~30kbとすることを特徴とする請求項4記載の予測方法。
  6.  上記シンテニー様領域は、比較に用いる少なくとも一対のゲノム塩基配列情報のうちの1種のゲノム塩基配列情報と、これら一対のゲノム塩基配列情報とは異なる第3のゲノム塩基配列情報とを用いてシンテニー領域/非シンテニー領域をあらかじめ決定し、このシンテニー領域をシンテニー様領域とすることを特徴とする請求項1記載の予測方法。
  7.  上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる相同遺伝子の数及び/又は特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較し、相同遺伝子の数が基準値以上である遺伝子クラスタ、及び/又は全遺伝子の数が基準値未満の遺伝子クラスタについて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記工程を行うことを特徴とする請求項1記載の予測方法。
  8.  相同遺伝子の数に関する上記基準値を3個とし、全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項7記載の予測方法。
  9.  上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに含まれる遺伝子の数が上記基準値の個数となるように、当該遺伝子クラスタに隣接する遺伝子を追加して当該遺伝子クラスタを修正し、当該基準値の個数の遺伝子からなる修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項1記載の予測方法。
  10.  全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項9記載の予測方法。
  11.  上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに対して、所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタを修正し、修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項1記載の予測方法。
  12.  上記遺伝子クラスタを特定する工程では、Smith-Watermanアルゴリズムで作成されるSmith-Watermanマトリックスにおいて、最大のスコアを示すセルからトレースバックすることで遺伝子クラスタを特定することを特徴とする請求項1記載の予測方法。
  13.  上記遺伝子クラスタを特定する工程では、上記特定された遺伝子クラスタに含まれるセルのスコアに0を代入し、このSmith-Watermanマトリックスについてトレースバックすることで他の遺伝子の並びが保存されている領域を特定し、特定された領域について再度Smith-Watermanアルゴリズムにて遺伝子の並びが保存された領域を特定し、これを遺伝子クラスタとして特定することを特徴とする請求項12記載の予測方法。
  14.  上記遺伝子クラスタを特定する工程の後、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、遺伝子クラスタに対して所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタが当該基準値となるように拡張し、
     拡張した遺伝子クラスタを構成する各遺伝子について、他のゲノム塩基配列情報における比較対象の遺伝子クラスタを構成する遺伝子との間で相同性がある場合にはプラスのスコア、相同性がない場合にはマイナスのスコアを与え、
     遺伝子クラスタの中心に位置する遺伝子から端に向かって順にスコアの合計値を算出し、スコアの合計値が極大値となる遺伝子を遺伝子クラスタの境界として特定し、
     境界として特定された遺伝子により挟み込まれる領域を遺伝子クラスタとすることを特徴とする請求項1記載の予測方法。
  15.  全遺伝子の数を予め定めた上記基準値を15~65個とすることを特徴とする請求項14記載の予測方法。
  16.  入力手段、中央演算処理手段及び記憶手段を備えるコンピュータに対して、
     上記中央演算処理手段が少なくとも一対のゲノム塩基配列情報に含まれる遺伝子に関して相互に相同性検索し、これらゲノム塩基配列情報間で相同遺伝子の組み合わせ、相同遺伝子の組み合わせのなかでオーソログ遺伝子の組み合わせを同定する工程と、
     上記中央演算処理手段が上記相同性検索の結果に基づいて他のゲノム塩基配列情報との間で遺伝子の並びが保存されている領域を遺伝子クラスタとして特定する工程と、
     上記中央演算処理手段が上記工程で特定された遺伝子クラスタのうち、上記相同性検索の結果のうちオーソログ遺伝子の存在に基づいてシンテニー様領域を特定し、当該シンテニー様領域が占める割合に基づいて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程と
     を実行させる、二次代謝系遺伝子を含む遺伝子クラスタの予測プログラム。
  17.  上記中央演算処理手段は、上記シンテニー様領域に含まれる遺伝子数が上記遺伝子クラスタ全体に含まれる遺伝子数に対して、所定の割合以下である場合に当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであると判定することを特徴とする請求項16記載の予測プログラム。
  18.  上記所定の割合を25%とすることを特徴とする請求項17記載の予測プログラム。
  19.  上記シンテニー様領域は、少なくとも2以上のオーソログ遺伝子を含み、隣接するオーソログ遺伝子の距離が上記ゲノム塩基配列情報及び上記他のゲノム塩基配列情報のそれぞれにおいて所定の距離以内にある領域とすることを特徴とする請求項16記載の予測プログラム。
  20.  上記所定の距離は、10~30kbとすることを特徴とする請求項19記載の予測プログラム。
  21.  上記シンテニー様領域は、比較に用いる少なくとも一対のゲノム塩基配列情報のうちの1種のゲノム塩基配列情報と、これら一対のゲノム塩基配列情報とは異なる第3のゲノム塩基配列情報とを用いてシンテニー領域/非シンテニー領域をあらかじめ決定し、このシンテニー領域をシンテニー様領域とすることを特徴とする請求項16記載の予測プログラム。
  22.  上記遺伝子クラスタを特定する工程の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる相同遺伝子の数及び/又は特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較し、相同遺伝子の数が基準値以上である遺伝子クラスタ、及び/又は全遺伝子の数が基準値未満の遺伝子クラスタについて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記工程を行うことを特徴とする請求項16記載の予測プログラム。
  23.  相同遺伝子の数に関する上記基準値を3個とし、全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項22記載の予測プログラム。
  24.  上記遺伝子クラスタを特定する工程の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに含まれる遺伝子の数が上記基準値の個数となるように、当該遺伝子クラスタに隣接する遺伝子を追加して当該遺伝子クラスタを修正し、当該基準値の個数の遺伝子からなる修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項16記載の予測プログラム。
  25.  全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項24記載の予測プログラム。
  26.  上記遺伝子クラスタを特定する工程の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する工程では、判定対象の遺伝子クラスタに対して、所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタを修正し、修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項16記載の予測プログラム。
  27.  上記遺伝子クラスタを特定する工程では、Smith-Watermanアルゴリズムで作成されるSmith-Watermanマトリックスにおいて、最大のスコアを示すセルからトレースバックすることで遺伝子クラスタを特定することを特徴とする請求項16記載の予測プログラム。
  28.  上記遺伝子クラスタを特定する工程では、上記特定された遺伝子クラスタに含まれるセルのスコアに0を代入し、このSmith-Watermanマトリックスについてトレースバックすることで他の遺伝子の並びが保存されている領域を特定し、特定された領域について再度Smith-Watermanアルゴリズムにて遺伝子の並びが保存された領域を特定し、これを遺伝子クラスタとして特定することを特徴とする請求項27記載の予測プログラム。
  29.  上記遺伝子クラスタを特定する工程の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、遺伝子クラスタに対して所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタが当該基準値となるように拡張し、
     拡張した遺伝子クラスタを構成する各遺伝子について、他のゲノム塩基配列情報における比較対象の遺伝子クラスタを構成する遺伝子との間で相同性がある場合にはプラスのスコア、相同性がない場合にはマイナスのスコアを与え、
     遺伝子クラスタの中心に位置する遺伝子から端に向かって順にスコアの合計値を算出し、スコアの合計値が極大値となる遺伝子を遺伝子クラスタの境界として特定し、
     境界として特定された遺伝子により挟み込まれる領域を遺伝子クラスタとすることを特徴とする請求項16記載の予測プログラム。
  30.  全遺伝子の数を予め定めた上記基準値を15~65個とすることを特徴とする請求項29記載の予測プログラム。
  31.  入力手段、中央演算処理手段及び記憶手段を備え、
     上記中央演算処理手段が少なくとも一対のゲノム塩基配列情報に含まれる遺伝子に関して相互に相同性検索し、これらゲノム塩基配列情報間で相同遺伝子の組み合わせ、相同遺伝子の組み合わせのなかでオーソログ遺伝子の組み合わせを同定する相同性検索手段と、
     上記中央演算処理手段が上記相同性検索手段の結果に基づいて他のゲノム塩基配列情報との間で遺伝子の並びが保存されている領域を遺伝子クラスタとして特定する遺伝子クラスタ特定手段と、
     上記中央演算処理手段が上記遺伝子クラスタ特定手段で特定された遺伝子クラスタのうち、上記相同性検索手段の結果のうちオーソログ遺伝子の存在に基づいてシンテニー様領域を特定し、当該シンテニー様領域が占める割合に基づいて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する判定手段と
     から構成される、二次代謝系遺伝子を含む遺伝子クラスタの予測装置。
  32.  上記中央演算処理手段は、上記シンテニー様領域に含まれる遺伝子数が上記遺伝子クラスタ全体に含まれる遺伝子数に対して、所定の割合以下である場合に当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであると判定することを特徴とする請求項31記載の予測装置。
  33.  上記所定の割合を25%とすることを特徴とする請求項32記載の予測装置。
  34.  上記シンテニー様領域は、少なくとも2以上のオーソログ遺伝子を含み、隣接するオーソログ遺伝子の距離が上記ゲノム塩基配列情報及び上記他のゲノム塩基配列情報のそれぞれにおいて所定の距離以内にある領域とすることを特徴とする請求項31記載の予測装置。
  35.  上記所定の距離は、10~30kbとすることを特徴とする請求項34記載の予測装置。
  36.  上記シンテニー様領域は、比較に用いる少なくとも一対のゲノム塩基配列情報のうちの1種のゲノム塩基配列情報と、これら一対のゲノム塩基配列情報とは異なる第3のゲノム塩基配列情報とを用いてシンテニー領域/非シンテニー領域をあらかじめ決定し、このシンテニー領域をシンテニー様領域とすることを特徴とする請求項31記載の予測装置。
  37.  上記遺伝子クラスタ特定手段による処理の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる相同遺伝子の数及び/又は特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較し、相同遺伝子の数が基準値以上である遺伝子クラスタ、及び/又は全遺伝子の数が基準値未満の遺伝子クラスタについて、当該遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記判定手段の処理を行うことを特徴とする請求項31記載の予測装置。
  38.  相同遺伝子の数に関する上記基準値を3個とし、全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項37記載の予測装置。
  39.  上記遺伝子クラスタ特定手段による処理の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記判定手段の処理を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記判定手段は、判定対象の遺伝子クラスタに含まれる遺伝子の数が上記基準値の個数となるように、当該遺伝子クラスタに隣接する遺伝子を追加して当該遺伝子クラスタを修正し、当該基準値の個数の遺伝子からなる修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項31記載の予測装置。
  40.  全遺伝子の数に関する上記基準値を35個とすることを特徴とする請求項39記載の予測装置。
  41.  上記遺伝子クラスタ特定手段による処理の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、全遺伝子の数又は長さが基準値未満の遺伝子クラスタについて、遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記判定手段の処理を行い、
     遺伝子クラスタが二次代謝系遺伝子を含む遺伝子クラスタであるか判定する上記判定手段では、判定対象の遺伝子クラスタに対して、所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタを修正し、修正された遺伝子クラスタについてシンテニー様領域を特定することを特徴とする請求項31記載の予測装置。
  42.  上記遺伝子クラスタ特定手段では、Smith-Watermanアルゴリズムで作成されるSmith-Watermanマトリックスにおいて、最大のスコアを示すセルからトレースバックすることで遺伝子クラスタを特定することを特徴とする請求項31記載の予測装置。
  43.  上記遺伝子クラスタ特定手段では、上記特定された遺伝子クラスタに含まれるセルのスコアに0を代入し、このSmith-Watermanマトリックスについてトレースバックすることで他の遺伝子の並びが保存されている領域を特定し、特定された領域について再度Smith-Watermanアルゴリズムにて遺伝子の並びが保存された領域を特定し、これを遺伝子クラスタとして特定することを特徴とする請求項42記載の予測装置。
  44.  上記遺伝子クラスタを特定する工程の後、上記中央演算処理手段は、特定した遺伝子クラスタに含まれる全遺伝子の数を予め定めた基準値を比較するか、特定した遺伝子クラスタの長さを予め定めた基準値と比較し、遺伝子クラスタに対して所定の数の遺伝子又は所定の長さの領域を追加して当該遺伝子クラスタが当該基準値となるように拡張し、
     拡張した遺伝子クラスタを構成する各遺伝子について、他のゲノム塩基配列情報における比較対象の遺伝子クラスタを構成する遺伝子との間で相同性がある場合にはプラスのスコア、相同性がない場合にはマイナスのスコアを与え、
     遺伝子クラスタの中心に位置する遺伝子から端に向かって順にスコアの合計値を算出し、スコアの合計値が極大値となる遺伝子を遺伝子クラスタの境界として特定し、
     境界として特定された遺伝子により挟み込まれる領域を遺伝子クラスタとすることを特徴とる請求項31記載の予測装置。
  45.  全遺伝子の数を予め定めた上記基準値を15~65個とすることを特徴とする請求項44記載の予測装置。
PCT/JP2013/075702 2012-09-24 2013-09-24 二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置 Ceased WO2014046284A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2014536957A JP5946149B2 (ja) 2012-09-24 2013-09-24 二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置
US14/427,349 US20150310168A1 (en) 2012-09-24 2013-09-24 Method for predicting gene cluster including secondary metabolism-related genes, prediction program, and prediction device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2012-210044 2012-09-24
JP2012210044 2012-09-24

Publications (1)

Publication Number Publication Date
WO2014046284A1 true WO2014046284A1 (ja) 2014-03-27

Family

ID=50341583

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2013/075702 Ceased WO2014046284A1 (ja) 2012-09-24 2013-09-24 二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置

Country Status (3)

Country Link
US (1) US20150310168A1 (ja)
JP (1) JP5946149B2 (ja)
WO (1) WO2014046284A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118522347A (zh) * 2024-07-23 2024-08-20 江西师范大学 基于同源性分析和区域加权的线粒体基因组重排量化方法

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10118945B2 (en) * 2015-10-30 2018-11-06 University Of Kansas Dereplication strain of Aspergillus nidulans
US10612032B2 (en) 2016-03-24 2020-04-07 The Board Of Trustees Of The Leland Stanford Junior University Inducible production-phase promoters for coordinated heterologous expression in yeast
CA3042726A1 (en) 2016-11-16 2018-05-24 The Board Of Trustees Of The Leland Stanford Junior University Systems and methods for identifying and expressing gene clusters
WO2019055816A1 (en) * 2017-09-14 2019-03-21 Lifemine Therapeutics, Inc. HUMAN THERAPEUTIC TARGETS AND MODULATORS THEREFOR
CN111508561B (zh) * 2019-07-04 2024-02-06 北京希望组生物科技有限公司 同源序列和同源序列中串联重复序列的检测方法、计算机可读介质和应用
WO2021158989A1 (en) * 2020-02-07 2021-08-12 Lodo Therapeutics Corporation Methods and apparatus for efficient and accurate assembly of long-read genomic sequences
WO2023091950A1 (en) * 2021-11-16 2023-05-25 Lifemine Therapeutics, Inc. Methods and systems for discovery of non-embedded target genes

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005514959A (ja) * 2002-01-24 2005-05-26 エコピア バイオサイエンシーズ インク 微生物から二次代謝産物を同定するための方法、システム及び情報リポジトリ
WO2012039484A1 (ja) * 2010-09-22 2012-03-29 独立行政法人産業技術総合研究所 遺伝子クラスタ及び遺伝子の探索、同定法およびそのための装置

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005514959A (ja) * 2002-01-24 2005-05-26 エコピア バイオサイエンシーズ インク 微生物から二次代謝産物を同定するための方法、システム及び情報リポジトリ
WO2012039484A1 (ja) * 2010-09-22 2012-03-29 独立行政法人産業技術総合研究所 遺伝子クラスタ及び遺伝子の探索、同定法およびそのための装置

Non-Patent Citations (6)

* Cited by examiner, † Cited by third party
Title
D'ALENCON, E.: "Extensive synteny conservation of holocentric chromosomes in Lepidoptera despite high rates of local genome rearrangements", PROCEEDINGS OF THE NATIONAL ACADEMY OF SCIENCES OF THE UNITED STATES OF AMERICA, vol. 107, no. 17, 27 April 2010 (2010-04-27), pages 7680 - 7685 *
HARUO IKEDA: "New System for the Production of Secondary Metabolites by a Versatile Host Genetically Optimized", KAGAKU TO SEIBUTSU, vol. 44, no. 6, 2006, pages 391 - 398 *
ITARU TAKEDA: "4Fp18 Prediction of secondary metabolite gene cluster using comparative genomics", ABSTRACTS OF THE ANNUAL MEETING OF THE SOCIETY FOR BIOTECHNOLOGY, 25 September 2012 (2012-09-25), JAPAN, pages 217 *
ITARU TAKEDA: "Hikaku Genome o Mochiita Niji Taishakei Idenshi Cluster no Yosoku", SOCIETY OF GENOME MICROBIOLOGY, 10 March 2012 (2012-03-10), pages 71 *
MEDEMA, M.H.: "antiSMASH: rapid identification, annotation and analysis of secondary metabolite biosynthesis gene clusters in bacterial and fungal genome sequences", NUCLEIC ACIDS RESEARCH, vol. 39, July 2011 (2011-07-01), pages W339 - 346 *
WEBER, T.: "CLUSEAN: a computer-based framework for the automated analysis of bacterial secondary metabolite biosynthetic gene clusters", JOURNAL OF BIOTECHNOLOGY, vol. 140, 10 March 2009 (2009-03-10), pages 13 - 17 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118522347A (zh) * 2024-07-23 2024-08-20 江西师范大学 基于同源性分析和区域加权的线粒体基因组重排量化方法

Also Published As

Publication number Publication date
JPWO2014046284A1 (ja) 2016-08-18
JP5946149B2 (ja) 2016-07-05
US20150310168A1 (en) 2015-10-29

Similar Documents

Publication Publication Date Title
JP5946149B2 (ja) 二次代謝系遺伝子を含む遺伝子クラスタの予測方法、予測プログラム及び予測装置
Liu et al. SMARTdenovo: a de novo assembler using long noisy reads
Zhu et al. Phylogenomics of 10,575 genomes reveals evolutionary proximity between domains Bacteria and Archaea
Angly et al. The GAAS metagenomic tool and its estimations of viral and microbial average genome size in four major biomes
Wheeler A resource for improved predictions of Trypanosoma and Leishmania protein three-dimensional structure
Auch et al. Genome BLAST distance phylogenies inferred from whole plastid and whole mitochondrion genome sequences
Sharma et al. Gene loss rather than gene gain is associated with a host jump from monocots to dicots in the smut fungus Melanopsichium pennsylvanicum
Spang et al. Complex archaea that bridge the gap between prokaryotes and eukaryotes
Stobbe et al. E-probe Diagnostic Nucleic acid Analysis (EDNA): a theoretical approach for handling of next generation sequencing data for diagnostics
Galardini et al. Evolution of intra-specific regulatory networks in a multipartite bacterial genome
Camiolo et al. The relation of codon bias to tissue-specific gene expression in Arabidopsis thaliana
Cerón-Romero et al. PhyloToL: a taxon/gene-rich phylogenomic pipeline to explore genome evolution of diverse eukaryotes
Pi et al. A genomics based discovery of secondary metabolite biosynthetic gene clusters in Aspergillus ustus
Fourie et al. Evidence for inter-specific recombination among the mitochondrial genomes of Fusarium species in the Gibberella fujikuroi complex
Zhang et al. Genome wide analysis of the transition to pathogenic lifestyles in Magnaporthales fungi
Panthee et al. Utilization of hybrid assembly approach to determine the genome of an opportunistic pathogenic fungus, Candida albicans TIMM 1768
Cope et al. Gene expression of functionally-related genes coevolves across fungal species: detecting coevolution of gene expression using phylogenetic comparative methods
WO2021158989A1 (en) Methods and apparatus for efficient and accurate assembly of long-read genomic sequences
Umemura et al. Fine de novo sequencing of a fungal genome using only SOLiD short read data: verification on Aspergillus oryzae RIB40
Renzi et al. Yeast metagenomics: analytical challenges in the analysis of the eukaryotic microbiome
O’Meara et al. DeORFanizing Candida albicans genes using coexpression
Li et al. A nearest neighbor approach for automated transporter prediction and categorization from protein sequences
O’Meara et al. CryptoCEN: A Co-Expression Network for Cryptococcus neoformans reveals novel proteins involved in DNA damage repair
CN114245922A (zh) 单一生物单元的序列信息的新型处理方法
Atias et al. Large-scale analysis of Arabidopsis transcription reveals a basal co-regulation network

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13838136

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2014536957

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 14427349

Country of ref document: US

122 Ep: pct application non-entry in european phase

Ref document number: 13838136

Country of ref document: EP

Kind code of ref document: A1