WO2017169736A1 - 情報処理装置及びプログラム - Google Patents
情報処理装置及びプログラム Download PDFInfo
- Publication number
- WO2017169736A1 WO2017169736A1 PCT/JP2017/010169 JP2017010169W WO2017169736A1 WO 2017169736 A1 WO2017169736 A1 WO 2017169736A1 JP 2017010169 W JP2017010169 W JP 2017010169W WO 2017169736 A1 WO2017169736 A1 WO 2017169736A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- individual
- data
- generation
- codon
- mutation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12M—APPARATUS FOR ENZYMOLOGY OR MICROBIOLOGY; APPARATUS FOR CULTURING MICROORGANISMS FOR PRODUCING BIOMASS, FOR GROWING CELLS OR FOR OBTAINING FERMENTATION OR METABOLIC PRODUCTS, i.e. BIOREACTORS OR FERMENTERS
- C12M1/00—Apparatus for enzymology or microbiology
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
Definitions
- the present invention relates to an information processing apparatus and program capable of increasing the production amount of a target protein.
- a technique for introducing a plurality of genes encoding a target protein when a target protein is produced by a microorganism or the like is known. Such genes are often used having the same DNA sequence. However, when a plurality of genes having the same DNA sequence are introduced, homologous recombination occurs between these genes and a part of the gene is lost.
- homologous recombination refers to recombination that occurs at a site (homologous site) whose DNA base sequence is very similar. This is conceptually represented in FIG. FIG. 1 (a) shows an example in which five genes having the same DNA sequence are introduced. Among these five genes, when homologous recombination occurs in the second half of the second gene to the first half of the fifth gene, as shown in FIG. As a result, the production efficiency of the target protein decreases.
- Patent Document 1 discloses a method for obtaining a synthetic nucleic acid molecule, which comprises (i) providing an amino acid sequence derived from an amino acid repeat region of a polypeptide; (ii) a plurality of samples each encoding the amino acid sequence Inferring a codon optimized nucleic acid sequence; (iii) aligning the plurality of sample codon optimized nucleic acid sequences by sequence homology and constructing a neighborhood binding tree comprising the plurality of sample codon optimized nucleic acid sequences; (Iv) selecting only one of the plurality of sample codon optimized nucleic acid sequences; and (v) obtaining a nucleic acid molecule comprising the selected sample codon optimized nucleic acid sequence. Has been.
- the present invention provides an information processing apparatus and program capable of suppressing homologous recombination and increasing the production amount of a target protein.
- a child individual population data acquisition unit for acquiring first generation child individual population data representing a first generation child individual population, and a predetermined evaluation criterion, the evaluation criterion regarding the codon fitness and the base sequence of the codon
- non-dominant sort processing is executed on the first generation integrated data, and the first generation integrated data is included in the first generation integrated data.
- a non-dominant sort execution unit that classifies all individual data for each rank in the Pareto optimal solution, and selects a predetermined number of the individual data in descending order of rank from all the individual data classified for each rank An information processing apparatus having an individual selection unit is provided.
- the present invention it is possible to suppress homologous recombination and increase the production amount of the target protein based on two different evaluation criteria.
- the individual selection unit selects, in order from the largest congestion distance, if the individual data having the same rank exists.
- the parent individual population data acquisition unit sets the individual data selected by the individual selection unit as second generation parent individual population data representing a second generation parent individual population, and the mutation processing unit, The processes by the superior sort execution unit and the individual selection unit are executed until a predetermined number of generations are reached.
- the evaluation standard regarding the codon matching degree is based on the minimum value of the codon matching index of CDS representing a base sequence that each individual has a plurality of base sequences to be subjected to amino acid translation.
- the higher the minimum value of the codon matching index included in the individual the higher the evaluation of the individual.
- the evaluation standard regarding the base sequence of the codon is based on the minimum value of the number of mismatched bases indicating the number of bases that do not match each other among the two CDSs included in each individual.
- the higher the minimum value of the number of mismatched bases the higher the evaluation of the individual.
- the evaluation standard regarding the base sequence of the codon is that the longest base sequence among the base sequences that continuously match at different sites between the respective CDSs or within one CDS among the CDS included in each individual. Is based on the length of the longest common character string. Preferably, the shorter the length of the longest common character string, the higher the individual is evaluated.
- the mutation processing unit has a second mutation different from the first mutation process and the first mutation process for each individual data included in the g generation parent individual population data representing the g generation parent individual population. Execute the process.
- the mutation processing unit replaces the codon included in the CDS with a codon having a higher frequency than the codon with a predetermined probability for all CDS included in the individual. Execute.
- the mutation processing unit is the longest common character that is the longest base sequence among the base sequences that continuously match at different sites between the respective CDSs or within one CDS among the CDS included in each individual.
- a second mutation process is executed to replace the codon overlapping with another codon with a predetermined probability.
- the first mutation process or the second mutation process is selected at random.
- a cross processing unit that performs a cross process on the individuals included in the first generation parent individual population data, the cross process includes a g generation parent individual group representing a g generation parent individual group A predetermined even number of individual data is extracted from the data, two pieces of individual data are selected from the extracted individual data, and a crossover process is executed on the selected two pieces of individual data.
- the intersection processing unit determines an intersection point from a boundary of the codon included in the CDS included in the CDS included in the first individual data and the second individual data which are the selected two individual data, and the intersection The codons included in the first individual data and the second individual data are exchanged at a point.
- the computer is a data generated based on the data representing the amino acid sequence, the number of genes, and the codon frequency table, the first representing the first generation parent individual population including a predetermined number of individual data.
- a parent individual group data acquisition unit for acquiring generation parent individual group data; a mutation processing unit for executing a mutation process on an individual included in the first generation parent individual group data;
- a child individual group data acquisition unit for acquiring first generation child individual group data representing a first generation child individual group, a predetermined evaluation criterion, based on a codon fitness and an evaluation criterion related to the base sequence of the codon
- a non-dominant sort process is performed on the first generation integrated data obtained by integrating the first generation parent individual population data and the first generation child individual population data, and the first generation integrated data
- a non-dominant sort execution unit that classifies all contained individual data for each rank in the Pareto optimal solution, and selects a predetermined number of the individual data in descending order of rank from all the individual data classified for each rank
- An information processing program for functioning as an individual selection unit is provided.
- FIG. 1 When producing a target protein in a microorganism or the like, it is a conceptual diagram showing a conventional technique of introducing a plurality of genes encoding the target protein, (a) is an example of introducing five genes having the same DNA sequence, (B) shows the result of the homologous recombination occurring in the second half of the second gene to the first half of the fifth gene, and the number of genes being reduced to two of the five genes.
- FIG. 2 is a diagram illustrating an example of a hardware configuration of the information processing apparatus 1.
- FIG. It is an exemplary functional block diagram of information processor 1 concerning one embodiment of the present invention. It is a figure for demonstrating congestion distance, (a) is a conceptual diagram of congestion distance, (b) represents the calculation formula of congestion distance.
- the abscissa represents the codon fitness evaluation criterion (minimum value of the CDS codon suitability index (CAI) representing a base sequence that each individual has a plurality of base sequences to be subjected to amino acid translation), It is the graph which plotted the calculation result of the 1st generation, the 10th generation, and the 250th generation by making the evaluation standard (minimum value of the number of mismatched bases) about a base sequence into a vertical axis. It is a figure showing the process result in the Example of this invention. In FIG.
- the evaluation criteria related to the codon suitability (the minimum value of the codon fit index (CAI) of the CDS representing the base sequences that each individual has a plurality of base sequences to be subjected to amino acid translation) are plotted on the horizontal axis, It is the graph which plotted the calculation result of the 1st generation, the 10th generation, and the 250th generation by making the evaluation standard (length of the longest common character string) regarding a base sequence into a vertical axis, respectively.
- FIG. 2 is a conceptual diagram showing gene sequence design according to an embodiment of the present invention.
- a gene sequence group that does not induce homologous recombination is designed and introduced into a microorganism or the like, thereby increasing the production amount of the target protein.
- the data representing the target protein and the data representing the N genes are used as input data, and calculation processing is performed based on an algorithm (hereinafter referred to as the present algorithm).
- the present algorithm uses a multi-objective genetic algorithm for the purpose of simultaneously optimizing a plurality of objective functions having mutually conflicting evaluation criteria.
- the genes 1 to 5 are genes having different base sequences.
- the information processing apparatus 1 includes a processing unit 10, a storage unit 20, an operation unit 30, a display unit 40, and a communication unit 50.
- the processing unit 10 executes various arithmetic processes, and is configured by, for example, a CPU.
- the storage unit 20 stores various data and programs, and includes, for example, a memory, an HDD, an SSD, or the like.
- the program may be preinstalled at the time of shipment of the information processing apparatus 1, may be downloaded as an application from a site on the Web, or may be transferred from another information processing apparatus by wireless communication.
- the operation unit 30 operates the information processing apparatus 1 and includes, for example, a motion recognition device using a touch panel, a keyboard, a voice input unit, a camera, and the like.
- the display unit 40 displays various images (including still images and moving images) and includes, for example, a touch panel display, an organic EL display, electronic paper, and other displays.
- the communication unit 50 transmits / receives various data to / from other information processing apparatuses, and is configured by an arbitrary I / O.
- the bus 100 is composed of a serial bus, a parallel bus, and the like, and electrically connects each part to enable transmission / reception of various data.
- the information processing apparatus 1 is, for example, a multifunction information terminal, such as a PC, a server, a smartphone, a tablet terminal, or a smart watch.
- the information processing apparatus 1 includes an operation unit 30, a display unit 40, a communication unit 50, a processing unit 10, and a storage unit 20.
- the processing unit 10 includes an individual generation unit 101, a parent individual group data acquisition unit 102, a child individual group data acquisition unit 103, an intersection processing unit 104, a mutation processing unit 105, a non-dominant sort execution unit 106, and an individual selection unit 107.
- the storage unit 20 includes an amino acid sequence data storage unit 201, a gene number data storage unit 202, a codon frequency table data storage unit 203, a calculation data storage unit 204, and an evaluation criterion storage unit 205.
- the individual generation unit 101 acquires data representing the amino acid sequence, the number of genes, and the codon frequency table from the amino acid sequence data storage unit 201, the gene number data storage unit 202, and the codon frequency table data storage unit 203, respectively, and the amino acid sequence of the same protein P pieces of individual data representing an individual randomly generated under the restriction of coding.
- p is a positive parameter and can be an arbitrary number.
- the parent individual group data acquisition unit 102 uses the p pieces of individual data generated by the individual generation unit 101 for the present algorithm, and is used as g-th generation parent individual group data representing the g-th generation parent individual group. get. Further, the calculation in this algorithm is to execute a loop calculation that repeatedly executes a predetermined flow a plurality of times as will be described later, and the parent individual group data acquisition unit 102 has a first generation, a second generation,... First-generation parent individual population data representing the g-th generation parent individual population, second-generation parent individual population data, g-th generation parent individual population data are acquired.
- g is a positive number and represents the number of calculation loops in this algorithm.
- the intersection processing unit 104 extracts e (even predetermined number) individual data from the g-th generation parent individual population data, selects two individual data from the extracted individual data, and selects the selected 2 Crossing processing is executed for individual individual data. Then, two pieces of individual data are selected from the pieces of individual data that have not yet undergone the crossover process, and the crossover process is executed. Such a process is repeated for all the extracted e pieces of individual data. Specifically, a crossing point is determined from the boundary of codons included in the CDS included in the first individual data and the second individual data that are the two selected individual data, and the first individual data is determined from the crossing point as a boundary. And the codons included in the second individual data are replaced.
- the selection of the two pieces of individual data is executed at random using, for example, a random number table or the like.
- e is the maximum even number not exceeding “p ⁇ Pc”.
- Pc is a parameter and can be any value greater than 0 and less than 1.
- a method for extracting e pieces of individual data from the g generation parent individual group data including p pieces of individual data is not particularly limited. For example, a “binary tournament selection method” can be used.
- the mutation processing unit 105 performs a mutation process on all individuals included in the g-th generation parent individual population data.
- a first mutation process and a second mutation process different from the first mutation process are executed for each individual data included in the g generation parent individual population data.
- the first mutation process and the second mutation process are randomly determined for each individual data. If the first mutation process is determined, the codons included in the CDS are replaced with codons with a higher frequency than the codons with a predetermined probability Pm for all CDS included in each individual data.
- it is determined as the second mutation process it is the longest base sequence among the CDS included in each individual data, among the base sequences that continuously match at different sites between the respective CDSs or within one CDS. The codon overlapping the longest common character is replaced with another codon with a predetermined probability Pm. Details of these processes will be described later.
- the child individual group data acquisition unit 103 acquires g-th generation child individual group data representing a g-th generation child individual group including an individual for which the mutation processing by the mutation processing unit 105 has been executed.
- the number of individual data included in the g generation child individual population data is equal to the number of individual data included in the g generation parent individual population data, which is p. This is because the mutation process by the mutation processing unit 105 has been executed for all individuals included in the g-th generation parent individual population data.
- the non-dominant sort execution unit 106 integrates the g-th generation parent individual population data and the g-th child individual population data based on a predetermined evaluation criterion, which is an evaluation criterion regarding the codon fitness and the base sequence of the codon.
- a non-dominant sort process is performed on the g-th generation integrated data.
- the individual selection unit 107 selects a predetermined number of pieces of individual data in descending order from the g-th generation integrated data classified for each front (for each rank) in the Pareto optimal solution. For example, p can be adopted as a predetermined number. Then, the parent individual group data acquisition unit 102 acquires p pieces of individual data selected by the individual selection unit 107 as g + 1 generation parent individual group data representing the g + 1 generation parent individual group.
- the individual selection unit 107 may select in order from the largest crowding distance (crowding distance) if individual data having the same rank exists.
- the congestion distance is an average distance between two solutions on both sides of a certain solution.
- FIG. 5A conceptually shows this. Then, the congestion distance is calculated by the calculation formula of FIG.
- the congestion distance corresponds to the average of the lengths of the periphery of the quadrangle indicated by the broken line in FIG.
- the amino acid sequence data storage unit 201 stores data representing an amino acid sequence.
- the amino acid sequence represents the sequence of amino acids in the protein.
- the gene number data storage unit 202 stores data representing the number of genes.
- the number of genes represents the number of CDS included in the individual data.
- the codon frequency table data storage unit 203 stores data representing a codon frequency table.
- the codon frequency table is a table summarizing the codon usage frequency in the host cell.
- the calculation data storage unit 204 includes an individual generation unit 101, a parent individual group data acquisition unit 102, a child individual group data acquisition unit 103, an intersection processing unit 104, a mutation processing unit 105, a non-dominant sort execution unit 106, an individual selection unit 107, and the like. It stores the calculation results in various processes according to.
- the evaluation standard storage unit 205 is a predetermined evaluation standard, and stores evaluation standards related to codon fitness and codon base sequence. Specifically, the evaluation standard regarding the codon matching degree is based on the minimum value of the codon matching index of CDS representing a base sequence that each individual has a plurality of base sequences to be subjected to amino acid translation. Hereinafter, this standard is referred to as a first evaluation standard. In the first evaluation criterion, the individual is evaluated higher as the minimum value of the codon matching index included in the individual is larger. The first evaluation criterion regarding the base sequence of the codon is based on the minimum value of the number of mismatched bases representing the number of bases that do not match each other among the two CDSs included in each individual.
- this standard is referred to as a second evaluation standard.
- the individual is evaluated higher as the minimum value of the number of mismatched bases is larger.
- the second of the evaluation criteria related to the base sequence of codons is the longest base sequence among the CDSs included in each individual, and the base sequence that continuously matches between different CDSs or in different sites within one CDS. Is based on the length of the longest common character string.
- this standard is referred to as a third evaluation standard. In the third evaluation criterion, the individual is evaluated higher as the length of the longest common character string is shorter.
- FIG. 6 is a diagram showing an example of a flowchart for carrying out gene sequence design according to an embodiment of the present invention.
- the process shown in FIG. 6 is a process executed prior to the main routine shown in FIG.
- the processing illustrated in FIG. 6 is referred to as preprocessing.
- the processing unit 10 acquires amino acid sequence data and gene number data from the amino acid sequence data storage unit 201 and the gene number data storage unit 202. Then, the data is stored in a storage unit such as a cache memory (not shown).
- the processing unit 10 acquires codon frequency table data from the codon frequency table data storage unit 203. Then, the data is stored in a storage unit such as a cache memory (not shown).
- the individual generation unit 101 generates p pieces of individual data representing individuals generated randomly under the restriction that the amino acid sequences of the same protein are encoded. For example, 100 pieces of random individual data may be generated.
- CDS protein coding regions
- G, I, V, E, and Q shown in FIG. 7 are amino acid sequences represented by the amino acid sequence data acquired by the processing unit 10 from the amino acid sequence data storage unit 201 in S11 of FIG.
- Each CDS has a different base sequence.
- the parent individual group data acquisition unit 102 acquires p individual data randomly generated by the individual generation unit 101 in S13 as first generation parent individual group data representing the first generation parent individual group.
- the parent individual population data is an archive population that is stored in the processing in the present algorithm.
- the processing unit 10 sets the variable g to 1.
- g is a code representing a g-th generation parent individual group.
- g takes a value from 1 to G (a predetermined generation number G described later).
- the processing unit 10 acquires the first generation parent individual group data from the parent individual group data acquisition unit 102.
- the intersection processing unit 104 and the mutation processing unit 105 execute the intersection processing and the mutation process on the individual data included in the first generation parent individual group data. Note that the crossing process is optional and can be omitted as necessary. Hereinafter, the crossover process and the mutation process will be described with reference to FIGS.
- FIG. 9 is a diagram illustrating an example of a flowchart for performing the intersection processing according to the embodiment of the present invention.
- the intersection processing unit 104 sets a variable i to 0.
- a technique used for such extraction is not particularly limited, and for example, a “binary tournament selection method” can be used.
- e is the maximum even number not exceeding “p ⁇ Pc”.
- Pc is a parameter and can be any value greater than 0 and less than 1.
- the intersection processing unit 104 selects two pieces of individual data at random from (e ⁇ i) pieces of individual data.
- the selection of two pieces of individual data is executed at random using, for example, a random number table or the like.
- intersection processing unit 104 performs the intersection process on the two pieces of individual data selected in S323.
- the intersection process will be specifically described with reference to FIG.
- the two individual data selected in S323 are defined as first individual data and second individual data, respectively.
- the first individual data and the second individual data each have three CDSs and have different base sequences.
- An intersection point is determined from these individual data.
- One intersection point is selected from the codon-codon boundary. Such a determination may be made at random.
- the intersection points in the first individual data and the second individual data are the same place.
- the codons included in the first individual data and the second individual data are exchanged at the intersection point.
- such processing is referred to as intersection processing.
- intersection processing unit 104 increases the variable i by two.
- the mutation process is executed for all individuals included in the g-th generation parent individual population data.
- the mutation process is executed for the p pieces of individual data included in the g-th generation.
- a total of p individual data including the e individual data for which the intersection process has been executed and the pe individual data for which the intersection process has not been executed is combined. Mutation processing is executed for.
- FIG. 11 is a diagram illustrating an example of a flowchart for performing a mutation process according to an embodiment of the present invention.
- the mutation processing unit 105 randomly determines whether to execute the first mutation process or the second mutation process on each individual data included in the g-th generation parent individual population data.
- the mutation process is executed on all of the p pieces of individual data included in the g generation parent individual population data.
- the second mutation process is a mutation process different from the first mutation process.
- the mutation processing unit 105 determines whether or not the determination result in S221 is the first mutation process. And if a determination result is YES, it will progress to S223a and will perform a 1st variation
- the mutation processing unit 105 performs a first mutation process on the individual data. Specifically, for all CDS included in the individual data, each codon is replaced with a codon having a higher frequency than the codon with a predetermined probability Pm.
- the higher frequency codon is obtained from the codon frequency table data acquired from the codon frequency table data storage unit 203 in S12 in the preprocessing of FIG.
- the first mutation process will be described with reference to FIG.
- FIG. 12A shows an example in which three CDS are included in the individual data.
- the broken line in FIG. 12A represents the range of codons to be subjected to the mutation process with the probability Pm.
- “GGC” that is the first codon included in the third CDS, CDS-3, is replaced with a codon that is more frequent than GGC with a probability Pm.
- higher frequency codons are obtained from codon frequency table data.
- “GGT” and “GGA” exist as codons having a higher frequency than “GGC”.
- any one codon is selected at random and replaced with “GGC”.
- Such substitution is not performed.
- Such replacement is performed for all codons included in the individual data.
- such a process is referred to as a first mutation process.
- the first mutation process is intended to increase the minimum CAI value according to the first evaluation criterion described later.
- the mutation processing unit 105 executes a second mutation process on the individual data. Specifically, among the CDS included in the individual data, a codon that overlaps with the longest common character that is the longest base sequence among base sequences that continuously match at different sites within each CDS or within one CDS, Replace with another codon with a predetermined probability Pm.
- the longest common character string will be described with reference to FIG.
- the individual data shown in FIG. 14 includes three CDSs as an example.
- the longest character string that matches continuously is called the longest common character string.
- “GGCATCGTCGA” (the part underlined by a solid line) is the longest common character string, and its length (number of characters) is 11.
- GTCGAGCAG the part with the underline of the broken line
- the longest common character string corresponds to a concept called “the longest common substring” in computer science.
- FIG. 12 (b) shows an example in which three CDS are included in the individual data.
- a mutation process is executed with a probability Pm for the codon overlapping the longest common character string.
- the longest common character string is “GGCATCGTCGA” (the part with a solid underline).
- the first to fourth codons are codons that overlap the longest common character string. is there.
- the broken line in FIG.12 (b) represents the range of the codon used as the object for which a variation
- “GGC”, which is the first codon included in CDS-3, is replaced with another codon with probability Pm.
- FIG. 12B “GGT”, “GGA”, and “GGG” exist as codons different from “GGC”.
- any one codon is selected at random and replaced with “GGC”.
- substitution is not performed.
- substitution may not be possible.
- Such substitution is performed for the codon overlapping the longest common character string.
- a process is referred to as a second mutation process.
- the second mutation process is intended to increase the number of mismatched bases according to a second evaluation criterion described later, and to reduce the longest common character string.
- the calculation result is output to the calculation data storage unit 204.
- the child individual group data acquisition unit 103 generates g-th generation child individual group data.
- a method for generating the g-th generation child individual group data for each execution of the intersection process will be described.
- the child individual population data acquisition unit 103 newly adds p individual data in which p individual data included in the g-th generation parent individual population data has been subjected to mutation processing. It is set as g generation offspring individual population data.
- the child individual population data acquisition unit 103 performs e individual data on which crossing processing has been executed among the p individual data included in the g generation parent individual population data. Then, the p individual data obtained by mutating all the p individual data including the pe individual data for which the cross process has not been executed is newly set as the g-th generation child individual population data.
- the processing unit 10 integrates the g generation parent individual population data and the g generation child individual population data to generate g generation integrated data. As a result, 2p pieces of individual data are included in the g-th generation integrated data.
- the non-dominant sort execution unit 106 is non-dominant with respect to the g-th generation integrated data based on the evaluation criteria determined in advance and the evaluation criteria regarding the codon fitness and the base sequence of the codon. Perform a sort. Then, 2p pieces of individual data are classified for each front (for each rank) in the Pareto optimal solution.
- the individual selecting unit 107 selects a predetermined number of pieces of individual data in descending order from the g-th generation integrated data classified for each front (for each rank) in the Pareto optimal solution.
- the individual selection unit 107 may select, in order from the largest congestion distance, when individual data having the same rank exists when selecting a predetermined number of individual data.
- p can be adopted as a predetermined number.
- the parent individual group data acquisition unit 102 generates and acquires p pieces of individual data selected by the individual selection unit 107 as g + 1 generation parent individual group data representing the g + 1 generation parent individual group.
- such evaluation criteria will be described with reference to FIGS. 13 and 14.
- evaluation criteria when selecting p pieces of individual data from 2p pieces of individual data that have been subjected to non-dominant sorting, evaluation criteria from two viewpoints are used.
- This viewpoint is a viewpoint derived for the purpose of suppressing homologous recombination and increasing the production amount of the target protein.
- the first aspect relates to “codon fitness”. Specifically, it is based on the minimum value of the codon-matching index of CDS, which is a base sequence possessed by each individual and represents the base sequence to be subjected to amino acid translation. This is the first evaluation criterion.
- the second aspect relates to the “codon base sequence”. Furthermore, the second viewpoint is divided into “number of mismatched bases” and “longest common character string”.
- the second evaluation criterion is based on the minimum value of the number of mismatched bases among the two CDSs included in the individual data. Further, the third evaluation criterion is based on the length of the longest common character string in the CDS included in the individual data.
- the first evaluation criterion which is the first viewpoint, relates to “codon fitness”.
- the “codon suitability” is assumed to be higher as more frequently used codons are included in the CDS included in the individual data.
- the minimum value hereinafter referred to as the minimum CAI value
- CAI Codon Adaptation Index
- the minimum CAI value can be obtained by the following equation.
- Ci number of CDS included in individual data
- Ci i-th CDS
- Ci CAI of Ci
- the second evaluation criterion which is the first of the second viewpoints, relates to “the number of mismatched bases”.
- the minimum value of the number of mismatched bases (hereinafter referred to as the minimum number of mismatched bases) is used as an evaluation criterion.
- the mismatch base is a base in which bases constituting a codon do not match by comparing two CDSs (hereinafter referred to as CDS pairs) Ci and Cj among x CDSs included in individual data. It is.
- the number of mismatched bases among the bases constituting Ci and Cj is five. Therefore, the number of mismatched bases of the CDS pair (Ci and Cj) is 5.
- the minimum number of mismatched bases can be determined by the following formula.
- Ci number of CDS included in individual data
- Ci i-th CDS
- Cj jth CDS NN
- Ci, Cj Number of mismatched bases between Ci and Cj
- the third evaluation criterion which is the second of the second viewpoints, relates to the “longest common character string”.
- the “longest common character string” is a sequence in which a character string representing a codon included in a different site within each CDS or one CDS is compared with a character string included in another CDS. Is the longest of the matching strings.
- the first evaluation criterion which is the first viewpoint
- the second evaluation criterion and the third evaluation criterion which are the second viewpoint, it is possible to select individual data having a large ratio of different base sequences and suppress the occurrence of homologous recombination.
- the individual selection unit 107 selects p pieces of individual data in descending order from the g-th generation integrated data (including 2p pieces of individual data) classified for each front (for each rank) in the Pareto optimal solution. To do. Then, the parent individual group data acquisition unit 102 newly sets the selected p pieces of individual data as the g + 1th generation parent individual group data, and generates g + 1st generation parent individual group data.
- the processing unit 10 determines whether or not the variable g exceeds a predetermined generation number G. If the determination result is NO, the process proceeds to S28. On the other hand, if the determination result in S ⁇ b> 27 is YES, the main routine is terminated and the calculation result is output to the calculation data storage unit 204.
- the process proceeds to S28, the variable g is incremented (that is, 1 is added to the variable g), and the process returns to S21 again.
- the parent individual group data acquisition unit 102 acquires the second generation parent individual group data generated in S26. This process is repeated until the variable g reaches a predetermined generation number G. In other words, the processes in S21 to S26 are repeated 250 times.
- FIG. 15 is a graph plotted with the first evaluation standard on the horizontal axis and the second evaluation standard on the vertical axis.
- the plots represented by circles in the graph represent the first generation
- the plots represented by squares represent the 10th generation
- the plots represented by triangles represent the results of the 250th generation.
- the first evaluation criterion has a higher evaluation as the minimum CAI value is larger.
- the point plotted on the right side of the horizontal axis is better, and the point plotted on the left side of the horizontal axis is evaluated. Is bad.
- FIG. 16 is a graph plotted with the first evaluation standard on the horizontal axis and the third evaluation standard on the vertical axis.
- the meaning of each plot represented by a circle, a square, and a triangle is the same as in FIG.
- the third evaluation criterion is better evaluated as the point plotted below the vertical axis and plotted above the vertical axis. It can be said that evaluation is so bad that a point.
- the selection in S26 of the main routine in FIG. 8 uses the first evaluation standard and the second evaluation standard, or the first evaluation standard and the third evaluation standard, and the graphs shown in FIG. 15 and FIG. One of these may be obtained.
- both the first evaluation standard and the second evaluation standard, and the first evaluation standard and the third evaluation standard are used, individual data with high evaluation is obtained from the graph in FIG. 15 and the graph in FIG. 16 for each generation. It is possible to identify and assign points on an arbitrary basis, and select individual data having a high sum of points in these two graphs. Alternatively, instead of a two-dimensional graph as shown in FIGS.
- a three-dimensional graph is obtained with the first evaluation criterion on the x-axis, the second evaluation criterion on the y-axis, and the third evaluation criterion on the z-axis.
- Individual data obtained by creating a high evaluation for each of the three evaluation criteria may be selected.
- the storage unit 20 can be in a form of cloud computing provided in an information processing apparatus such as an external PC or server without being provided in the information processing apparatus 1.
- an external information processing apparatus transmits data necessary for each calculation to the information processing apparatus 1.
- ASIC application specific integrated circuit
- FPGA field-programmable gate array
- DRP Dynamic ReConfigurable Processor
- NGA-II which is a multi-purpose genetic algorithm
- p optimal solutions can be obtained together as in the present algorithm.
- simulated annealing or “(single purpose) genetic algorithm” which is a kind of combination optimization algorithm may be used. In this case, however, p optimal solutions cannot be obtained collectively, so it is necessary to repeat the calculation at least p times or more to obtain p optimal solutions.
- the calculation results of these two or more algorithms may be mixed. In this case, an arbitrary integer ⁇ less than or equal to p is set, ⁇ individuals are selected from the calculation results by a certain algorithm, p ⁇ individuals are selected from the calculation results by another algorithm, and these are combined p Individual individuals may be used.
- the present invention provides First-generation parent individual population data representing a first-generation parent individual population data that is generated based on data representing an amino acid sequence, the number of genes, and a codon frequency table and includes a predetermined number of individual data
- a parent individual population data acquisition step to acquire A mutation process step for performing a mutation process on an individual included in the first generation parent individual population data
Landscapes
- Bioinformatics & Cheminformatics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Genetics & Genomics (AREA)
- Molecular Biology (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Computational Biology (AREA)
- Analytical Chemistry (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Theoretical Computer Science (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
相同組み換えを抑制し、目的タンパク質の生産量を高めることが可能な情報処理装置及びプログラムを提供するものである。 本発明によれば、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部と、前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部と、前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部と、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部と、前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部と、を有する情報処理装置が提供される。
Description
本発明は、目的タンパク質の生産量を高めることができる情報処理装置及びプログラムに関する。
微生物等に目的タンパク質を生産させる際に、目的タンパク質をコードする遺伝子を複数個導入する手法が知られている。かかる遺伝子は、同じDNA配列を有するものが利用されることが多い。しかし、同じDNA配列を有する複数個の遺伝子を導入すると、これらの遺伝子間で相同組み換えが生じ、遺伝子の一部が欠損してしまう。ここで、相同組み換えとは、DNAの塩基配列がよく似た部位(相同部位)で起こる組み換えのことである。これを概念的に表したのが図1である。図1(a)は、同じDNA配列を有する5個の遺伝子を導入した例を示す。かかる5個の遺伝子のうち、2個目の遺伝子の後半部分~5個目の遺伝子の前半部分において相同組み換えが生じると、図1(b)に示されるように、遺伝子の数が2つまで減少してしまい、目的タンパク質の生産効率が低下してしまう。
特許文献1には、合成核酸分子を取得するための方法であって、(i)ポリペプチドのアミノ酸繰り返し領域由来のアミノ酸配列を提供する工程;(ii)前記アミノ酸配列をそれぞれコードする複数のサンプルコドン最適化核酸配列を推測する工程;(iii)前記複数のサンプルコドン最適化核酸配列を、配列相同性により整列させ、前記複数のサンプルコドン最適化核酸配列を含む近隣結合ツリーを構築する工程;(iv)前記複数のサンプルコドン最適化核酸配列の1つのみを選択する工程;ならびに、(v)前記選択されたサンプルコドン最適化核酸配列を含む核酸分子を取得する工程を含む、方法が開示されている。
本発明は、相同組み換えを抑制し、目的タンパク質の生産量を高めることが可能な情報処理装置及びプログラムを提供するものである。
本発明によれば、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部と、前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部と、前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部と、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部と、前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部と、を有する情報処理装置が提供される。
本発明によれば、異なる2つの評価基準に基づいて、相同組み換えを抑制し、目的タンパク質の生産量を高めることが可能となる。
以下、本発明の種々の実施形態を例示する。以下に示す実施形態は互いに組み合わせ可能である。
好ましくは、前記個体選択部は、前記予め定められた数の前記個体データを選択するときに、前記ランクが同じ前記個体データが存在する場合には、混雑距離が大きいものから順に選択する。
好ましくは、前記親個体集団データ取得部は、前記個体選択部により選択された前記個体データを、第2世代の親個体集団を表す第2世代親個体集団データとし、前記変異処理部、前記非優越ソート実行部及び前記個体選択部による処理を、予め定められた世代数となるまで実行する。
好ましくは、前記コドン適合度に関する評価基準は、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする。
好ましくは、前記個体に含まれる前記コドン適合インデックスの最小値が大きいほど、前記個体の評価を高くする。
好ましくは、前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる2つの前記CDSのうち、互いに一致しない塩基の数を表す不一致塩基数の最小値を基準とする。
好ましくは、前記不一致塩基数の最小値が大きいほど、前記個体の評価を高くする。
好ましくは、前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる前記CDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字列の長さを基準とする。
好ましくは、前記最長共通文字列の長さが短いほど、前記個体を高く評価する。
好ましくは、前記変異処理部は、第g世代の親個体集団を表す第g世代親個体集団データに含まれる各個体データに対し、第1変異処理及び前記第1変異処理とは異なる第2変異処理を実行する。
好ましくは、前記変異処理部は、前記各個体に含まれる全てのCDSに対し、前記CDSに含まれる前記コドンを、予め定められた確率で前記コドンより高頻度のコドンに置換する第1変異処理を実行する。
好ましくは、前記変異処理部は、前記各個体に含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なる前記コドンを、予め定められた確率で他のコドンに置換する第2変異処理を実行する。
好ましくは、前記第1変異処理又は前記第2変異処理は、ランダムに選択される。
好ましくは、前記第1世代親個体集団データに含まれる個体に対し、交差処理を実行する交差処理部を有し、前記交差処理は、第g世代の親個体集団を表す第g世代親個体集団データから予め定められた偶数個の個体データを抽出し、前記抽出された個体データから2個の個体データを選択し、前記選択された2個の個体データに対して交差処理を実行する。
好ましくは、前記交差処理部は、前記選択された2個の個体データである第1個体データ及び第2個体データに含まれる前記CDSに含まれる前記コドンの境界から交差ポイントを決定し、前記交差ポイントを境として前記第1個体データと前記第2個体データに含まれる前記コドンを入れ替える。
好ましくは、コンピュータを、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部、前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部、前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部、前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部、として機能させるための情報処理プログラムが提供される。
好ましくは、前記個体選択部は、前記予め定められた数の前記個体データを選択するときに、前記ランクが同じ前記個体データが存在する場合には、混雑距離が大きいものから順に選択する。
好ましくは、前記親個体集団データ取得部は、前記個体選択部により選択された前記個体データを、第2世代の親個体集団を表す第2世代親個体集団データとし、前記変異処理部、前記非優越ソート実行部及び前記個体選択部による処理を、予め定められた世代数となるまで実行する。
好ましくは、前記コドン適合度に関する評価基準は、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする。
好ましくは、前記個体に含まれる前記コドン適合インデックスの最小値が大きいほど、前記個体の評価を高くする。
好ましくは、前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる2つの前記CDSのうち、互いに一致しない塩基の数を表す不一致塩基数の最小値を基準とする。
好ましくは、前記不一致塩基数の最小値が大きいほど、前記個体の評価を高くする。
好ましくは、前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる前記CDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字列の長さを基準とする。
好ましくは、前記最長共通文字列の長さが短いほど、前記個体を高く評価する。
好ましくは、前記変異処理部は、第g世代の親個体集団を表す第g世代親個体集団データに含まれる各個体データに対し、第1変異処理及び前記第1変異処理とは異なる第2変異処理を実行する。
好ましくは、前記変異処理部は、前記各個体に含まれる全てのCDSに対し、前記CDSに含まれる前記コドンを、予め定められた確率で前記コドンより高頻度のコドンに置換する第1変異処理を実行する。
好ましくは、前記変異処理部は、前記各個体に含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なる前記コドンを、予め定められた確率で他のコドンに置換する第2変異処理を実行する。
好ましくは、前記第1変異処理又は前記第2変異処理は、ランダムに選択される。
好ましくは、前記第1世代親個体集団データに含まれる個体に対し、交差処理を実行する交差処理部を有し、前記交差処理は、第g世代の親個体集団を表す第g世代親個体集団データから予め定められた偶数個の個体データを抽出し、前記抽出された個体データから2個の個体データを選択し、前記選択された2個の個体データに対して交差処理を実行する。
好ましくは、前記交差処理部は、前記選択された2個の個体データである第1個体データ及び第2個体データに含まれる前記CDSに含まれる前記コドンの境界から交差ポイントを決定し、前記交差ポイントを境として前記第1個体データと前記第2個体データに含まれる前記コドンを入れ替える。
好ましくは、コンピュータを、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部、前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部、前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部、前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部、として機能させるための情報処理プログラムが提供される。
<実施形態>
以下、図面を用いて本発明の実施形態について説明する。以下に示す実施形態中で示した各種特徴事項は、互いに組み合わせ可能である。
以下、図面を用いて本発明の実施形態について説明する。以下に示す実施形態中で示した各種特徴事項は、互いに組み合わせ可能である。
<本発明の一実施形態に係る遺伝子配列設計>
図2は、本発明の一実施形態に係る遺伝子配列設計を表す概念図である。一実施形態に係る遺伝子配列設計は、相同組み換えを誘発しない遺伝子配列群を設計し、微生物等に導入することで、目的タンパク質の生産量を高めるものである。図2(a)に示されるように、目的タンパク質を表すデータ及びN個の遺伝子を表すデータを入力データとし、アルゴリズム(以下、本アルゴリズムという)に基いて計算処理し、かかる計算結果である目的タンパク質をコードするN個の遺伝子配列群を表すデータを出力する。ここで、本アルゴリズムは、相競合する評価基準を持つ複数の目的関数を同時に最適化することを目的とする多目的遺伝的アルゴリズムを利用する。これにより、図2(b)に示されるように、例えば導入された遺伝子が5個である場合、かかる5個の遺伝子データに相同組み換えが生じず、全ての遺伝子から目的タンパク質が生産される。ここで、図2(b)では、1~5までの遺伝子は、互いに塩基配列が異なる遺伝子である。
図2は、本発明の一実施形態に係る遺伝子配列設計を表す概念図である。一実施形態に係る遺伝子配列設計は、相同組み換えを誘発しない遺伝子配列群を設計し、微生物等に導入することで、目的タンパク質の生産量を高めるものである。図2(a)に示されるように、目的タンパク質を表すデータ及びN個の遺伝子を表すデータを入力データとし、アルゴリズム(以下、本アルゴリズムという)に基いて計算処理し、かかる計算結果である目的タンパク質をコードするN個の遺伝子配列群を表すデータを出力する。ここで、本アルゴリズムは、相競合する評価基準を持つ複数の目的関数を同時に最適化することを目的とする多目的遺伝的アルゴリズムを利用する。これにより、図2(b)に示されるように、例えば導入された遺伝子が5個である場合、かかる5個の遺伝子データに相同組み換えが生じず、全ての遺伝子から目的タンパク質が生産される。ここで、図2(b)では、1~5までの遺伝子は、互いに塩基配列が異なる遺伝子である。
<ハードウェア構成>
次に、本発明の一実施形態に係る情報処理装置1のハードウェア構成の例について、図3を用いて説明する。情報処理装置1は、処理部10、記憶部20、操作部30、表示部40及び通信部50を有する。処理部10は、種々の演算処理を実行するものであり、例えば、CPU等により構成される。記憶部20は、種々のデータやプログラムを記憶するものであり、例えば、メモリ、HDD又はSSD等により構成される。ここで、プログラムは、情報処理装置1の出荷時点においてプリインストールされていてもよく、Web上のサイトからアプリケーションとしてダウンロードしてもよく、無線通信により他の情報処理装置から転送されてもよい。操作部30は、情報処理装置1を操作するものであり、例えば、タッチパネル、キーボード、音声入力部、カメラ等を利用した動き認識装置等により構成される。表示部40は、種々の画像(静止画及び動画を含む)を表示するものであり、例えば、タッチパネルディスプレイ、有機ELディスプレイ、電子ペーパーその他のディスプレイで構成される。通信部50は、他の情報処理装置と種々のデータを送受信するものであり、任意のI/Oにより構成される。バス100はシリアルバス、パラレルバス等で構成され、各部を電気的に接続し、種々のデータの送受信を可能にするものである。
次に、本発明の一実施形態に係る情報処理装置1のハードウェア構成の例について、図3を用いて説明する。情報処理装置1は、処理部10、記憶部20、操作部30、表示部40及び通信部50を有する。処理部10は、種々の演算処理を実行するものであり、例えば、CPU等により構成される。記憶部20は、種々のデータやプログラムを記憶するものであり、例えば、メモリ、HDD又はSSD等により構成される。ここで、プログラムは、情報処理装置1の出荷時点においてプリインストールされていてもよく、Web上のサイトからアプリケーションとしてダウンロードしてもよく、無線通信により他の情報処理装置から転送されてもよい。操作部30は、情報処理装置1を操作するものであり、例えば、タッチパネル、キーボード、音声入力部、カメラ等を利用した動き認識装置等により構成される。表示部40は、種々の画像(静止画及び動画を含む)を表示するものであり、例えば、タッチパネルディスプレイ、有機ELディスプレイ、電子ペーパーその他のディスプレイで構成される。通信部50は、他の情報処理装置と種々のデータを送受信するものであり、任意のI/Oにより構成される。バス100はシリアルバス、パラレルバス等で構成され、各部を電気的に接続し、種々のデータの送受信を可能にするものである。
<機能ブロック図>
次に、情報処理装置1の機能について、図4の機能ブロック図を用いて説明する。情報処理装置1は、例えば、多機能情報端末であり、PC、サーバ、スマートフォン、タブレット端末、スマートウォッチ等である。情報処理装置1は、操作部30、表示部40及び通信部50と、処理部10と、記憶部20を備える。処理部10は、個体生成部101、親個体集団データ取得部102、子個体集団データ取得部103、交差処理部104、変異処理部105、非優越ソート実行部106、個体選択部107を備える。また、記憶部20は、アミノ酸配列データ記憶部201、遺伝子数データ記憶部202、コドン頻度表データ記憶部203、計算データ記憶部204、評価基準記憶部205を備える。
次に、情報処理装置1の機能について、図4の機能ブロック図を用いて説明する。情報処理装置1は、例えば、多機能情報端末であり、PC、サーバ、スマートフォン、タブレット端末、スマートウォッチ等である。情報処理装置1は、操作部30、表示部40及び通信部50と、処理部10と、記憶部20を備える。処理部10は、個体生成部101、親個体集団データ取得部102、子個体集団データ取得部103、交差処理部104、変異処理部105、非優越ソート実行部106、個体選択部107を備える。また、記憶部20は、アミノ酸配列データ記憶部201、遺伝子数データ記憶部202、コドン頻度表データ記憶部203、計算データ記憶部204、評価基準記憶部205を備える。
操作部30、表示部40及び通信部50の各機能については、図3の説明を参照されたい。
<処理部10>
次に、処理部10の機能について説明する。個体生成部101は、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータをそれぞれアミノ酸配列データ記憶部201、遺伝子数データ記憶部202及びコドン頻度表データ記憶部203から取得し、同じタンパク質のアミノ酸配列をコードするという制約下でランダムに生成した個体を表す個体データをp個生成するものである。ここで、pは正の数のパラメータであり、任意の数とすることができる。
次に、処理部10の機能について説明する。個体生成部101は、アミノ酸配列、遺伝子数及びコドン頻度表を表すデータをそれぞれアミノ酸配列データ記憶部201、遺伝子数データ記憶部202及びコドン頻度表データ記憶部203から取得し、同じタンパク質のアミノ酸配列をコードするという制約下でランダムに生成した個体を表す個体データをp個生成するものである。ここで、pは正の数のパラメータであり、任意の数とすることができる。
親個体集団データ取得部102は、個体生成部101が生成したp個の個体データを、本アルゴリズムに利用するデータであって、第g世代の親個体集団を表す第g世代親個体集団データとして取得する。さらに、本アルゴリズムにおける計算は、後述するように所定のフローを複数回繰り返し実行するループ計算を実行するものであり、親個体集団データ取得部102は、第1世代、第2世代、・・・第g世代の親個体集団を表す第1世代親個体集団データ、第2世代親個体集団データ・・・第g世代親個体集団データを取得する。ここで、gは正の数であり、本アルゴリズムにおける計算のループ数を表す。
交差処理部104は、第g世代親個体集団データからe個(予め定められた偶数個)の個体データを抽出し、抽出された個体データから2個の個体データを選択し、選択された2個の個体データに対して交差処理を実行するものである。そして、まだ交差処理が行われていない個体データの中から2個の個体データを選択し、交差処理を実行する。かかる処理を、抽出されたe個の個体データの全てに対して繰り返す。具体的には、選択された2個の個体データである第1個体データ及び第2個体データに含まれるCDSに含まれるコドンの境界から交差ポイントを決定し、交差ポイントを境として第1個体データと第2個体データに含まれるコドンを入れ替える。ここで、2個の個体データの選択は、例えば乱数表等を利用してランダムに実行される。ここで、eは「p×Pc」を超えない最大の偶数である。なお、Pcはパラメータであり、0より大きく1より小さい任意の値とすることができる。p個の個体データを含む第g世代親個体集団データからe個の個体データを抽出する手法は特に限定されないが、例えば「binary tournament selection法」を用いることができる。
変異処理部105は、第g世代親個体集団データに含まれる全ての個体に対し、変異処理を実行するものである。本実施形態では、第g世代親個体集団データに含まれる各個体データに対し、第1変異処理及び前記第1変異処理とは異なる第2変異処理を実行する。具体的には、各個体データに対し、第1変異処理及び第2変異処理をランダムに決定する。そして、第1変異処理と決定された場合、各個体データに含まれる全てのCDSに対し、CDSに含まれるコドンを、予め定められた確率Pmでかかるコドンより高頻度のコドンに置換する。また、第2変異処理と決定された場合、各個体データに含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なるコドンを、予め定められた確率Pmで他のコドンに置換する。これらの処理の詳細については後述する。
子個体集団データ取得部103は、変異処理部105による変異処理が実行された個体を含む第g世代の子個体集団を表す第g世代子個体集団データを取得する。ここで、第g世代子個体集団データに含まれる個体データの数は、第g世代親個体集団データに含まれる個体データの数と等しく、p個である。これは、変異処理部105による変異処理が、第g世代親個体集団データに含まれる全ての個体に実行されたためである。
非優越ソート実行部106は、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、第g世代親個体集団データ及び第g子個体集団データを統合した第g世代統合データに対し、非優越ソート処理を実行するものである。第g世代統合データに含まれる個体データの数は、2p(=p+p)個である。そして、全個体をパレート最適解におけるフロント毎(ランク毎)に分類する。
個体選択部107は、パレート最適解におけるフロント毎(ランク毎)に分類された第g世代統合データから、ランクの高い順に定められた数の個体データを選択するものである。例えば、予め定められた数として、pを採用することができる。そして、親個体集団データ取得部102は、個体選択部107により選択されたp個の個体データを、第g+1世代の親個体集団を表す第g+1世代親個体集団データとして取得する。
ここで、個体選択部107は、定められた数の個体データを選択するときに、ランクが同じ個体データが存在する場合には、混雑距離(Crowding Distance)が大きいものから順に選択することとしてもよい。ここで、混雑距離とは、ある解の両側にある2つの解の平均距離である。これを概念的に表したのが図5(a)である。そして、図5(b)の計算式により、混雑距離が計算される。ここで、混雑距離は、図5(a)において破線で示される四角形の周囲の長さの平均に相当する。
<記憶部20>
次に、記憶部20の機能について説明する。アミノ酸配列データ記憶部201は、アミノ酸配列を表すデータを記憶するものである。アミノ酸配列は、タンパク質中のアミノ酸の配列を表すものである。
次に、記憶部20の機能について説明する。アミノ酸配列データ記憶部201は、アミノ酸配列を表すデータを記憶するものである。アミノ酸配列は、タンパク質中のアミノ酸の配列を表すものである。
遺伝子数データ記憶部202は、遺伝子数を表すデータを記憶するものである。ここで、本実施形態では、遺伝子数は、個体データに含まれるCDSの数を表すものとする。
コドン頻度表データ記憶部203は、コドン頻度表を表すデータを記憶するものである。コドン頻度表は、宿主細胞におけるコドンの使用頻度をまとめた表である。
計算データ記憶部204は、個体生成部101、親個体集団データ取得部102、子個体集団データ取得部103、交差処理部104、変異処理部105、非優越ソート実行部106、個体選択部107等による種々の処理における計算結果を記憶するものである。
評価基準記憶部205は、予め定められた評価基準であって、コドン適合度及びコドンの塩基配列に関する評価基準を記憶するものである。具体的には、コドン適合度に関する評価基準は、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする。以下、かかる基準を第1評価基準という。第1評価基準においては、個体に含まれるコドン適合インデックスの最小値が大きいほど、個体が高く評価される。そして、コドンの塩基配列に関する評価基準のうちの1つ目は、各個体に含まれる2つのCDSのうち、互いに一致しない塩基の数を表す不一致塩基数の最小値を基準とする。以下、かかる基準を第2評価基準という。第2評価基準においては、不一致塩基数の最小値が大きいほど、個体が高く評価される。コドンの塩基配列に関する評価基準のうちの2つ目は、各個体に含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字列の長さを基準とする。以下、かかる基準を第3評価基準という。第3評価基準においては、最長共通文字列の長さが短いほど、個体が高く評価される。
次に、以上説明した種々の機能、処理及び基準の詳細について、図6~図13を用いて説明する。
<前処理>
図6は、本発明の一実施形態に係る遺伝子配列設計を実施するためのフローチャートの一例を示す図である。図6に示される処理は、図8に示されるメインルーチンに先立ち実行される処理である。以下、図6に示される処理を前処理という。
図6は、本発明の一実施形態に係る遺伝子配列設計を実施するためのフローチャートの一例を示す図である。図6に示される処理は、図8に示されるメインルーチンに先立ち実行される処理である。以下、図6に示される処理を前処理という。
まず、S11において、処理部10は、アミノ酸配列データ記憶部201及び遺伝子数データ記憶部202から、アミノ酸配列データ及び遺伝子数データを取得する。そして、図示しないキャッシュメモリ等の記憶部にデータを記憶する。
次に、S12において、処理部10は、コドン頻度表データ記憶部203からコドン頻度表データを取得する。そして、図示しないキャッシュメモリ等の記憶部にデータを記憶する。
次に、S13において、個体生成部101は、同じタンパク質のアミノ酸配列をコードするという制約下でランダムに生成した個体を表す個体データをp個生成する。例えば、ランダムな個体データを100個生成してもよい。
(個体データ)
ここで、図7を用いて個体データについて説明する。図7に示されるように、本実施形態では、1つの個体を、同じアミノ酸をコードする複数のタンパクコード領域(CDS)として表現する。図7に示される個体データでは、CDSがx個である。これは、図6のS11において処理部10が遺伝子数データ記憶部202から取得した遺伝子数データが表す遺伝子の数である。各CDSはそれぞれ同じアミノ酸をコードする。ここで、図7に示されるG,I,V,E,Qは、図6のS11において処理部10がアミノ酸配列データ記憶部201から取得したアミノ酸配列データが表すアミノ酸配列である。また、各CDSは、それぞれ塩基配列が異なっている。
ここで、図7を用いて個体データについて説明する。図7に示されるように、本実施形態では、1つの個体を、同じアミノ酸をコードする複数のタンパクコード領域(CDS)として表現する。図7に示される個体データでは、CDSがx個である。これは、図6のS11において処理部10が遺伝子数データ記憶部202から取得した遺伝子数データが表す遺伝子の数である。各CDSはそれぞれ同じアミノ酸をコードする。ここで、図7に示されるG,I,V,E,Qは、図6のS11において処理部10がアミノ酸配列データ記憶部201から取得したアミノ酸配列データが表すアミノ酸配列である。また、各CDSは、それぞれ塩基配列が異なっている。
図6に戻り、前処理についてさらに説明する。S14において、親個体集団データ取得部102は、S13において個体生成部101がランダムに生成したp個の個体データを第1世代の親個体集団を表す第1世代親個体集団データとして取得する。親個体集団データは、本アルゴリズムにおける処理において保存されるアーカイブ母集団である。そして、第1世代親個体集団データを取得すると、前処理を終了する。
<メインルーチン>
次に、図8を用いて、本アルゴリズムにおけるメインルーチンについて説明する。まず、S20において、処理部10は、変数gを1にセットする。ここで、gは第g世代の親個体集団を表す符号である。gは、1~G(後述する予め定められた世代数G)までの値をとる。
次に、図8を用いて、本アルゴリズムにおけるメインルーチンについて説明する。まず、S20において、処理部10は、変数gを1にセットする。ここで、gは第g世代の親個体集団を表す符号である。gは、1~G(後述する予め定められた世代数G)までの値をとる。
次に、S21において、処理部10は、親個体集団データ取得部102から第1世代親個体集団データを取得する。
次に、S22において、交差処理部104及び変異処理部105は、第1世代親個体集団データに含まれる個体データに対して交差処理及び変異処理を実行する。なお、交差処理は任意であり、必要に応じて省略することができる。以下、図9~図12を用いて交差処理及び変異処理について説明する。
<交差処理>
まず、図9及び図10を用いて交差処理について説明する。図9は、本発明の一実施形態に係る交差処理を実施するためのフローチャートの一例を示す図である。まず、S321において、交差処理部104は、変数iを0にセットする。
まず、図9及び図10を用いて交差処理について説明する。図9は、本発明の一実施形態に係る交差処理を実施するためのフローチャートの一例を示す図である。まず、S321において、交差処理部104は、変数iを0にセットする。
次に、S322において、交差処理部104は、処理部10又は親個体集団データ取得部102から、第g世代親個体集団データ(図8におけるメインルーチンでg=1の場合は第1世代親個体集団データ)を取得する。そして、第g世代親個体集団データに含まれるp個の個体データから、(e-i)個(現時点ではi=0のためにe個)の個体データをランダムに抽出する。かかる抽出に利用する手法は特に限定されないが、例えば「binary tournament selection法」を用いることができる。ここで、eは「p×Pc」を超えない最大の偶数である。なお、Pcはパラメータであり、0より大きく1より小さい任意の値とすることができる。
次に、S323において、交差処理部104は、(e-i)個の個体データからランダムに2個の個体データを選択する。2個の個体データの選択は、例えば乱数表等を利用してランダムに実行される。
次に、S324において、交差処理部104は、S323にて選択された2個の個体データに対して交差処理を実行する。ここで、交差処理について、図10を用いて具体的に説明する。
図10に示されるように、S323にて選択された2個の個体データをそれぞれ第1個体データ及び第2個体データとする。図10の例では、第1個体データ及び第2個体データはそれぞれ3つのCDSを有し、異なる塩基配列を有する。これらの個体データから、交差ポイントを決定する。交差ポイントは、コドンとコドンの境界から1箇所選ばれる。かかる決定はランダムに行われてもよい。本実施形態では、第1個体データと第2個体データにおける交差ポイントは同じ場所とする。そして、交差ポイントを境として、第1個体データと第2個体データに含まれるコドンを入れ替える。本実施形態では、かかる処理を交差処理という。
図9に戻り、交差処理についてさらに説明する。S325において、交差処理部104は、変数iを2増やす。
次に、S326において、交差処理部104は、変数i=eであるか否かを判定する。そして、判定結果がNOであれば、再びS323に戻る。一方、判定結果がYESであれば、交差処理を終了し、かかる計算結果を計算データ記憶部204へ出力する。ここで、現時点ではi=2であり、eが2よりも大きいとすると、S326からS323へ戻ることになる。そして、まだ交差処理が実行されていない(e-2)個の個体データからランダムに2個の個体データを選択する。かかる処理を、S326における判定結果がYES、つまり、e個の個体データ全てに対して交差処理が実行されるまで繰り返す。なお、前述のとおり、かかる交差処理は任意であり、必要に応じて省略することができる。
<変異処理>
次に、図11及び図12を用いて、変異処理について説明する。変異処理は、第g世代親個体集団データに含まれる全ての個体に対して実行される。ここで、S22において交差処理が実行されていない場合には、第g世代に含まれるp個の個体データに対して変異処理を実行する。一方、S22において交差処理が実行された場合には、交差処理が実行されたe個の個体データと、交差処理が実行されていないp-e個の個体データを合わせた計p個の個体データに対して変異処理を実行する。
次に、図11及び図12を用いて、変異処理について説明する。変異処理は、第g世代親個体集団データに含まれる全ての個体に対して実行される。ここで、S22において交差処理が実行されていない場合には、第g世代に含まれるp個の個体データに対して変異処理を実行する。一方、S22において交差処理が実行された場合には、交差処理が実行されたe個の個体データと、交差処理が実行されていないp-e個の個体データを合わせた計p個の個体データに対して変異処理を実行する。
図11は、本発明の一実施形態に係る変異処理を実施するためのフローチャートの一例を示す図である。まず、S221において、変異処理部105は、第g世代親個体集団データに含まれる各個体データに対し、第1変異処理又は第2変異処理のいずれを実行するかをランダムに決定する。本実施形態では、第g世代親個体集団データに含まれるp個の個体データの全てに対して変異処理を実行するものとする。ここで、第2変異処理は、第1変異処理とは異なる変異処理である。
次に、S222において、変異処理部105は、S221における決定結果が第1変異処理であるか否かを判定する。そして、判定結果がYESであれば、S223aに進み、第1変異処理を実行する。一方、判定結果がNOであれば、S223bに進み、第2変異処理を実行する。
(第1変異処理)
次に、S223aにおいて、変異処理部105は、個体データに対して第1変異処理を実行する。具体的には、個体データに含まれる全てのCDSに対し、各コドンを予め定められた確率Pmでかかるコドンより高頻度のコドンに置換する。ここで、より高頻度のコドンは、図6の前処理におけるS12でコドン頻度表データ記憶部203から取得したコドン頻度表データより得る。ここで、図12(a)を用いて第1変異処理について説明する。
次に、S223aにおいて、変異処理部105は、個体データに対して第1変異処理を実行する。具体的には、個体データに含まれる全てのCDSに対し、各コドンを予め定められた確率Pmでかかるコドンより高頻度のコドンに置換する。ここで、より高頻度のコドンは、図6の前処理におけるS12でコドン頻度表データ記憶部203から取得したコドン頻度表データより得る。ここで、図12(a)を用いて第1変異処理について説明する。
図12(a)は、個体データに3つのCDSが含まれる例を示す。図12(a)に示されるように、第1変異処理では、個体データに含まれる3つのCDSについて、全てのコドン(5×3=15個のコドン)に対して確率Pmで変異処理を実行する。なお、図12(a)中の破線は、確率Pmで変異処理が実行される対象となるコドンの範囲を表すものである。一例として、3つ目のCDSであるCDS-3に含まれる最初のコドンである「GGC」を、確率PmでGGCより高頻度なコドンに置換する。ここで、より高頻度なコドンは、コドン頻度表データから得る。図12(a)の例では、「GGC」より高頻度なコドンは、「GGT」及び「GGA」が存在する。このように、より高頻度なコドンが複数ある場合には、いずれか1つのコドンをランダムに選び、「GGC」と置換する。なお、「GGC」より高頻度なコドンが存在しない場合、かかる置換はされない。このような置換を、個体データに含まれる全てのコドンに対して実行する。本実施形態では、このような処理を第1変異処理という。ここで、第1変異処理は、後述する第1評価基準に係る最小CAI値を大きくすることを意図するものである。
(第2変異処理)
図11に戻り、変異処理についてさらに説明する。S223bにおいて、変異処理部105は、個体データに対して第2変異処理を実行する。具体的には、個体データに含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なるコドンを、予め定められた確率Pmで他のコドンに置換する。ここで、図14を用いて、最長共通文字列について説明する。
図11に戻り、変異処理についてさらに説明する。S223bにおいて、変異処理部105は、個体データに対して第2変異処理を実行する。具体的には、個体データに含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なるコドンを、予め定められた確率Pmで他のコドンに置換する。ここで、図14を用いて、最長共通文字列について説明する。
「最長共通文字列」
図14に示される個体データは、一例として3つのCDSを含むものである。ここで、各CDSに含まれる5個のコドンを表す文字列(3個の塩基(=文字)×5=15文字)を、他のCDS又は1つのCDS内部の異なる部位に含まれる文字列と対比して、連続して一致する文字列の中で最も長いものを最長共通文字列という。図14の例では、「GGCATCGTCGA」(実線の下線が付された部分)が最長共通文字列となり、その長さ(文字数)は11である。なお、「GTCGAGCAG」(破線の下線が付された部分)も共通文字列であるが、長さが9であり、最長ではないので最長共通文字列とならない。なお、最長共通文字列は、計算機科学における最長共通部分文字列(The longest common substring)と呼ばれている概念に相当する。
図14に示される個体データは、一例として3つのCDSを含むものである。ここで、各CDSに含まれる5個のコドンを表す文字列(3個の塩基(=文字)×5=15文字)を、他のCDS又は1つのCDS内部の異なる部位に含まれる文字列と対比して、連続して一致する文字列の中で最も長いものを最長共通文字列という。図14の例では、「GGCATCGTCGA」(実線の下線が付された部分)が最長共通文字列となり、その長さ(文字数)は11である。なお、「GTCGAGCAG」(破線の下線が付された部分)も共通文字列であるが、長さが9であり、最長ではないので最長共通文字列とならない。なお、最長共通文字列は、計算機科学における最長共通部分文字列(The longest common substring)と呼ばれている概念に相当する。
図12(b)は、個体データに3つのCDSが含まれる例を示す。図12(b)に示されるように、第2変異処理では、個体データに含まれる3つのCDSについて、CDSに含まれる5個のコドンを表す文字列(3個の塩基(=文字)×5=15文字)のうち、最長共通文字列と重なるコドンに対して確率Pmで変異処理を実行する。ここで、図12(b)の例では、最長共通文字列は「GGCATCGTCGA」(実線の下線が付された部分)である。図12(b)の例では、2つ目及び3つ目のCDSであるCDS-2及びCDS-3に含まれるコドンのうち、1~4つ目のコドンが最長共通文字列と重なるコドンである。なお、図12(b)中の破線は、確率Pmで変異処理が実行される対象となるコドンの範囲を表すものである。一例として、CDS-3に含まれる最初のコドンである「GGC」を、確率Pmで他のコドンに置換する。図12(b)の例では、「GGC」とは異なるコドンとして、「GGT」、「GGA」及び「GGG」が存在する。このように、他のコドンが複数ある場合には、いずれか1つのコドンをランダムに選び、「GGC」と置換する。なお、「GGC」以外のコドンが存在しない場合には、かかる置換はされない。例えば、特定のアミノ酸をコードするコドンが1種類しか存在しないときには、置換ができない場合があるためである。このような置換を、最長共通文字列と重なるコドンに対して実行する。本実施形態では、このような処理を第2変異処理という。ここで、第2変異処理は、後述する第2評価基準に係る不一致塩基数を大きくし、最長共通文字列を小さくすることを意図するものである。
そして、第1変異処理及び第2変異処理が終了すると、かかる計算結果を計算データ記憶部204へ出力する。
図8に戻り、メインルーチンについてさらに説明する。S22において、第g世代親個体集団データに対して変異処理部105による変異処理、必要に応じて、交差処理部104による交差処理が実行された後、S23に進む。
次に、S23において、子個体集団データ取得部103は、第g世代子個体集団データを生成する。以下、交差処理の実行の有無毎に、第g世代子個体集団データの生成の仕方について説明する。
1.S22において変異処理のみが実行された場合
子個体集団データ取得部103は、第g世代親個体集団データに含まれるp個の個体データが全て変異処理されたp個の個体データを、新たに第g世代子個体集団データとする。
子個体集団データ取得部103は、第g世代親個体集団データに含まれるp個の個体データが全て変異処理されたp個の個体データを、新たに第g世代子個体集団データとする。
2.S22において変異処理及び交差処理が実行された場合
子個体集団データ取得部103は、第g世代親個体集団データに含まれるp個の個体データのうち、交差処理が実行されたe個の個体データと、交差処理が実行されていないp-e個の個体データを合わせた計p個の個体データが全て変異処理されたp個の個体データを、新たに第g世代子個体集団データとする。
子個体集団データ取得部103は、第g世代親個体集団データに含まれるp個の個体データのうち、交差処理が実行されたe個の個体データと、交差処理が実行されていないp-e個の個体データを合わせた計p個の個体データが全て変異処理されたp個の個体データを、新たに第g世代子個体集団データとする。
次に、S24において、処理部10は、第g世代親個体集団データ及び第g世代子個体集団データを統合し、第g世代統合データを生成する。これにより、第g世代統合データには2p個の個体データが含まれることとなる。
次に、S25において、非優越ソート実行部106は、予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、第g世代統合データに対して非優越ソートを実行する。そして、2p個の個体データをパレート最適解におけるフロント毎(ランク毎)に分類する。
次に、S26において、個体選択部107は、パレート最適解におけるフロント毎(ランク毎)に分類された第g世代統合データから、ランクの高い順に定められた数の個体データを選択する。なお、個体選択部107は、定められた数の個体データを選択するときに、ランクが同じ個体データが存在する場合には、混雑距離が大きいものから順に選択することとしてもよい。ここで、予め定められた数として、pを採用することができる。そして、親個体集団データ取得部102は、個体選択部107により選択されたp個の個体データを、第g+1世代の親個体集団を表す第g+1世代親個体集団データとして生成し、取得する。以下、図13及び図14を用いて、かかる評価基準について説明する。
<評価基準>
本実施形態では、非優越ソートを実行した2p個の個体データからp個の個体データを選択するに際し、2つの観点の評価基準を利用する。かかる観点は、相同組み換えを抑制し、目的タンパク質の生産量を高めることを目的として導き出された観点である。1つ目の観点は、「コドン適合度」に関するものである。具体的には、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする。これが第1評価基準である。そして、2つ目の観点は、「コドンの塩基配列」に関するものである。さらに、2つ目の観点は、「不一致塩基数」及び「最長共通文字列」に分かれる。そして、個体データに含まれる2つのCDSのうち、不一致塩基数の最小値を基準とするのが第2評価基準である。また、個体データに含まれるCDSのうち、最長共通文字列の長さを基準とするのが第3評価基準である。以下、これら3つの評価基準の意義について、それぞれ説明する。
本実施形態では、非優越ソートを実行した2p個の個体データからp個の個体データを選択するに際し、2つの観点の評価基準を利用する。かかる観点は、相同組み換えを抑制し、目的タンパク質の生産量を高めることを目的として導き出された観点である。1つ目の観点は、「コドン適合度」に関するものである。具体的には、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする。これが第1評価基準である。そして、2つ目の観点は、「コドンの塩基配列」に関するものである。さらに、2つ目の観点は、「不一致塩基数」及び「最長共通文字列」に分かれる。そして、個体データに含まれる2つのCDSのうち、不一致塩基数の最小値を基準とするのが第2評価基準である。また、個体データに含まれるCDSのうち、最長共通文字列の長さを基準とするのが第3評価基準である。以下、これら3つの評価基準の意義について、それぞれ説明する。
(第1評価基準:コドン適合度)
第1の観点である第1評価基準は、「コドン適合度」に関するものである。ここで、「コドン適合度」とは、個体データに含まれるCDS中に利用頻度の高いコドンが多く含まれているほど高くなるものとする。具体的には、各個体データに含まれるCDSのコドン適合インデックス(Codon Adaptation Index(以下、CAIという))の最小値(以下、最小CAI値という)を基準とする。CAIは、例えば以下の式で求めることができる。
第1の観点である第1評価基準は、「コドン適合度」に関するものである。ここで、「コドン適合度」とは、個体データに含まれるCDS中に利用頻度の高いコドンが多く含まれているほど高くなるものとする。具体的には、各個体データに含まれるCDSのコドン適合インデックス(Codon Adaptation Index(以下、CAIという))の最小値(以下、最小CAI値という)を基準とする。CAIは、例えば以下の式で求めることができる。
L:個体データに含まれるコドンの数
fi:i番目のコドンの頻出度
max(fj):最頻出である同義コドン(j番目のコドン)の頻出度
ここで、同義コドンとは、同じアミノ酸をコードするコドンであって、異なる配列を持ったコドンのことである。
そして、上記の式で求めたCAIを用いて、以下の式で最小CAI値を求めることができる。
ここで、あるCDSのCAIが高いほど、そのCDSには利用頻度の高いコドンが多く含まれている(逆に言うとCDSに含まれるレアコドンの数が少ない)ことを示す。そして、ある個体データが、CAIが極端に低い(換言すると、レアコドンが多く含まれた)CDSを持っていると、そのCDSは効率的に翻訳されない可能性がある。したがって、最小CAI値を第1評価基準として用い、最小CAI値が大きいほど、かかる個体データの評価を高くすることにより、CAIが極端に低いCDSを持つ個体データを最適化の過程で取り除くことが可能になる。したがって、第1変異処理により、最小CAI値を大きくすることで、より好ましいシミュレーション結果を得ることができる。
(第2評価基準:不一致塩基数)
次に、図13を用いて第2評価基準について説明する。第2の観点のうちの1つ目である第2評価基準は、「不一致塩基数」に関するものである。具体的には、不一致塩基数の最小値(以下、最小不一致塩基数という)を評価基準に用いる。ここで、不一致塩基とは、個体データに含まれるx個のCDSのうち、2つのCDS(以下、CDSペアという)Ci及びCjを対比して、コドンを構成する塩基が不一致となる塩基のことである。図13の例では、Ci及びCjを構成する塩基のうち、不一致塩基の数が5個となっている。したがって、かかるCDSペア(Ci及びCj)の不一致塩基数は5となる。最小不一致塩基数は、以下の式で求めることができる。
次に、図13を用いて第2評価基準について説明する。第2の観点のうちの1つ目である第2評価基準は、「不一致塩基数」に関するものである。具体的には、不一致塩基数の最小値(以下、最小不一致塩基数という)を評価基準に用いる。ここで、不一致塩基とは、個体データに含まれるx個のCDSのうち、2つのCDS(以下、CDSペアという)Ci及びCjを対比して、コドンを構成する塩基が不一致となる塩基のことである。図13の例では、Ci及びCjを構成する塩基のうち、不一致塩基の数が5個となっている。したがって、かかるCDSペア(Ci及びCj)の不一致塩基数は5となる。最小不一致塩基数は、以下の式で求めることができる。
ここで、ある個体データが、不一致塩基数が極端に低い(換言すると、塩基配列がよく似た)CDSペアを持っていると、そのCDSペアの間で相同組み換えが生じる可能性が高くなる。これは、相同組み換えは、塩基配列がよく似た部位(相同部位)で生じるためである。したがって、最小不一致塩基数を第2評価基準として用い、最小不一致塩基数が大きい(換言すると、塩基配列が異なる割合が大きい)ほど、かかる個体データの評価を高くすることにより、塩基配列がよく似た個体データを最適化の過程で取り除くことが可能になる。したがって、第2変異処理により、不一致塩基数を大きくすることで、より好ましいシミュレーション結果を得ることができる。
(第3評価基準:最長共通文字列)
次に、図14を用いて第3評価基準について説明する。第2の観点のうちの2つ目である第3評価基準は、「最長共通文字列」に関するものである。すでに述べたように、「最長共通文字列」とは、各CDS又は1つのCDS内部の異なる部位に含まれるコドンを表す文字列を、他のCDSに含まれる文字列と対比して、連続して一致する文字列の中で最も長いもののことである。
次に、図14を用いて第3評価基準について説明する。第2の観点のうちの2つ目である第3評価基準は、「最長共通文字列」に関するものである。すでに述べたように、「最長共通文字列」とは、各CDS又は1つのCDS内部の異なる部位に含まれるコドンを表す文字列を、他のCDSに含まれる文字列と対比して、連続して一致する文字列の中で最も長いもののことである。
ここで、「全く同じ塩基配列」がゲノム近傍にあると、相同組み換えが生じる可能性が高くなる。これは、前述の通り、相同組み換えは、塩基配列がよく似た部位(相同部位)で生じるためである。したがって、「最長共通文字列」の長さを第3評価基準として用い、最長共通文字列の長さが短いほど、かかる個体データの評価を高くすることにより、「全く同じ塩基配列」が高い割合で含まれる個体データを最適化の過程で取り除くことが可能になる。したがって、第2変異処理により、最長共通文字列を小さくすることで、より好ましいシミュレーション結果を得ることができる。
以上説明したように、第1の観点である第1評価基準を用いることにより、利用頻度の高いコドンが多く含まれるCDSを有する個体データを選択することが可能となる。また、第2の観点である第2評価基準及び第3評価基準を用いることにより、塩基配列が異なる割合が大きい個体データを選択し、相同組み換えの発生を抑制することが可能となる。
図8に戻り、メインルーチンについてさらに説明する。S26において、個体選択部107は、パレート最適解におけるフロント毎(ランク毎)に分類された(2p個の個体データを含む)第g世代統合データから、ランクの高い順にp個の個体データを選択する。そして、親個体集団データ取得部102は、選択されたp個の個体データを新たに第g+1世代の親個体集団データとし、第g+1世代親個体集団データを生成する。
次に、S27において、処理部10は、変数gが予め定められた世代数Gを超えるか否かを判定する。そして、かかる判定結果がNOであれば、S28に進む。一方、S27における判定結果がYESであれば、メインルーチンを終了し、かかる計算結果を計算データ記憶部204へ出力する。
S27における判定結果がNOであれば、S28に進み、変数gをインクリメントし(つまり、変数gに1を加え)、再びS21に戻る。ここで、現時点では変数g=2であるので、親個体集団データ取得部102は、S26において生成された第2世代親個体集団データを取得する。かかる処理を、変数gが予め定められた世代数Gとなるまで繰り返し実行する。換言すると、S21~S26における処理を250回繰り返し実行する。
以上説明したメインルーチンを繰り返し実行することにより、3つの評価基準に基いて選択されたp個の個体データは、繰り返し回数が増えるほど、遺伝子配列群として好ましいものとなっていく。
<実施例>
以下、本アルゴリズムを用いた遺伝子配列設計につき、実施例について説明する。かかる実施例では、シミュレーションとして、ヒトのインスリンA鎖(アミノ酸配列:GIVEQCCTSICSLYQLENYCN)をコードする10個のCDSを設計した。種々のパラメータについては、以下の通りである。
予め定められた確率Pm(変異率)=0.05
Pc(交差率)=0.5
第g世の個体集団データ(親個体集団データ、子個体集団データ)に含まれる個体データの数p=100
予め定められた世代数G(最大世代数)=250
以下、本アルゴリズムを用いた遺伝子配列設計につき、実施例について説明する。かかる実施例では、シミュレーションとして、ヒトのインスリンA鎖(アミノ酸配列:GIVEQCCTSICSLYQLENYCN)をコードする10個のCDSを設計した。種々のパラメータについては、以下の通りである。
予め定められた確率Pm(変異率)=0.05
Pc(交差率)=0.5
第g世の個体集団データ(親個体集団データ、子個体集団データ)に含まれる個体データの数p=100
予め定められた世代数G(最大世代数)=250
以下、図15及び図16を用いて、本シミュレーションにおける計算結果について、第1評価基準を横軸に、第2評価基準を縦軸にとってプロットしたグラフと、第1評価基準を横軸に、第3評価基準を縦軸にとってプロットしたグラフについて説明する。
図15は、第1評価基準を横軸に、第2評価基準を縦軸にとってプロットしたグラフである。ここで、グラフ中にて丸で表されるプロットは第1世代、四角形で表されるプロットは第10世代、三角形で表されるプロットは第250世代における計算結果を示す。なお、1つのプロットは1つの設計結果(=個体データ)に対応する。すでに述べたように、第1評価基準は最小CAI値が大きいほど評価が高いので、グラフ中では横軸の右側にプロットされた点ほど評価が良く、横軸の左側にプロットされた点ほど評価が悪いといえる。また、第2評価基準は、最小不一致塩基数が大きいほど評価が高いので、グラフ中では縦軸の上側にプロットされた点ほど評価が良く、縦軸の下側にプロットされた点ほど評価が悪いといえる。図15に示されるように、世代数が大きくなるにしたがって(換言すると、図8におけるメインルーチンの繰り返し回数が増えるにしたがって)、個体集団データ全体として好ましいものとなっていることが読み取れる。
図16は、第1評価基準を横軸に、第3評価基準を縦軸にとってプロットしたグラフである。丸、四角形及び三角形で表される各プロットの意味は、図15と同様である。ここで、第3評価基準は、最長共通文字列の長さが短いほど評価が高いので、グラフ中では縦軸の下側にプロットされた点ほど評価が良く、縦軸の上側にプロットされた点ほど評価が悪いといえる。図16に示されるように、世代数が大きくなるにしたがって(換言すると、図8におけるメインルーチンの繰り返し回数が増えるにしたがって)、個体集団データ全体として好ましいものとなっていることが読み取れる。
以上、種々の実施形態について説明したが、本発明はこれらに限定されない。
例えば、図8におけるメインルーチンのS26における選択は、第1評価基準及び第2評価基準、又は、第1評価基準及び第3評価基準のいずれか一方を用い、図15及び図16に示されるグラフの一方を得ることとしてもよい。また、第1評価基準及び第2評価基準、及び、第1評価基準及び第3評価基準の両方を用いる場合は、世代毎に図15におけるグラフと図16におけるグラフからそれぞれ評価の高い個体データを特定し、任意の基準でポイントを付与し、これら2つのグラフにおけるポイントの合計が高い個体データを選択してもよい。もしくは、図15及び図16のように2次元のグラフではなく、第1評価基準をx軸に、第2評価基準をy軸に、第3評価基準をz軸にして、3次元のグラフを作成することにより3つの評価基準のそれぞれについて高い評価を得た個体データを選択してもよい。
また、記憶部20は、情報処理装置1の内部に設けずに、外部のPC又はサーバ等の情報処理装置に設けるクラウドコンピューティングの態様とすることができる。この場合、計算の度に必要なデータを外部の情報処理装置が情報処理装置1に送信する。
また、情報処理装置1の機能を実装したASIC(application specific integrated circuit)、FPGA(field-programmable gate array)、DRP(Dynamic ReConfigurable Processor)として提供することもできる。また、コンピュータに、情報処理装置1の機能を実現するためのプログラムとして提供することもできる。この場合、かかるプログラムをインターネット等を介して配信することもできる。
さらに、本アルゴリズムとして、多目的遺伝的アルゴリズムである「NSGA-II」を利用することもできる。これは、本アルゴリズムと同様に、p個の最適解をまとめて得ることができるためである。また、組み合わせ最適化アルゴリズムの一種である「シミュレーテッドアニーリング」や「(単目的の)遺伝的アルゴリズム」を利用してもよい。ただし、この場合には、p個の最適解をまとめて得ることができないので、計算を少なくともp回以上繰り返し、p個の最適解を得る必要がある。さらに、これら2つ以上のアルゴリズムの計算結果を混合してもよい。この場合、p以下の任意の整数αを設定し、あるアルゴリズムによる計算結果からα個の個体を選択し、他のアルゴリズムによる計算結果からp-α個の個体を選択し、これらを結合したp個の個体を用いることとしてもよい。
さらに、本発明は、
アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得ステップと、
前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理ステップと、
前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得ステップと、
予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行ステップと、
前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択ステップと、
を有する遺伝子配列設計方法
として捉えることもできる。
アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得ステップと、
前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理ステップと、
前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得ステップと、
予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行ステップと、
前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択ステップと、
を有する遺伝子配列設計方法
として捉えることもできる。
1:情報処理装置、10:処理部、20:記憶部、30:操作部、40:表示部、50:通信部、100:バス、101:個体生成部、102:親個体集団データ取得部、103:子個体集団データ取得部、104:交差処理部、105:変異処理部、106:非優越ソート実行部、107:個体選択部、201:アミノ酸配列データ記憶部、202:遺伝子数データ記憶部、203:コドン頻度表データ記憶部、204:計算データ記憶部、205:評価基準記憶部205
Claims (16)
- アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部と、
前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部と、
前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部と、
予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部と、
前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部と、
を有する情報処理装置。 - 前記個体選択部は、前記予め定められた数の前記個体データを選択するときに、前記ランクが同じ前記個体データが存在する場合には、混雑距離が大きいものから順に選択する、
請求項1に記載の情報処理装置。 - 前記親個体集団データ取得部は、前記個体選択部により選択された前記個体データを、第2世代の親個体集団を表す第2世代親個体集団データとし、
前記変異処理部、前記非優越ソート実行部及び前記個体選択部による処理を、予め定められた世代数となるまで実行する、
請求項1又は請求項2に記載の情報処理装置。 - 前記コドン適合度に関する評価基準は、各個体が複数有する塩基配列であって、アミノ酸翻訳の対象となる塩基配列を表すCDSのコドン適合インデックスの最小値を基準とする、
請求項1~請求項3のいずれか1項に記載の情報処理装置。 - 前記個体に含まれる前記コドン適合インデックスの最小値が大きいほど、前記個体の評価を高くする、
請求項4に記載の情報処理装置。 - 前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる2つの前記CDSのうち、互いに一致しない塩基の数を表す不一致塩基数の最小値を基準とする、
請求項1~請求項5のいずれか1項に記載の情報処理装置。 - 前記不一致塩基数の最小値が大きいほど、前記個体の評価を高くする、
請求項6に記載の情報処理装置。 - 前記コドンの塩基配列に関する評価基準は、前記各個体に含まれる前記CDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字列の長さを基準とする、
請求項1~請求項7のいずれか1項に記載の情報処理装置。 - 前記最長共通文字列の長さが短いほど、前記個体を高く評価する、
請求項8に記載の情報処理装置。 - 前記変異処理部は、
第g世代の親個体集団を表す第g世代親個体集団データに含まれる各個体データに対し、第1変異処理及び前記第1変異処理とは異なる第2変異処理を実行する、
請求項1~請求項9のいずれか1項に記載の情報処理装置。 - 前記変異処理部は、
前記各個体に含まれる全てのCDSに対し、前記CDSに含まれる前記コドンを、予め定められた確率で前記コドンより高頻度のコドンに置換する第1変異処理を実行する、
請求項10に記載の情報処理装置。 - 前記変異処理部は、
前記各個体に含まれるCDSのうち、それぞれのCDS間又は1つのCDS内部の異なる部位で連続して一致する塩基配列のうち最長の塩基配列である最長共通文字と重なる前記コドンを、予め定められた確率で他のコドンに置換する第2変異処理を実行する、
請求項10又は請求項11に記載の情報処理装置。 - 前記第1変異処理又は前記第2変異処理は、ランダムに選択される、
請求項10~請求項12のいずれか1項に記載の情報処理装置。 - 前記第1世代親個体集団データに含まれる個体に対し、交差処理を実行する交差処理部を有し、
前記交差処理は、
第g世代の親個体集団を表す第g世代親個体集団データから予め定められた偶数個の個体データを抽出し、前記抽出された個体データから2個の個体データを選択し、前記選択された2個の個体データに対して交差処理を実行する、
請求項1~請求項13のいずれか1項に記載の情報処理装置。 - 前記交差処理部は、
前記選択された2個の個体データである第1個体データ及び第2個体データに含まれる前記CDSに含まれる前記コドンの境界から交差ポイントを決定し、
前記交差ポイントを境として前記第1個体データと前記第2個体データに含まれる前記コドンを入れ替える、
請求項14に記載の情報処理装置。 - コンピュータを、
アミノ酸配列、遺伝子数及びコドン頻度表を表すデータに基いて生成されたデータであって、予め定められた数の個体データを含む第1世代の親個体集団を表す第1世代親個体集団データを取得する親個体集団データ取得部、
前記第1世代親個体集団データに含まれる個体に対し、変異処理を実行する変異処理部、
前記変異処理が実行された個体を含む第1世代の子個体集団を表す第1世代子個体集団データを取得する子個体集団データ取得部、
予め定められた評価基準であって、コドン適合度及び前記コドンの塩基配列に関する評価基準に基いて、前記第1世代親個体集団データ及び前記第1世代子個体集団データを統合した第1世代統合データに対して非優越ソート処理を実行し、前記第1世代統合データに含まれる全個体データをパレート最適解におけるランク毎に分類する非優越ソート実行部、
前記ランク毎に分類された全個体データから、前記ランクの高い順に予め定められた数の前記個体データを選択する個体選択部、
として機能させるための情報処理プログラム。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2016-070976 | 2016-03-31 | ||
| JP2016070976A JP2019095819A (ja) | 2016-03-31 | 2016-03-31 | 情報処理装置及びプログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2017169736A1 true WO2017169736A1 (ja) | 2017-10-05 |
Family
ID=59964367
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2017/010169 Ceased WO2017169736A1 (ja) | 2016-03-31 | 2017-03-14 | 情報処理装置及びプログラム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP2019095819A (ja) |
| WO (1) | WO2017169736A1 (ja) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2021532439A (ja) * | 2018-07-30 | 2021-11-25 | ナンジン、ジェンスクリプト、バイオテック、カンパニー、リミテッドNanjing Genscript Biotech Co., Ltd. | コドン最適化 |
| CN116307296A (zh) * | 2023-05-22 | 2023-06-23 | 南京航空航天大学 | 云上资源优化配置方法 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0638772A (ja) * | 1991-09-30 | 1994-02-15 | P C C Technol:Kk | 植物における外来遺伝子の発現方法 |
| JP2007172306A (ja) * | 2005-12-22 | 2007-07-05 | Yamaha Motor Co Ltd | 多目的最適化装置、多目的最適化方法および多目的最適化プログラム |
-
2016
- 2016-03-31 JP JP2016070976A patent/JP2019095819A/ja active Pending
-
2017
- 2017-03-14 WO PCT/JP2017/010169 patent/WO2017169736A1/ja not_active Ceased
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0638772A (ja) * | 1991-09-30 | 1994-02-15 | P C C Technol:Kk | 植物における外来遺伝子の発現方法 |
| JP2007172306A (ja) * | 2005-12-22 | 2007-07-05 | Yamaha Motor Co Ltd | 多目的最適化装置、多目的最適化方法および多目的最適化プログラム |
Non-Patent Citations (1)
| Title |
|---|
| LUYI WANG: "Tamokuteki Identeki Algorithm ni Okeru Kai no Seido to Habahirosa no Kojo no Kento", SHUSHI RONBUN, 23 January 2010 (2010-01-23), pages 1 - 24, Retrieved from the Internet <URL:http://www.is.doshisha.ac.jp/academic/papers/pdf/09/2009mthesis/20091uy.pdf> [retrieved on 20170523] * |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2021532439A (ja) * | 2018-07-30 | 2021-11-25 | ナンジン、ジェンスクリプト、バイオテック、カンパニー、リミテッドNanjing Genscript Biotech Co., Ltd. | コドン最適化 |
| JP7542443B2 (ja) | 2018-07-30 | 2024-08-30 | ナンジン ジェンスクリプト バイオテック カンパニー,リミテッド | コドン最適化 |
| CN116307296A (zh) * | 2023-05-22 | 2023-06-23 | 南京航空航天大学 | 云上资源优化配置方法 |
| CN116307296B (zh) * | 2023-05-22 | 2023-09-29 | 南京航空航天大学 | 云上资源优化配置方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2019095819A (ja) | 2019-06-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110136773A (zh) | 一种基于深度学习的植物蛋白质互作网络构建方法 | |
| Bidlo | On routine evolution of complex cellular automata | |
| Poladian et al. | Multi-objective evolutionary algorithms and phylogenetic inference with multiple data sets | |
| CN106096327A (zh) | 基于Torch监督式深度学习的基因性状识别方法 | |
| Zhang et al. | Compression of deep neural networks: bridging the gap between conventional-based pruning and evolutionary approach | |
| Dondi et al. | Orthology correction for gene tree reconstruction: Theoretical and experimental results | |
| CN110990353B (zh) | 日志提取方法、日志提取装置及存储介质 | |
| Bruneau et al. | A clustering package for nucleotide sequences using Laplacian Eigenmaps and Gaussian Mixture Model | |
| Nayeem et al. | Multiobjective formulation of multiple sequence alignment for phylogeny inference | |
| Feng et al. | Artificial intelligence in bioinformatics: Automated methodology development for protein residue contact map prediction | |
| CN113611354A (zh) | 一种基于轻量级深度卷积网络的蛋白质扭转角预测方法 | |
| Lin et al. | An efficient hybrid Taguchi-genetic algorithm for protein folding simulation | |
| CN110955702B (zh) | 一种基于改进遗传算法的模式数据挖掘方法 | |
| Kawamura et al. | A hybrid approach for optimal feature subset selection with evolutionary algorithms | |
| Daniels et al. | MRFy: remote homology detection for beta-structural proteins using Markov random fields and stochastic search | |
| JP2019095819A (ja) | 情報処理装置及びプログラム | |
| Du et al. | Genetic algorithms | |
| Ray et al. | Disease associated protein complex detection: a multi-objective evolutionary approach | |
| Khanum et al. | Reflected adaptive differential evolution with two external archives for large-scale global optimization | |
| CN113297293A (zh) | 一种基于约束优化进化算法的自动化特征工程方法 | |
| Kaya et al. | A novel multi-objective genetic algorithm for multiple sequence alignment | |
| Xu et al. | Molecular de novo design through transformer-based reinforcement learning | |
| Hadjiivanov et al. | Epigenetic evolution of deep convolutional models | |
| Yang et al. | Optimization of classification algorithm based on gene expression programming | |
| Prousalis et al. | A Survey on Sequence Alignment Algorithms and State-of-the-Art Aligners |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 17774280 Country of ref document: EP Kind code of ref document: A1 |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 17774280 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: JP |


