EP4558990A1 - Verfahren zur optimierung einer nukleotidsequenz durch austausch synonymer codons für die expression einer aminosäuresequenz in einem zielorganismus - Google Patents
Verfahren zur optimierung einer nukleotidsequenz durch austausch synonymer codons für die expression einer aminosäuresequenz in einem zielorganismusInfo
- Publication number
- EP4558990A1 EP4558990A1 EP23750552.4A EP23750552A EP4558990A1 EP 4558990 A1 EP4558990 A1 EP 4558990A1 EP 23750552 A EP23750552 A EP 23750552A EP 4558990 A1 EP4558990 A1 EP 4558990A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- codon
- tuple
- amino acid
- sequence
- nucleotide sequence
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
- C12N15/67—General methods for enhancing the expression
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
- G16B25/10—Gene or protein expression profiling; Expression-ratio estimation or normalisation
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B50/00—ICT programming tools or database systems specially adapted for bioinformatics
Definitions
- the present invention lies in the field of producing synthetic nucleotide sequences and their use for producing proteins by introducing these nucleotide sequences into an expression system with a suitable host organism which expresses the protein encoded by the nucleotide sequence.
- the present invention relates in particular to methods in which a nucleotide sequence is optimized for expression in a predetermined host organism.
- An “expression system” is a biological system that is capable of carrying out protein biosynthesis in a targeted and controlled manner, i.e. producing, i.e. “expressing”, certain proteins based on the template of a nucleotide sequence.
- Heterologous expression is understood to mean the expression of a gene or part of it in a host organism that does not naturally possess this gene or gene fragment.
- the corresponding nucleotide sequences are created using genetic technology or recombinant DNA technology, for example with the help of of vectors or genome editing, into the host organism, whereupon it is multiplied and stimulated to overproduce the protein.
- a “homologous expression” therefore refers to the expression of a gene or part of it in a host organism or a system from which it originally comes.
- Heterologous protein expression can be carried out in many types of host organisms.
- the host organism can e.g. B. be a bacterium, a fungus, a yeast, an insect cell, a mammalian cell or a plant cell.
- a frequently occurring problem in heterologous protein expression is a low transcription and translation rate of a foreign nucleotide sequence in a specific host organism, also referred to here as the target organism.
- the cause is u. a. the degeneration of the genetic code, which leads to the fact that for most of the amino acids to be incorporated in translation, several codons with the same meaning, also referred to here as “synonymous codons”, are available.
- a codon is a sequence of three consecutive nucleobases of a nucleic acid sequence, i.e. a “base triplet” that represents one Amino acid can encode. There are a total of 64 possible codons, of which 61 code for the 20 canonical proteinogenic amino acids, and three more code for stop codons.
- codons in the genes of an organism are not arranged randomly, but rather that the observed frequencies of codon pairs can, contrary to expectations, deviate from the product of the respective individual frequencies and are therefore statistically “underrepresented” or “overrepresented”. This context is also referred to as the “codon context,” which can have an additional influence on translation efficiency.
- codon optimization by exchanging individual codons of the nucleotide sequence to be expressed for synonymous codons with higher frequencies in the target organism (also referred to as “codon optimization” or “codon usage optimization”), as well as through transmission the frequencies of codons of highly expressed genes in the target organism on the target protein (also referred to as “codon adaptation index”).
- codon context optimization whereby the codons of the nucleic acid sequence to be expressed are adapted to the codon context by replacing them with overrepresented or underrepresented synonymous codons of the target organism can be adapted without changing the encoded amino acid sequence.
- WO 2020/024917 Al discloses a computer-implemented one
- nucleic acid sequences etc. a. on the basis of Codon Adaptation Index and codon context can be optimized for the expression of a protein in a host using a computer-aided NSGA-II algorithm.
- a computer-aided method in which a predetermined nucleic acid sequence is optimized for expression in a predetermined target organism using a quality function.
- the quality function can take into account, among other things, the codon use and the codon context as a quality criterion.
- WO 2008/000632 A1 a further method is proposed in which new coding sequences are generated from a predetermined nucleotide sequence, which encodes a predetermined amino acid sequence, by exchanging one or more synonymous codons in several repetition steps and based on a fitness value, which, among other things. a. taking the codon context of the host organism into account.
- WO 2018/104385 A1 a method for determining an optimized nucleotide sequence, which encodes a predetermined amino acid sequence and is optimized for expression in a specific target organism, is known, wherein a large number of candidate nucleotide sequences are generated and using a statistical machine learning algorithm be rated .
- the subject of the present invention is a method for optimizing a nucleotide sequence for the expression of a predetermined amino acid sequence in at least one target organism.
- the expression can in principle be heterologous or homologous. It is preferably a heterologous expression.
- the nucleotide sequence comprises a large number of base triplets, with at least one change position in the nucleotide sequence being a base triplet that contains an amino acid of the predetermined one Amino acid sequence encoded is replaced by a synonymous base triplet that encodes the same amino acid of the predetermined amino acid sequence in order to optimize the nucleotide sequence for expression in the at least one target organism.
- a change position according to the invention here comprises a direct sequence of n base triplets, which forms a first codon n-tuple and encodes a sequence section of n amino acids of the predetermined amino acid sequence, which forms an amino acid n-tuple, the amino acid n-tuple having a predetermined amount of amino acid n-tuple events in the genome of the at least one target organism or part thereof and / or in genomes of viruses or parts thereof capable of infecting the at least one target organism.
- the method according to the invention includes that at least one of the n base triplets from the direct succession of the at least one change position is replaced by a synonymous base triplet, the synonymous base triplet being chosen so that a second codon n-tuple results, which in relation to the quantity the amino acid n-tuple events have a higher relative codon n-tuple frequency in the genome or the part thereof of the at least one target organism and / or in the genomes or the parts thereof of the viruses capable of infecting the at least one target organism as the first codon n-tuple.
- n is a natural number greater than or equal to two and in particular less than or equal to a total number N of amino acids of the predetermined amino acid sequence.
- the invention is based on the knowledge that the influence of a direct sequence of n base triplets, referred to here as a codon n-tuple, the translation efficiency of a sequence of n amino acids encoded by the n base triplets, here referred to as an amino acid n-tuple, in a specific target organism can be expressed by the relative frequency with which the direct sequence of n base triplets encodes the direct sequence of n amino acids within the genome of a host organism, referred to here as relative codon n-tuple frequency.
- the inventors have recognized that the ratio of the absolute frequency of a codon n-tuple and the absolute frequency of the corresponding amino acid n-tuple, which is encoded by the codon n-tuple, is an advantageous measure to quantify the suitability of a given nucleic acid sequence for expression in a specific target organism.
- the relative frequency can have values between 0 and 1 or Assume 0% and 100%, where the relative frequency is equal to 0, if a particular codon n-tuple in the genome or part of it of the target organism is not used at all to encode the corresponding amino acid n-tuple.
- the relative frequency is equal to 1 or .
- any codon n-tuple can be selected which encodes this amino acid n-tuple.
- the associated synonymous codon n-tuples can each be assigned a uniform relative frequency of 1/i, where i is the number of synonymous codon n-tuples that encode the amino acid n-tuple.
- codon n-tuples encoding this amino acid n-tuple can be assigned a relative frequency of 0. Another possibility is to use this change position or to exclude this amino acid n-tuple from the optimization.
- the method according to the invention can also enable the biosynthesis of proteins that previously could not or hardly be expressed heterologously. In this way, the method according to the invention leads to an improvement in efficiency and sustainability in biotechnological Protein production for scientific, medical and technical or industrial purposes.
- n is less than or equal to 50, less than or equal to 40, less than or equal to 30, less than or equal to 20 or less than or equal to 10.
- n 2.
- n 3.
- n 6.
- n is greater than or equal to three.
- predetermined means that the set of amino acid n-tuple events is determined by the genome or the proteome of the at least one target organism or a portion thereof or the genomes of the at least one viruses or parts thereof capable of targeting a target organism. It is therefore understood that the determination of the number of events, hereinafter also referred to as the absolute frequency, with which a specific amino acid n-tuple occurs in the genome of the at least one target organism or a part or in the genomes of viruses or parts thereof capable of infecting the at least one target organism represents a step that can be carried out during the process, but does not have to be carried out.
- the information about the absolute frequency with which an amino acid n -Tuple is encoded in the genome of the at least one target organism or part thereof and/or in genomes of viruses or parts thereof capable of infecting the at least one target organism another way, e.g. B. from databases or the like, can be included in the method according to the invention.
- the determination of the set of events with which the amino acid n-tuple in the genome of at least one target organism or a part thereof or is encoded in genomes of viruses or parts thereof capable of infecting the at least one target organism is carried out as a method step.
- the same basically applies to determining the number of events, i.e. H . the absolute frequency with which a specific codon n-tuple occurs in the genome of at least one target organism or a part thereof or occurs in genomes of viruses or parts thereof capable of infecting the at least one target organism, and/or for the resulting relative frequency of the codon n-tuple according to the invention.
- the absolute frequency and/or the relative frequency of essentially every combinatorially possible codon n-tuple in the genome of the at least one target organism or a part thereof or in genomes of viruses or parts thereof capable of infecting the at least one target organism are deposited in a database and are included in the method according to the invention in the form of database information.
- the synonymous base triplet is selected in the method according to the invention specifically from the point of view that the second codon n-tuple fulfills the condition required according to the invention of a higher relative codon n-tuple frequency compared to the first codon n-tuple.
- Selecting the synonymous base triplet can therefore in particular include determining and/or evaluating the relative codon n-tuple frequency of the second codon n-tuple.
- the term “determine” can as stated above, e.g. B. in the form of a calculation of the relative codon n-tuple frequency of the second codon n-tuple or in the form of a data comparison, the inclusion of database information or the like.
- a preferred method implementation therefore includes at least the following steps: a) determining the at least one change position; b) replacing the at least one base triplet of the at least one change position with the synonymous base triplet; c) Determine the relative codon n-tuple frequency of the resulting second codon n-tuple.
- the method further comprises a step d) evaluating the relative codon n-tuple frequency of the second codon n-tuple determined in step c), the evaluation z.
- a comparison with the relative codon n-tuple frequency of the first codon n-tuple can be carried out and/or based on a target criterion such as a minimum value or the like.
- Steps b) and c) and if necessary. d) can also be repeated until a second codon n-tuple results, which has the higher relative codon n-tuple frequency required according to the invention than the first codon n-tuple.
- the determined set of amino acid n-tuple events or The determined relative frequency of the codon n-tuple does not necessarily have to be based on the complete genome of the at least one target organism or the complete genomes of the viruses capable of infecting the at least one target organism. Rather, in certain process variants it can be sufficient and advantageous if only parts of the genome or of the genomes from the determined absolute amino acid n-tuple frequency or the relative codon n-tuple frequency are included, especially since the Most genomes each contain a large portion of non-coding regions, which are less relevant for the inventive measurement of the suitability of a given nucleic acid sequence for expression in the target organism.
- the set of amino acid n-tuple events results from several, preferably all, protein-coding genes and/or proteins of the at least one target organism or the viruses or viruses capable of infecting at least one target organism. is based on several protein-coding genes and/or proteins of at least one target organism or the viruses capable of infecting at least one target organism are determined. Proteins constitutively expressed by the target organism or proteins with high transient expression or high abundance are particularly suitable for this. Particularly preferably, at least 25%, at least 50%, at least 75%, at least 80%, at least 90% or at least 95% of the coding part of the genome of the target organism is included in the determination of the amount of amino acid n-tuple events.
- the relative codon n-tuple frequency also results from a set of events of the first or second codon n-tuple in several, preferably all, protein-coding genes and / or proteins of the at least one target organism or the viruses or viruses capable of infecting at least one target organism. is based on several protein-coding genes and/or proteins of at least one target organism or of viruses capable of infecting at least one target organism is determined, which is based on the amount of amino acid n-tuple events. Particularly preferred are at least 25%, at least 50%, at least 75%, at least 80%, at least 90% or at least 95% of the coding part of the genome of the target organism is covered by the determination of the relative codon n-tuple frequency.
- the relative frequency of the first or second codon n-tuple is a value arithmetically calculated from absolute frequencies.
- the relative frequency of the first and/or second codon n-tuple can also be determined by a different probability distribution or another probability measure can be expressed or replaced, e.g. B. as an interval estimate. That's how it is, for example. B. possible, from chance observations, e.g. B.
- the relative codon n-tuple frequency is then z. B. represented by the interval center or .
- the relative codon n-tuple frequency can also be calculated using other values from the interval. be represented. In addition to the interval center, e.g. B. the smallest value of the Interval, especially for very conservative estimates, or the average value of a weight function (e.g. -1/x) on the interval represents the relative codon n-tuple frequency.
- the at least one change position comprises a direct sequence of n base triplets, which forms a first codon n-tuple and encodes a sequence section of n amino acids of the predetermined amino acid sequence, which forms an amino acid n-tuple , whereby at least one of the n base triplets of the direct sequence is replaced by a synonymous base triplet, which is selected using an estimating function in such a way that a second codon n-tuple results, which has the amino acid n-tuple with a greater probability in the genome or a part thereof of the at least one target organism and/or in genomes or parts thereof of viruses capable of infecting the at least one target organism coded as the first codon n-tuple.
- Suitable estimation functions that can be implemented for the method according to the invention are known to those skilled in the art.
- “at least one change position” includes the possibility that the method includes several or a plurality of change positions, in each of the change positions at least one of the n base triplets of the direct succession through a synonymous base triplet is replaced and the synonymous base triplets are chosen so that at least some of the resulting second codon n-tuples have a higher relative codon n-tuple frequency than the respective first codon n-tuples.
- the change positions can for this purpose in one or more of the change positions also several or all of the n base triplets are each replaced by a synonymous base triplet.
- At least in the change position with the first codon n-tuple, which has the lowest relative codon n-tuple frequency of all change positions at least one of the base triplets of the direct sequence is replaced by a synonymous base triplet, which is chosen so that the resulting second codon n-tuple has a higher relative codon n-tuple frequency than the first codon n-tuple.
- the inventors have recognized that the codon n-tuple with the lowest relative frequency in the nucleotide sequence regularly has a limiting effect on the entire translation process, so that replacing this first codon n-tuple with a second codon n-tuple with a higher relative frequency Frequency can have a particularly beneficial effect on translation and folding and can therefore lead to a particularly strong improvement in the expression of soluble proteins in the target organism.
- the method according to the invention provides for the possibility that the direct successions of the n base triplets of at least two change positions overlap, with the base triplet, which is replaced by the synonymous base triplet, of at least these two change positions is included at the same time.
- the method according to the invention can further provide that the synonymous base triplet is selected so that in one of the two change positions the resulting second codon n-tuple has a lower relative codon n-tuple frequency and in the other of the two Change positions, the resulting second codon n-tuple has a higher relative codon n-tuple frequency than the respective first codon n-tuple.
- the method according to the invention can also provide for such overlapping change positions that the resulting second codon n-tuple in both change positions has a higher relative codon n-tuple frequency than the respective first codon n-tuple . It is still possible that the relative codon n-tuple frequency decreases in both change positions.
- n n base triplets from more than two change positions overlap and the base triplet, which is replaced by the synonymous base triplet, includes more than two change positions at the same time is . It is then possible for the synonymous base triplet to be chosen so that in at least one of the change positions the resulting second codon n-tuple has a lower relative codon n-tuple frequency and in the other of the change positions the resulting second codon n-tuple has a higher relative codon n-tuple frequency than the respective first codon n-tuple.
- the relative codon n-tuple frequency of the second codon n-tuples compared to the respective first codon n-tuples is at least approximately 1%, 5% or 10% and/or at most approximately 40%, 30 % or 20% of the change items is reduced.
- the inventors have further recognized that in order to improve the expression rate of soluble protein from the nucleotide sequence in the target organism, it may be significantly more important to use the method according to the invention to increase particularly low relative codon n-tuple frequencies than isolated or even the majority achieve particularly high relative codon n-tuple frequencies.
- the method according to the invention therefore looks preferred
- the synonymous base triplets are chosen so that the relative codon n-tuples frequency of the second codon n-tuples has the greatest possible minimum value, i.e. H . a greatest possible global minimum is achieved or at least not more than 50%, preferably not more than
- the synonymous base triplets are preferably chosen so that an average of the relative codon n-tuple frequencies of the second codon n-tuples reaches a maximum value or at least not more than 50%, preferably not more than 40% or not more than 30%, preferably not more than 20%, particularly preferably not more than 10% below an achievable maximum value.
- the optimization method can be used to provide nucleotide sequences that are particularly well adapted to expression in the target organism and can therefore be expressed particularly reliably and with high expression rates of soluble protein in the target organism.
- the basic condition according to the invention is that at least one of the n base triplets from the direct succession of the at least one change position is replaced by a synonymous base triplet, which is specifically chosen so that at least a second codon n-tuple results, which is related to the set of amino acid n-tuple events has a higher relative codon n-tuple frequency in the genome or the part thereof of the at least one target organism and / or in the genomes or the parts thereof of the viruses capable of infecting the at least one target organism as the first codon-n-tuple, inevitably always fulfilled when the criteria mentioned are reached.
- the relative frequency of the first codon n-tuples can in principle be used as a control parameter for implementing the method according to the invention or, for example, to compare an intermediate or final result of the optimization with the initial state. Furthermore, in these embodiments, the relative frequency of the first codon n-tuples can remain open, since the goal of the optimization is not based on which initial sequence was started.
- the mean value contains a degressive weighting of the relative codon n-tuple frequencies of the first and second codon n-tuples, which is configured so that a high relative Codon n-tuple frequency has a disproportionate influence on the mean value compared to a lower relative codon n-tuple frequency. whose calculation has .
- an increase in a low relative codon n-tuple frequency can have a greater impact on the mean than an increase in a medium or high relative codon n-tuple frequency.
- the method can additionally or alternatively also be designed so that the relative codon n-tuple frequency reaches a maximum value in at least some of the change positions.
- the natural number n must be the same natural number for the n base triplets, the codon n-tuple and the amino acid n-tuple of a change position.
- the change position e.g. B. comprises a direct sequence of three base triplets, this also codes for a sequence section of three amino acids of the given amino acid sequence, i.e. H .
- the first and second codon n-tuple are each a codon 3-tuple and the amino acid n-tuple is correspondingly an amino acid 3-tuple.
- various change positions e.g. B. at least two change positions, in the number n differ or that n for different change positions, e.g. B.
- n one Change position can be varied during the process, i.e. H .
- the change position can be enlarged or reduced during the process.
- a large change position can be divided into several small change positions and vice versa. This can e.g. B. be advantageous in areas of the nucleotide sequence that encode amino acid n-tuples that occur rarely or not at all in the genome of the target organism or organisms or viruses. in the parts of it, resort to smaller n with more reliable statistics.
- nucleotide sequences that are harmful for expression in the target organism such as. B. Restriction interfaces are very likely to be excluded by using larger n. Since such sequence motifs usually have no basis in the target organism's own genome, i.e. H . If the relative codon n-tuple frequency in the genome of the target organism or organisms approaches zero, the method according to the invention implicitly leads to their systematic exclusion.
- harmful sequences are e.g. B. up to a length of 3n-2 inclusive is automatically excluded from the optimized nucleotide sequence if n is greater than or equal to 3.
- the base triplets are replaced by the synonymous base triplets in several iteration steps using a computer-aided optimization method. In this way, it is particularly possible to successively add nucleotide sequences with a large number of overlapping change positions optimize.
- at least one of the n base triplets can be replaced by a synonymous base triplet in all change positions or only in part of the change positions.
- the at least one of the n base triplets is replaced by a synonymous base triplet in only part of the change positions.
- the base triplet can be comprised of one change position or several overlapping change positions.
- Replacing the base triplets with the synonymous base triplets can, for example, be iterated so often until one of the target criteria already mentioned above is achieved, for example the relative codon n-tuple frequencies of the second codon n-tuples have a maximum possible minimum value, a maximum possible Average, in particular a largest possible weighted average, or reach a maximum value or approach these values in the dimensions defined above.
- the method in these embodiments can also include one or more iteration steps that lead away from the respective target criterion, in particular by determining the relative frequency of the second codon n-tuple in one or more of the change positions compared to the first codon -n- tuples are at least temporarily reduced by an iteration step in order to obtain local maxima of minimum value or To be able to overcome mean values that stand in the way of achieving the target criterion. It is therefore also provided that in one or more of the change positions, the at least one of the n base triplets can be replaced several times by different synonyms until the target criterion is reached Base triplets can be replaced. In particular, there is no provision for determining a change position to a specific second codon n-tuple after an iteration step has been carried out.
- the computer-aided optimization method preferably includes an approximation method, in particular a simulated cooling method, also referred to as “simulated annealing”. These methods are particularly useful for finding an approximate solution for a nucleotide sequence that is optimal for the target organism with regard to the relative codon n-tuple frequencies proven to be suitable and advantageous, especially when longer nucleotide sequences, for example with 30 codons or more, with a large number of overlapping change positions due to their complexity preclude the complete checking of all possible synonymous base triplets and mathematical optimization methods.
- other heuristic approximation methods such as For example, a deluge algorithm or a genetic algorithm is possible. Other suitable approximation methods are known to those skilled in the art.
- other computer-aided optimization methods such as B. artificial intelligence (AI)-based applications are conceivable.
- the change positions together comprise at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40% or at least 50% of the base triplets of the nucleotide sequence that encode an amino acid of the predetermined amino acid sequence .
- the change positions together comprise at least 60%, at least 70% or at least 80%, particularly preferably at least 90% or at least 95% of the base triplets of the nucleotide sequence which contain an amino acid of the predetermined Encoding amino acid sequence.
- the method according to the invention ensures a particularly reliable optimization of the nucleotide sequence for expression in the target organism.
- change positions can also contain base triplets that are not replaced by a synonymous base triplet, i.e. H . It is not necessary that each base triplet included in the change positions is replaced by a synonymous base triplet.
- the synonymous base triplets are preferably chosen so that the relative codon n-tuple frequencies of the second codon n-tuples are at least 10%, at least 15%, at least 20%, at least 25%, at least 30% or at least 35% , preferably in at least 40% or at least 45% of the change positions, preferably in at least 50%, at least 55%, at least 60%, at least 65% or at least 70% of the change positions, particularly preferably in at least 75%, at least 80%, at least 85 % or at least 90% of the change positions have a higher relative codon n-tuple frequency than the respective first codon n-tuples.
- At least one target organism includes the possibility that the nucleotide sequence is optimized for the expression of the specified amino acid sequence in a plurality of different target organisms at the same time.
- the method according to the invention provides for this, for example that the relative codon n-tuple frequency results from the set of events of a codon n-tuple based on the set of events of the corresponding amino acid n-tuple in the genomes or parts thereof of the majority of the different target organisms.
- an optimized nucleotide sequence can be provided using the method according to the invention better for heterologous expression in different
- the at least one target organism can in principle be a predetermined host cell or any predetermined organism that is suitable for the expression of the predetermined amino acid sequence.
- the host cell can be a prokaryotic or a eukaryotic host cell.
- the host cell may be a host cell suitable for culture in liquid or solid media.
- the host cell can be a cell that is part of a multicellular tissue or a multicellular organism such as a plant, an animal or a human, in particular a transgenic one.
- the host cell can be microbial or non-microbial.
- a microbial host cell can be a bacterial, yeast or fungal cell.
- Suitable bacterial host cells include both Gram-positive and Gram-negative bacteria.
- suitable bacterial host cells are bacteria from the genera Bacillus, Actinomycetis, Escherichia, Streptomyces and lactic acid bacteria such as Lactobacillus, Streptococcus, Lactococcus, Oenococcus, Leuconostoc, Pediococcus, Carnobacterium, Propionibacterium, Enterococcus and Bifidobacterium.
- Bacillus subtilis Bacillus amyloliquefaciens
- the host cell can also be a eukaryotic microorganism such as a yeast or a fungus, in particular a filamentous one.
- yeasts as host cells belong to the genera Saccharomyces, Kluyveromyces, Candida, Pichia, Schizosaccharomyces, Hansenula, Kloeckera, Schwanniomyces, and Yarrowia.
- Particularly preferred Debaromyces host cells are Saccharomyces cerevisiae and Kluyveromyces lactis.
- the host cell of the present invention is a cell of a filamentous fungus.
- Filamentous fungi include all filamentous forms of the subdivisions Eumycota and Oomycota.
- Filamentous mushrooms are characterized by a mycelial wall composed of chitin, cellulose, glucan, chitosan, mannan and other complex polysaccharides. Vegetative growth occurs through hyphal elongation and carbon degradation is obligatorily aerobic.
- the filamentous fungi whose strains can be used as host cells in the present invention include, among others, strains of the genera Acremonium, Aspergillus, Aureobasidium, Cryptococcus, Filibasidium, Fusarium Humicola, Magnaporthe, Mucor, Myceliophthora, Neocallimastix, Neurospora, Paecilomyces, Penicillium , Piromyces, Schizophyllum, Chrysosporium, Talaromyces, Thermoascus, Thielavia, Tolypocladium and Trichoderma.
- filamentous fungi are selected from the group consisting of Aspergillus niger, Aspergillus oryzae, Aspergillus sojae, Trichoderma reesei and Penicillium chrysogenum. Examples of suitable host strains are known to those skilled in the art.
- Suitable non-microbial host cells are, for example
- Mammalian host cells such as hamster cells (e.g. Chinese hamster
- CHO ovary
- BHK Baby Hamster Kidney
- Mouse cells monkey cells or human cells or cell lines such as HeLa or HEK293
- Insect cells such as Drosophila cells or Lepidoptera cell lines Hi5, S f21
- Plant cells such as B.
- Non-pathogenic Leishmania are suitable for protein expression.
- Such non-microbial cells are particularly suitable for the production of mammalian or human proteins for use in mammalian or human therapy.
- the predetermined amino acid sequence is preferably a protein or a part thereof, which is in particular naturally a eukaryotic protein.
- the method according to the invention has proven to be particularly advantageous for the expression of eukaryotic proteins, for example an insect, plant or mammalian protein, in bacterial, in particular prokaryotic, expression systems such as Escherichia coli.
- the nucleotide sequence is optimized for the expression of the specified amino acid sequence in a plurality of different target organisms at the same time, it is possible that the different target organisms have very different genome sizes, with large genomes usually having proteomes with a larger number of amino acids. n-tuples encode as smaller genomes. This can lead to the relative codon n-tuple frequencies being disproportionately influenced by large genomes and thus the optimization of the nucleotide sequence inevitably favors expression in target organisms with large genomes more than in target organisms with small genomes.
- the relative codon n-tuple frequency contains a genome-dependent weighting, which is configured, for example, in such a way that a difference in size in the genomes or parts thereof, in particular a different extent of the coding regions, of the different target organisms is at least partially compensated. This ensures that the nucleotide sequence optimized according to the invention is more suitable for expression in the various target organisms.
- the method according to the invention is not only able to significantly improve the quantitative protein yield in the heterologous expression system compared to conventional optimization methods, but in particular also to significantly increase the proportion of soluble protein in the yield.
- This is a particular advantage of the method according to the invention, since the dissolved form of a protein generally represents the native, biochemically active state, which is of central importance, especially in high-value proteins for scientific, medical-pharmaceutical and biotechnological purposes.
- a special characteristic of the method according to the invention is that after the optimization of the nucleotide sequence, the amino acid sequence expressed in the at least one target organism has a greater solubility and/or is present in a larger proportion in dissolved form than before the optimization.
- the optimization of the nucleotide sequence for expression in the at least one target organism can alternatively or in addition to the optimization based on the genome of the target organism or parts thereof also based on genomes or parts thereof of those capable of infecting the at least one target organism
- Viruses occur.
- the viruses may include bacteriophages.
- the optimization is based on a plurality or large number of virus genomes or Sharing of this occurs because a single viral genome or Viral transcriptome generally does not have the required size and therefore does not have the necessary statistical significance for effective optimization of the nucleotide sequence based on the relative codon n-tuple frequency according to the invention.
- the inventors therefore summarize the genomes or Transcriptomes of different viruses are combined to form a type of “supergenome” or “supertranscriptome,” which is used as the basis for determining the relative codon n-tuple frequency.
- the method preferably includes these
- the inventors make use of the knowledge that the genomes of viruses or Phages are already naturally optimized for high-throughput protein expression in the infected target organism.
- the inventors have recognized a particular advantage that viruses or Phage genomes often have a reduced base complexity, which means that an mRNA that is transcribed from a nucleotide sequence optimized according to the invention forms little or no secondary structures.
- the optimization method according to the invention is in practice significantly superior to other methods that use artificial algorithms for mRNA secondary structure optimization.
- the method can include steps that serve to reduce or exclude nucleotide sequences and/or motifs that are present within the nucleotide sequence and/or are randomly generated by replacing base triplets in the change positions and which can adversely affect expression in the target organism.
- such unfavorable nucleotide sequences and/or motifs are at least partially removed from the nucleotide sequence.
- Non-limiting examples of unfavorable nucleotide sequences and/or motifs include cis-acting mRNA destabilizing motifs, RNase splice sites, ribosome binding sites, repetitive elements, and restriction enzyme recognition sequences.
- B restriction enzyme recognition sequences.
- nucleotide sequences that are harmful for expression in the target organism are already inherently excluded for the most part by the method according to the invention, in particular by the length of the codon n-tuples with n greater than or equal to 3.
- an additional step such as: B. Optimization of the mRNA secondary structure or GC content, removal of mRNA destabilizing motifs, ribosome binding sites, repetitive elements and/or recognition sequences of restriction enzymes can be excluded from the process.
- the subject of the present invention is also the use of a method according to the method described above optimized nucleotide sequence for producing synthetic DNA and/or for protein expression in a target organism.
- a further subject of the present invention is a nucleic acid molecule, in particular an isolated one, which comprises an optimized nucleotide sequence which was obtained by one of the methods described here.
- the nucleic acid is DNA.
- a vector is provided which comprises the, in particular isolated, nucleic acid molecule. Nucleic acid molecules optimized according to the invention can be clearly distinguished from conventionally optimized sequences using a sequence comparison. In this regard, reference is also made to the following comparative examples.
- a further subject of the invention is a recombinant host cell which contains the above-mentioned, in particular isolated, nucleic acid molecule or the above-mentioned vector.
- the present invention also relates to a method for expressing a, in particular recombinant, protein in a target organism, which comprises providing a nucleotide sequence which encodes the protein and which is optimized according to the above method.
- the method may comprise one or more of the following steps: synthesizing a nucleic acid molecule comprising the optimized nucleic acid sequence; Introducing the nucleic acid molecule into the target organism; and cultivating the target organism under conditions that enable expression of the protein from the optimized nucleic acid sequence.
- the expression is preferably carried out at least partially at a temperature less than or equal to 30 ° C, less than or equal to 25 ° C or less than or equal to 20 ° C.
- nucleotide sequences optimized according to the invention significantly favor heterologous protein expression at relatively low temperatures compared to conventionally optimized nucleotide sequences.
- the method according to the invention is particularly suitable for the expression of sensitive high-value proteins and at the same time leads to greater sustainability through energy saving potential.
- Another subject of the present invention is a computer program with program code means.
- the program code means of the computer program are set up to carry out a method according to the above description when the computer program is executed on a computer.
- the computer program can include an interface to a DNA and/or RNA synthesis device.
- the subject of the present invention is also a computer-readable storage medium on which the aforementioned computer program is stored in computer-readable form.
- a further subject of the invention is a device for optimizing and/or producing a nucleotide sequence for the expression of a predetermined amino acid sequence in at least one target organism.
- the device has a computing device which is set up to carry out one of the above-mentioned methods.
- the device can in particular be a DNA and/or RNA synthesis device, also referred to as a “DNA/RNA synthesizer”.
- SEQ ID NO:1 Nucleotide sequence of the I. sakaiensis PETase (wild-type sequence) coding for amino acids (aa) 28-290;
- SEQ ID NO: 2 Synthetically produced nucleotide sequence of the I. sakaiensis PETase (encoding for AS 28-290) with a double Strep tag at the C-terminus after conventional optimization for expression in E. coli according to the prior art (reference);
- SEQ ID NO: 5 nucleotide sequence of the A. thaliana OTP86-DYW domain (AS 826-960) (wild-type sequence);
- SEQ ID NO: 6 Synthetically produced nucleotide sequence of the A. thaliana OTP86-DYW domain (AS 826-960) with double Strep tag and Tobacco etch virus (TEV) cleavage site at the N-terminus after conventional optimization for expression in E. coli according to the state of the art (reference);
- SEQ ID NO: 9 Protein sequence of citrine (encoding aa 1-239) as a predetermined amino acid sequence for expression in H. sapiens;
- SEQ ID NO: 10 Synthetically produced nucleotide sequence of citrine (encoding AS 1-239) with FLAG tag and double Strep tag at the N-terminus after conventional optimization for expression in H. sapiens according to the prior art (reference), 5 '- flanked by a Kozak sequence (GCCACC);
- SEQ ID NO: 12 nucleotide sequence (wild type) of the H. sapiens STING1 ER exit protein 1 ("STEEP1") coding for AS 1-222;
- SEQ ID NO: 13 Synthetically produced nucleotide sequence of H. sapiens STEEP1 (encoding aa 1-222) with FLAG tag and double Strep tag at the N-terminus after conventional optimization for expression in H. sapiens according to the state of the art (reference ) , 5'- flanked by a Kozak sequence (GCCACC);
- SEQ ID NO: 15 nucleotide sequence (wild type) of the H. sapiens nitric oxide synthase-interacting protein (NOSIP) coding for AS 1-304;
- SEQ ID NO: 16 Synthetically produced nucleotide sequence from H. sapiens NOSIP (encoding AS 1-304) with FLAG tag and double Strep tag at the N-terminus after conventional optimization using codon adaptation Index and mRNA secondary structure optimization for expression in H. sapiens according to the prior art (reference), 5'-flanked by a Kozak sequence (GCCACC);
- SEQ ID NO: 17 Synthetically produced nucleotide sequence of H. sapiens NOSIP (encoding AS 1-304) with FLAG tag and double Strep tag at the N-terminus after conventional optimization according to WO 2020/024917 Al for expression in Homo sapiens State of the art (reference), 5'-flanked by a Kozak sequence (GCCACC);
- SEQ ID NO: 19 Protein sequence of EqFP611 (AS 1-231) as a specified amino acid sequence for expression in S. elongatus;
- SEQ ID NO: 20 Synthetically produced nucleotide sequence of EqFP611 (encoding aa 1-231) with double Strep tag at the N-terminus after conventional optimization for expression in S. elongatus according to the state of the art (reference), 5 '-flanked by a restriction site Ndel and 3'- flanked by a transcription terminator and a Kpnl restriction site;
- FIG. 1 shows a flowchart with a schematic sequence of an embodiment of the method according to the invention
- Fig. 2 SDS-PAGE (A) and quantitative evaluation (B) of the heterologous expression of Ideonella sakaiensis PETase in Escherichia coli at 20 ° C;
- FIG. 3 SDS-PAGE (A) and quantitative evaluation (B) of the heterologous expression of Ideonella sakaiensis PETase in Escherichia coli at 30 °C;
- Fig. 4 is a graphical representation of the relative
- FIG. 8 Western blot (A) and quantitative evaluation (B) of the expression of H. sapiens STEEP1 in HeLa cells;
- FIG. 9 Western blot (A) and quantitative evaluation (B) of the expression of H. sapiens STEEP1 in HEK293 cells;
- FIG. 10 Western blot (A) and quantitative evaluation (B) of the expression of H. sapiens NOSIP in HeLa cells with SEQ ID NO: 16 (reference) and SEQ ID NO: 18;
- FIG. 11 Western blot (A) and quantitative evaluation (B) of the expression of H. sapiens NOSIP in HeLa cells with SEQ ID NO: 17 (reference) and SEQ ID NO: 18;
- FIG. 12 Western blot (A) and quantitative evaluation (B) of the expression of H. sapiens NOSIP in HEK293 Cells with SEQ ID NO: 16 (reference) and SEQ ID NO: 18.
- Comparative Example 1 Heterologous expression of Ideonella sakaiensis PET hydrolase (PETase) in Escherichia coli
- nucleotide sequence of the PETase from I. sakaiensis which codes for the amino acid positions 28-290 (molecular weight 27.9 kDa) (SEQ ID NO: 1), was used for heterologous expression in E. coli according to the method according to the invention and optimized as a reference according to the method according to WO 2020/024917 Al.
- three pET28a expression plasmids were purchased from Genscript, which encode the amino acid sequence of the PETase with a double Strep tag at the C-terminus under an inducible T7 promoter.
- the amino acid sequence was supplemented N-terminally with Met (start codon) and the amino acids Ala and Ser.
- Fig. 1 shows in this context a schematic sequence of an exemplary implementation of the method 100 according to the invention in a computer-implemented embodiment for optimizing the PETase from I. sakai ensi s.
- the field 102 represents the input of the nucleotide sequence to be optimized into the computer.
- the entire coding sequence was continuously in change positions, i.e. H . each with a codon offset between adjacent change positions.
- N-2 codon-3 tuples or N-l codon 2 tuples based on the total number N of amino acids of the given amino acid sequence, N-2 codon-3 tuples or N-l codon 2 tuples.
- the change positions are thus determined by entering the nucleotide sequence to be optimized.
- E. coli protein-coding genes were created using a DNA sequence database.
- a suitable DNA sequence database is e.g. B. GenBank (Nucleic Acids Research 41, 2013, D36-42).
- the Reference Sequence (RefSeq) database (The NCBI Handbook, 2nd edition, Chapter 18: The Reference Sequence (RefSeq) Database, Bethesda (MD), National Center for Biotechnology Information , USA, 2013).
- the absolute frequency of each combinatorially possible codon n-tuple as well as the absolute frequency of each combinatorially possible amino acid n-tuple within the protein-coding genes determined, where n in one
- unwanted nucleotide sequences such as the TATA box “TATAA” or the ribosomal binding site “AGGAGG”, which are known to those skilled in the art that they can affect expression in E. coli, were entered in field 108.
- Other unwanted sequence motifs included AAAAAA, TTTTT, AGGAGGT, TATAAA, ATCTGTT, GGAGGT and GGGTGGT.
- the base triplets in the change positions of the wild-type sequence were then successively replaced in field 110 in a large number of interaction steps 112 until the relative codon n-tuple frequency of all codon n-tuples Tuple in the nucleotide sequence achieved the largest possible weighted average while minimizing the number of unwanted nucleotide sequences. Only the start codon was excluded from the optimization, although it is in principle possible to also include the start codon and/or stop codon in the optimization.
- the weighted average was additionally offset against an expression for the occurrence of undesirable sequence motifs.
- a value F E was determined, which corresponds to the number of unwanted sequence motifs in the nucleotide sequence to be optimized multiplied by -1.
- the optimized nucleotide sequences SEQ ID NO: 3 and SEQ ID NO: 4 were output in field 114, which were then synthesized accordingly.
- nucleotide sequence optimized according to the invention differs significantly from the wild-type sequence even at the nucleotide level.
- the expression cultures were prepared from 1 mL preculture and 99 mL TB medium.
- the expression of recombinant PETase in the cultures were induced at an OD600 of 0.6 by adding IPTG at a final concentration of 1 mM. Expression took place in one variant at 20 °C for 14 hours and in another variant at 30 °C for five hours.
- the OD600 of the expression cultures was determined and the same amount of cells were harvested from each of the cultures in order to normalize the protein yields based on the cell mass.
- the cell pellets were dissolved in 10 mL of buffer A (20 mM Tris-Cl, pH 7.5, 150 mM NaCl, 1 mM DTT) and disrupted using ultrasound. The cell lysate was then centrifuged at 20,000 g for one hour to separate the insoluble cell components as a pellet.
- the supernatant with the soluble fraction was mixed in an Eppendorf tube with 200 pL of streptactin beads equilibrated in buffer A (IBA Lifesciences, Göttingen, Germany). The beads were washed twice with 1 mL of buffer A in the Eppendorf tube by centrifugation and removing the supernatant. The bound proteins were eluted with 200 pL buffer A containing 10 mM desthiobiotin. The identity of the protein was analytically verified by SDS-polyacrylamide gel electrophoresis (SDS-PAGE). The amount of protein in the SDS-PAGE gel bands was quantified using Image J software (National Institutes of Health, USA). In addition, the protein concentration in the respective supernatants was determined using the Bradford assay (Thermo Fisher Scientific, Bremen, Germany) according to the manufacturer's instructions.
- Figure 2A shows an image of the SDS-PAGE analysis of expression at 20°C.
- Lane 1 contains a size marker
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 2 (reference)
- lane 3 contains the expression product of the plasmid with the inserted SEQ ID NO: 3
- lane 4 contains the expression product of the plasmid with the inserted SEQ ID NO: 4.
- the PETase was successfully expressed by all three plasmids as evidenced by the band at 30.4 kDa, which corresponds to the molecular weight of the protein including the double Strep tag.
- the bar diagram shown in Fig. 2B shows the relative protein yield of soluble PETase depending on the nucleotide sequence used in each case, where the The quantitatively determined amount of protein was normalized to the amount of protein from the reference experiment with SEQ ID NO: 2.
- the hatched columns show the result of the quantification using SDS-PAGE, the white columns show the result of the quantification using the Bradf ord assay.
- Figure 3A shows an image of the SDS-PAGE analysis of expression at 30°C.
- Lane 1 contains a size marker
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 2 (reference)
- lane 3 contains the expression product of the plasmid with the inserted SEQ ID NO: 4.
- the PETase was also identified here from both plasmids the band at 30.4 kDa was successfully expressed. Based on the band strength it can be seen that the plasmid with the nucleotide sequence SEQ ID NO: 4 optimized according to the invention led to a higher protein yield than the plasmid with the reference sequence SEQ ID NO: 2.
- 3B again shows the relative protein yield of soluble PETase depending on the nucleotide sequence used in each case, the quantitatively determined amount of protein being normalized to the amount of protein from the reference experiment with SEQ ID NO: 2.
- the hatched columns show the result of the quantification using SDS-PAGE, the white columns show the result of the quantification using the Bradf ord assay.
- the quantitative analysis shows a more than doubled expression of PETase using the nucleotide sequence SEQ ID NO: 4 optimized according to the invention compared to the sequence optimized according to WO 2020/024917 Al.
- the nucleotide sequence of the OTP86-DYW domain in amino acid positions 826-960 from A. thaliana (SEQ ID NO: 5) was used for heterologous expression in E. coli according to the method according to the invention and according to the method according to WO 2020/024917 Al optimized for reference.
- the OTP86-DYW domain is a sensitive plant protein that is known to be difficult to express in heterologous systems.
- OTP86-DYW domain For the expression of the OTP86-DYW domain in E. coli as a target organism, three pET41 expression plasmids were cloned, which contain the amino acid sequence 826-960 of the OTP86-DYW domain with a double Strep tag and a TEV protease cleavage site as an insert under an inducible T7- Promoter included. A Met (start codon) and a Gly were also added to the N-terminus of the amino acid sequence. The inserts were purchased from Genscript.
- the coding part of the genomes of the following viruses or phages capable of infecting E. coli was taken as a basis:
- the experiment for the sequence optimization according to the invention was carried out essentially as described in Comparative Example 1.
- the relative codon 3 tuple frequencies P were used as the midpoints of the Clopper-Pearson confidence interval with a confidence level of 95 % calculated.
- the number L of codons in the nucleotide sequence included in the optimization is 501 in this example, since the start codon and the subsequent glycine codon were not included in the optimization. Expression was carried out at 17°C with buffer A containing no DTT.
- the sequence listing also makes it clear here that there are already significant differences at the nucleotide level between the wild-type nucleotide sequence and the nucleotide sequences that were optimized according to the method according to the invention.
- nucleotide sequences optimized according to the method according to the invention differs from the sequence optimized according to the prior art.
- the sequence according to the state of the Technology according to WO 2020/024917 Al was optimized as a reference (SEQ ID NO: 6), only has 85.2% identical nucleotides with the optimized sequence according to the method according to the invention using the E. coli genome (SEQ ID NO: 7) and only 74.6% identical nucleotides with the optimized sequence according to the method according to the invention based on the genomes of the viruses or phages capable of infecting E. coli (SEQ ID NO: 8).
- FIG. 4 shows the relative frequency of the first codon 3 tuples in the wild-type sequence SEQ ID NO: 5 (A) and the relative frequency of the second codon 3 tuples in the nucleotide sequence SEQ ID NO: 7 (B) optimized according to the invention OTP86-DYW for the sequence section that corresponds to nucleotide positions 211-330 in the sequence listing.
- the sequence section shown contains codons number 71 to number 110 inclusive of OTP86-DYW.
- Three neighboring codons each form a change position with a codon 3 tuple, with the sequence section shown being continuously divided into 38 change positions or codon 3 tuples, which are referred to here as ni to n 38 . There is an offset of one codon between successive change positions.
- each codon 3 tuple (x-axis) is assigned by a horizontal line the corresponding relative frequency in percent (y-axis) with which the respective codon 3 tuple contains the corresponding amino acid 3 tuple protein-coding genes of E. coli.
- the vertical lines each show the range of the relative frequencies of all codon 3 tuples that are considered for a specific change position and which code for the corresponding amino acid n-tuple of the change position.
- the method according to the invention results in a significant increase in the relative codon 3 tuple frequency took place in a majority of the change positions shown.
- at least one of the base triplets was replaced by a synonymous base triplet;
- two or three base triplets were replaced by a synonymous base triplet in order to optimally increase the relative frequency of the second codon 3 tuples.
- the second codon 3 tuple correspond to the codon 3 tuple with the greatest relative frequency in E. coli.
- a second codon 3 tuple is formed in the change position n 33 , which has a lower relative frequency than the original first codon 3 tuple, in order to be able to increase the relative frequencies of the more critical first codon 3 tuples in the other change positions . In this way, the largest possible weighted average of the relative codon 3 tuple frequencies could be achieved.
- Fig. 5 shows an image of SDS-PAGE analysis from expression at 17 °C.
- Lane 1 contains a size marker
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 6 (reference)
- lane 3 contains the expression product of the plasmid with the inserted SEQ ID NO: 7
- lane 4 contains the expression product of the plasmid with the inserted SEQ ID NO: 8.
- the OTP86-DYW domain was successfully expressed by all three plasmids as evidenced by the band at 19.5 kDa.
- Fig. 5B shows the relative protein yield of soluble OTP86-DYW domain depending on the nucleotide sequence used in each case, the quantitatively determined amount of protein being normalized to the amount of protein from the reference experiment with SEQ ID NO: 6.
- the hatched columns show the result of the quantification using SDS-PAGE, the white columns show the result of the quantification using photometric UV absorption measurement at 260 and 280 nm (Nanodrop).
- the nucleotide sequence of the fluorescent protein citrine (SEQ ID NO: 9) was used for heterologous expression in human HeLa cells as a target organism according to the method according to the invention and according to a conventional methods based on the codon adaptation index and local mRNA secondary structure optimization as a reference.
- Citrine is a variant of the green fluorescent protein (GFP) from Aquaeoria vi ctoria and is commonly used for reporter assays and in fluorescence microscopy.
- GFP green fluorescent protein
- citrine in HeLa cells two pTwist CMV expression plasmids were purchased from Twist Bioscience (San Francisco, CA, USA), which contain the amino acid sequence of citrine with a FLAG tag followed by a double Strep tag at the N-terminus encode a constitutive cytomegalovirus promoter. The amino acid sequence was supplemented N-terminally with Met (start codon) and the amino acid Ala. A Kozak sequence was inserted before the start codon.
- One of the plasmids contained the citrine-encoding nucleotide sequence with FLAG and double Strep tag after optimization by the manufacturer's method (SEQ ID NO: 10) as a reference.
- nucleotide sequence optimized according to the invention differs significantly from the sequence optimized according to the prior art (SEQ ID NO: 10) at the nucleotide level.
- HeLa cells were cultured 24 hours before transfection in 6-well plates with DMEM high glucose medium (Biowest SAS, Nuaille, France) with 10% FCS (Biochrom AG - Berlin, Germany) and 1% penicillin/ Streptomycin (Biowest). The transfections were carried out with 2 pg plasmid and Rotifect (Carl Roth GmbH, Düsseldorf, Germany) according to the manufacturer's instructions. 70 hours after transfection, the medium was removed and the cells were treated with 1 mL washed with ice-cold phosphate-buffered saline (PBS) and resuspended in RIPA lysis buffer.
- PBS ice-cold phosphate-buffered saline
- the lysates were mixed with 6x SDS loading buffer and separated by size on a 15% SDS polyacrylamide gel.
- the protein samples on the gel were then transferred to a nitrocellulose membrane using Western blotting.
- the nonspecific binding sites of the membrane were blocked with 2% BSA and the membrane was incubated overnight with the primary antibodies against the FLAG-tagged expressed target protein or the housekeeping gene GAPDH (loading control).
- the membrane was washed with TBS Tween and incubated with horseradish peroxidase (HRP)-coupled secondary antibody against rabbit (FLAG) or mouse (GAPDH).
- HRP horseradish peroxidase
- the proteins were visualized using the ECL kit (Pierce, Waltham, MA, USA) and the bands were quantified using ImageQuantTL (Cytiva, Marlborough, MA, USA).
- ImageQuantTL Cosmetic, Marlborough, MA, USA.
- the band strength of citrine was set in the respective ratio to the band strength of the loading control GAPDH in order to normalize the amount of protein applied in relation to the amount of cells.
- the cell lysates were centrifuged for two minutes at 13,000 g and the fluorescence of citrine in the supernatant was measured in triplicates in a Tecan Spark Plate Reader at an excitation wavelength of 516 nm and an emission wavelength of 529 nm.
- the intensity of the fluorescence was in turn set in relation to the respective band intensity of the loading controls (GAPDH) in order to take into account the different cell densities of the cultures.
- GPDH band intensity of the loading controls
- Figure 6A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged citrine and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the inserted SEQ ID NO: 10 (reference)
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 11, which was optimized according to the invention.
- Citrine was successfully expressed by both plasmids as evidenced by the bands stained by the specific HRP-coupled secondary antibody.
- the relative protein yield of citrine is shown depending on the nucleotide sequence used, with the quantitatively determined amount of protein normalized with respect to the cell amount using the loading control GAPDH as an internal standard and based on the amount of protein from the reference experiment with SEQ ID NO: 10 was standardized.
- Fig. 7A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged citrine and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the inserted SEQ ID NO: 10 (reference)
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 11, which was optimized according to the invention.
- Citrine was also successfully expressed in HEK293 cells from both plasmids as evidenced by the bands stained by the specific HRP-coupled secondary antibody.
- the bar diagram shown in Fig. 7B shows the relative protein yield of citrine depending on the nucleotide sequence used, with the quantitatively determined amount of protein normalized with respect to the cell amount using the loading control GAPDH as an internal standard and based on the amount of protein from the reference experiment with SEQ ID NO: 10 was standardized.
- the nucleotide sequence of the H. sapi ens protein STEEP1 was optimized for expression in HeLa cells using the method according to the invention and in a conventional manner according to Comparative Example 3 as a reference.
- STEEP1 is a human protein found in the endoplasmic reticulum membrane. Mutations in STEEP1 are responsible for several diseases.
- Proteins from H. sapi ens are generally difficult to express in a host system.
- One of the plasmids contained the STEEP1-encoding nucleotide sequence with FLAG and double Strep tag after conventional optimization by the manufacturer (SEQ ID NO: 13).
- sequence listing shows significant differences between the wild-type nucleotide sequence (SEQ ID NO: 12) and the nucleotide sequence optimized according to the method according to the invention (SEQ ID NO: 14).
- FIG. 8A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged STEEP1 and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the inserted SEQ ID NO: 13 (reference)
- lane 2 contains the expression product of the plasmid with the SEQ ID NO: 14 optimized according to the invention.
- STEEP1 was coupled to both plasmids as shown by the specific HRP Secondary antibody stained bands were successfully expressed.
- Fig. 8B shows the relative protein yield of STEEP1 depending on the nucleotide sequence used in each case, the quantitatively determined protein amount being normalized to the cell amount using the loading control as described above and related to the protein amount from the reference experiment with SEQ ID NO: 13 .
- the hatched columns show the result of the quantification based on the band intensity of the Western blot.
- Comparative example 5 used optimized nucleotide sequences of STEEP1 expressed in HEK293 cells. Otherwise, the test was carried out essentially as described in Comparative Example 5.
- Fig. 9A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged STEEP1 and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the inserted SEQ ID NO: 13 (reference)
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 14 optimized according to the invention.
- STEEP1 was also successfully expressed in HEK293 cells from both plasmids, as evidenced by the bands stained by the specific HRP-coupled secondary antibody.
- Fig. 9B shows the relative protein yield of STEEP1 depending on the nucleotide sequence used in each case, the quantitatively determined protein amount being normalized to the cell amount using the loading control as described above and related to the protein amount from the reference experiment with SEQ ID NO: 13 .
- the hatched columns show the result of the quantification based on the band intensity of the Western blot.
- Comparative Example 7 Expression of H. sapi ens Nitric oxide synthase-interacting protein (NOS IP) in Heia cell culture
- nucleotide sequence of the H. sapi ens protein NOS IP was optimized for expression in HeLa using the method according to the invention and using the prior art as a reference.
- the reference optimizations were carried out in a variant according to comparative example 3 and in a second variant according to WO 2020/024917 Al.
- NOS IP modulates the activity and localization of nitrite oxide synthase, thereby regulating nitrite oxide production, which is crucial for the development of the human brain, eye and face.
- NOS IP for the expression of NOS IP in HeLa cells as a target organism, three pTwist CMV expression plasmids were purchased from Twist Bioscience, which contain the amino acid sequence of NOS IP with a FLAG tag followed by a double Strep tag at the N-terminus under a constitutive cytomegalovirus Promoter encode. The amino acid sequence was supplemented N-terminally with Met (start codon) and the amino acid Ala. A Kozak sequence was inserted before the start codon.
- One of the plasmids contained the NOS IP coding nucleotide sequence with FLAG and double Strep tag after optimization according to the state of the art by the manufacturer's optimization service as a reference (SEQ ID NO: 16).
- the second plasmid contained the NOSIP-encoding nucleotide sequence with FLAG and double Strep tag after optimization according to WO 2020/024917 A1 as a further reference (SEQ ID NO: 17).
- sequence listing also makes it clear here that there are already significant differences at the nucleotide level between the wild-type nucleotide sequence and the nucleotide sequence that was optimized according to the method according to the invention.
- nucleotide sequence optimized according to the invention also clearly differs from the sequences optimized according to the prior art (SEQ ID NO: 16, SEQ ID NO: 17) at the nucleotide level.
- Fig. 10 shows the comparison between SEQ ID NO: 16 and SEQ ID NO: 18.
- Fig. 10A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged NOSIP and stained with the HRP-coupled secondary antibody for the charging control GAPDH.
- Lane 1 contains the expression product of the plasmid with the inserted SEQ ID NO: 16 (Reference)
- lane 2 contains the expression product of the plasmid with the inserted SEQ ID NO: 18, which was optimized according to the invention.
- nucleotide sequence SEQ ID NO: 18 optimized according to the invention compared to the plasmid with the conventionally optimized one Nucleotide sequence SEQ ID NO: 16.
- Fig. 10B shows the relative protein yield of NOS IP depending on the nucleotide sequence used, whereby the quantitatively determined protein amount is first normalized to the cell amount using the loading control GAPDH as described above and then to the protein amount from the reference experiment with SEQ ID NO : 16 was obtained.
- the hatched columns show the result of the quantification based on the band intensity of the Western blot.
- Fig. 11 shows the comparison between SEQ ID NO: 17 and SEQ ID NO: 18.
- Fig. 11A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged NOS IP and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the reference sequence SEQ ID NO: 17 optimized according to WO 2020/024917 A1
- lane 2 contains the expression product of the plasmid with the sequence SEQ ID NO: 18 optimized according to the invention.
- NOS IP was successfully expressed by both plasmids as evidenced by the bands stained by the specific HRP-coupled secondary antibody. It can already be seen from the band strength that the plasmid with the nucleotide sequence SEQ ID NO: 18 optimized according to the invention led to an increased protein yield compared to the plasmid with the conventionally optimized nucleotide sequence SEQ ID NO: 17.
- Fig. 11B shows the relative protein yield of NOS IP depending on the respective nucleotide sequence used, whereby the quantitatively determined protein amount is first normalized to the cell amount using the loading control GAPDH as described above and then to the protein amount from the reference experiment with SEQ ID NO : 17 was obtained.
- Fig. 12A shows an image of the Western blot stained with the HRP-coupled secondary antibody for the FLAG-tagged NOS IP and stained with the HRP-coupled secondary antibody for the loading control GAPDH.
- Lane 1 contains the expression product of the plasmid with the conventionally optimized SEQ ID NO: 16 (reference)
- lane 2 contains the expression product of the plasmid with the SEQ ID NO: 18 optimized according to the invention.
- NOS IP could also be successfully expressed in HEK293 cells with both plasmids, as evidenced by the bands colored by the specific HRP-coupled secondary antibody, with the band strength showing a significantly better expression of the nucleotide sequence SEQ ID NO: 18 optimized according to the invention.
- the relative protein yield of NOS IP is shown depending on the nucleotide sequence used in each case, with the quantitatively determined protein amount, as described above, first being normalized to the cell amount using the loading control GAPDH and then to the protein amount from the reference experiment with SEQ ID NO : 16 was obtained.
- the hatched columns show the result of the quantification based on the band intensity of the Western blot.
- eqFP611 is a red fluorescent protein (RFP) from Entacmaea quadricolor and is commonly used for reporter assays and in fluorescence microscopy.
- RFP red fluorescent protein
- eqFP611 in S . el ongatus As the target organism, two pSyn-6 expression plasmids (Thermo Fisher, Waltham, MA, USA) were purchased from Genscript, which encodes the amino acid sequence of eqFP611 with a double Strep tag at the N-terminus under a constitutive psbAl promoter.
- Genscript Thermo Fisher, Waltham, MA, USA
- One of the plasmids contained the eqFP611-encoding nucleotide sequence with a double Strep tag after optimization according to WO 2020/024917 A1 (SEQ ID NO: 20).
- sequence listing shows that the nucleotide sequence optimized according to the invention for the heterologous expression of eqFP611 in S. elongatus (SEQ ID NO: 21) differs significantly from the sequence optimized according to the prior art (SEQ ID NO: 20) at the nucleotide level.
- An analog comparison calculation resulted in a weighted average of the relative codon 3 tuple frequencies for the reference sequence according to WO 2020/024917 Al (SEQ ID NO:20) of F w -496.2.
- eqFP611 For the expression of eqFP611 in S. elongatus, the manufacturer's protocol of the "GeneArt algal protein expression system" (Thermo Fisher) was followed. Further treatment of the cyanobacterial biomass is carried out as described in Comparative Example 1 for E. coli. The expressed eqFP611 protein is purified via streptactin and quantified by SDS-PAGE. Additionally, fractions of the eqFP611 cell lysates are centrifuged for two minutes at 13,000 g and the eqFP611 fluorescence in the supernatant is measured in a Tecan Spark Plate Reader at an excitation wavelength of 559 nm and an emission wavelength of 611 nm in triplicates.
- the invention is not limited to these by the description based on the exemplary embodiments. Rather, the invention includes every new feature and every combination of features, which in particular includes every combination of features in the patent claims and the description, even if this feature or this combination of features itself is not explicitly stated in the patent claims, the description or the exemplary embodiments.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Medical Informatics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Theoretical Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Genetics & Genomics (AREA)
- Chemical & Material Sciences (AREA)
- Molecular Biology (AREA)
- Bioethics (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Biochemistry (AREA)
- Public Health (AREA)
- Software Systems (AREA)
- Epidemiology (AREA)
- Data Mining & Analysis (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Wood Science & Technology (AREA)
- General Engineering & Computer Science (AREA)
- Organic Chemistry (AREA)
- Biomedical Technology (AREA)
- Zoology (AREA)
- Library & Information Science (AREA)
- Plant Pathology (AREA)
- Microbiology (AREA)
- Analytical Chemistry (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Preparation Of Compounds By Using Micro-Organisms (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE102022118459.5A DE102022118459A1 (de) | 2022-07-22 | 2022-07-22 | Verfahren zur optimierung einer nukleotidsequenz für die expression einer aminosäuresequenz in einem zielorganismus |
| PCT/EP2023/070275 WO2024018050A1 (de) | 2022-07-22 | 2023-07-21 | Verfahren zur optimierung einer nukleotidsequenz durch austausch synonymer codons für die expression einer aminosäuresequenz in einem zielorganismus |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4558990A1 true EP4558990A1 (de) | 2025-05-28 |
Family
ID=87554859
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23750552.4A Pending EP4558990A1 (de) | 2022-07-22 | 2023-07-21 | Verfahren zur optimierung einer nukleotidsequenz durch austausch synonymer codons für die expression einer aminosäuresequenz in einem zielorganismus |
Country Status (8)
| Country | Link |
|---|---|
| EP (1) | EP4558990A1 (de) |
| JP (1) | JP2025530875A (de) |
| KR (1) | KR20250044702A (de) |
| CN (1) | CN119631132A (de) |
| AU (1) | AU2023310935A1 (de) |
| DE (1) | DE102022118459A1 (de) |
| IL (1) | IL318519A (de) |
| WO (1) | WO2024018050A1 (de) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE102022118459A1 (de) | 2022-07-22 | 2024-01-25 | Proteolutions UG (haftungsbeschränkt) | Verfahren zur optimierung einer nukleotidsequenz für die expression einer aminosäuresequenz in einem zielorganismus |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| DE10260805A1 (de) | 2002-12-23 | 2004-07-22 | Geneart Gmbh | Verfahren und Vorrichtung zum Optimieren einer Nucleotidsequenz zur Expression eines Proteins |
| US20070298503A1 (en) | 2006-05-04 | 2007-12-27 | Lathrop Richard H | Analyzing traslational kinetics using graphical displays of translational kinetics values of codon pairs |
| WO2008000632A1 (en) | 2006-06-29 | 2008-01-03 | Dsm Ip Assets B.V. | A method for achieving improved polypeptide expression |
| CN106650307B (zh) | 2016-09-21 | 2019-04-05 | 武汉伯远生物科技有限公司 | 一种基于密码子对使用频度的基因密码子优化方法 |
| US11848074B2 (en) | 2016-12-07 | 2023-12-19 | Gottfried Wilhelm Leibniz Universität Hannover | Codon optimization |
| TWI802728B (zh) | 2018-07-30 | 2023-05-21 | 大陸商南京金斯瑞生物科技有限公司 | 密碼子優化方法、包括其之系統及電子裝置、其核酸分子及使用其之蛋白質表現方法 |
| KR20230020991A (ko) | 2020-05-07 | 2023-02-13 | 트랜슬레이트 바이오 인코포레이티드 | 최적화된 뉴클레오티드 서열의 생성 |
| DE102022118459A1 (de) | 2022-07-22 | 2024-01-25 | Proteolutions UG (haftungsbeschränkt) | Verfahren zur optimierung einer nukleotidsequenz für die expression einer aminosäuresequenz in einem zielorganismus |
-
2022
- 2022-07-22 DE DE102022118459.5A patent/DE102022118459A1/de active Pending
-
2023
- 2023-07-21 CN CN202380055403.3A patent/CN119631132A/zh active Pending
- 2023-07-21 AU AU2023310935A patent/AU2023310935A1/en active Pending
- 2023-07-21 EP EP23750552.4A patent/EP4558990A1/de active Pending
- 2023-07-21 KR KR1020257005749A patent/KR20250044702A/ko active Pending
- 2023-07-21 JP JP2025526880A patent/JP2025530875A/ja active Pending
- 2023-07-21 IL IL318519A patent/IL318519A/en unknown
- 2023-07-21 WO PCT/EP2023/070275 patent/WO2024018050A1/de not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| AU2023310935A1 (en) | 2025-03-06 |
| CN119631132A (zh) | 2025-03-14 |
| KR20250044702A (ko) | 2025-04-01 |
| DE102022118459A9 (de) | 2024-03-28 |
| IL318519A (en) | 2025-03-01 |
| WO2024018050A1 (de) | 2024-01-25 |
| DE102022118459A1 (de) | 2024-01-25 |
| WO2024018050A9 (de) | 2024-08-22 |
| JP2025530875A (ja) | 2025-09-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| AT501955B1 (de) | Mutierte aox1-promotoren | |
| EP3321365B1 (de) | Neue aus pflanzen stammende cis-regulatorische elemente für die entwicklung pathogen-responsiver chimärer promotoren | |
| EP0616035A2 (de) | Transgener Pathogen-resistenter Organismus | |
| EP4558990A1 (de) | Verfahren zur optimierung einer nukleotidsequenz durch austausch synonymer codons für die expression einer aminosäuresequenz in einem zielorganismus | |
| DE3751100T2 (de) | Erhöhte Produktion von Proteinen in Bakterien durch Verwendung einer neuen Ribosombindungsstelle. | |
| Cortes et al. | Trichoderma harzianum marker-free strain construction based on efficient CRISPR/Cas9 recyclable system: A helpful tool for the study of biological control agents | |
| DE60219632T2 (de) | Verfahren zur herstellung von matrizen-dna und verfahren zur proteinherstellung in einem zellfreien proteinsynthesesystem unter verwendung davon | |
| EP1891220B1 (de) | Autoaktiviertes resistenzprotein | |
| CH640268A5 (en) | Process for the preparation of filamentous hybrid phages, novel hybrid phages and their use | |
| DE10252245A1 (de) | Verfahren zur Expression und Sekretion von Proteinen mittels der nicht-konventionellen Hefe Zygosaccharomyces bailii | |
| EP1504103B1 (de) | Promotoren mit veränderter transkriptionseffizienz aus der methylotrophen hefe hansenula polymorpha | |
| DE69435058T2 (de) | Multicloning-vektor, expressionsvektor und herstellung von fremdproteinen unter verwendung des expressionsvektors | |
| EP1570062B1 (de) | Optimierte proteinsynthese | |
| DE10205091B4 (de) | Verfahren zur Vorhersage der Expressionseffizienz in zellfreien Expressionssystemen | |
| EP1280893B1 (de) | Verfahren zum herstellen von proteinen in einer hefe der gattung arxula und dafür geeignete promotoren | |
| DE102010016387A1 (de) | Verfahren zur Herstellung süßer Proteine | |
| EP1235906B8 (de) | Verfahren zur mutagenese von nukleotidsequenzen aus pflanzen, algen, oder pilzen | |
| WO1999064567A1 (de) | Transformierte zell-linien, die heterologe g-protein-gekoppelte rezeptoren exprimieren | |
| WO2001005976A1 (de) | Mutiertes ribosomales protein l3 | |
| WO2023006995A1 (de) | Collinolacton-biosynthese und herstellung | |
| WO2000012748A1 (de) | Organismen zur extrazellulären herstellung von riboflavin | |
| WO2004076672A2 (de) | Neuer dominanter selektionsmarker zur transformation von pilzen | |
| WO1996017068A2 (de) | Pathogenresistente pflanzen und verfahren zu ihrer herstellung | |
| DD296107A5 (de) | Herstellungsverfahren fuer expressionsplasmide zur auspraegung reifer genprodukte in bakteriellen wirten | |
| DD261503A3 (de) | Verfahren zur Herstellung von Hefe-Vektoren des Yep-Typs |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250211 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20251210 |